Pith. sign in

REVIEW 1 cited by

RoBERTuito: a pre-trained language model for social media text in Spanish

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.09453 v3 pith:MX37NWQR submitted 2021-11-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelslanguagetasksmodelrobertuitopre-trainedspanishtext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since BERT appeared, Transformer language models and transfer learning have become state-of-the-art for Natural Language Understanding tasks. Recently, some works geared towards pre-training specially-crafted models for particular domains, such as scientific papers, medical documents, user-generated texts, among others. These domain-specific models have been shown to improve performance significantly in most tasks. However, for languages other than English such models are not widely available. In this work, we present RoBERTuito, a pre-trained language model for user-generated text in Spanish, trained on over 500 million tweets. Experiments on a benchmark of tasks involving user-generated text showed that RoBERTuito outperformed other pre-trained language models in Spanish. In addition to this, our model achieves top results for some English-Spanish tasks of the Linguistic Code-Switching Evaluation benchmark (LinCE) and has also competitive performance against monolingual models in English tasks. To facilitate further research, we make RoBERTuito publicly available at the HuggingFace model hub together with the dataset used to pre-train it.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A confidence-based variant of Datamaps ranks common Spanish variety examples above random baseline, and a new 1,762-tweet Cuban Spanish dataset is presented.

Pith tools