REVIEW 2 cited by
XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language models are ubiquitous in current NLP, and their multilingual capacity has recently attracted considerable attention. However, current analyses have almost exclusively focused on (multilingual variants of) standard benchmarks, and have relied on clean pre-training and task-specific corpora as multilingual signals. In this paper, we introduce XLM-T, a model to train and evaluate multilingual language models in Twitter. In this paper we provide: (1) a new strong multilingual baseline consisting of an XLM-R (Conneau et al. 2020) model pre-trained on millions of tweets in over thirty languages, alongside starter code to subsequently fine-tune on a target task; and (2) a set of unified sentiment analysis Twitter datasets in eight different languages and a XLM-T model fine-tuned on them.
Forward citations
Cited by 2 Pith papers
-
Speech Emotion Recognition via Entropy-Aware Score Selection
Entropy and varentropy thresholds on a wav2vec2 emotion model trigger a fallback to Whisper plus RoBERTa sentiment, yielding small average F1 gains on IEMOCAP and MSP-IMPROV.
-
Analyzing public sentiment to gauge key stock events and determine volatility in conjunction with time and options premiums
A claim that LightGBM plus social sentiment predicts stock direction around earnings with 70.1 percent accuracy is undermined by unspecified labels, potential look-ahead bias, and no released artifacts.
Discussion (0). Continue with ORCID to comment.