REVIEW 13 cited by
Goldfish: Monolingual Language Models for 350 Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite state-of-the-art performance on reasoning tasks, we find that these models still struggle with basic grammatical text generation in many languages. First, large multilingual models perform worse than bigrams for many languages (e.g. 24% of languages in XGLM 4.5B; 43% in BLOOM 7.1B) using FLORES perplexity as an evaluation metric. Second, when we train small monolingual models with only 125M parameters on 1GB or less data for 350 languages, these small models outperform large multilingual models both in perplexity and on a massively multilingual grammaticality benchmark. To facilitate future work on low-resource language modeling, we release Goldfish, a suite of over 1,000 small monolingual language models trained comparably for 350 languages. These models represent the first publicly-available monolingual language models for 215 of the languages included.
Forward citations
Cited by 13 Pith papers
-
Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
A study of 100,000 tetun.org logs shows users translate high-resource to Tetun mostly for education and science, a domain mix almost opposite to available Tetun corpora.
-
Syntactic Belief Update as the Driver of Garden Path Processing Difficulty
Syntactic belief update via generalized Rényi divergence on syntactic trees predicts garden path reading times better than lexical surprisal.
-
Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right
An empirical study showing that smaller monolingual language models are more vulnerable to adversarial attacks than larger multilingual models across 70 languages, though multilinguality alone does not guarantee security.
-
Ask a Local: Detecting Hallucinations With Specialized Model Divergence
A multilingual hallucination detector that flags words where a language-specialized model's perplexity diverges from the rest, achieving IoU around 0.3.
-
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
A Unicode script and category based character encoding with constrained merging achieves compression competitive with byte-level BPE while removing the byte-premium penalty for non-Latin scripts.
-
Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility
Language models fine-tuned to predict speech reductions and prosodic prominences from text perform above random baselines, and models pretrained on conversational data outperform those pretrained on written data in En...
-
Why do language models perform worse for morphologically complex languages?
A language-modeling performance gap between agglutinative and fusional languages largely disappears when training data is measured and scaled in bytes rather than tokens.
-
On the Limits of Model Merging for Multilinguality in Pre-Training
Merging any combination of monolingual pre-trained models leads to performance collapse due to interference, indicating that merging flexibility from fine-tuning does not extend to pre-training.
-
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
AI infrastructure systematically disadvantages Bengali speakers through four compounding structural barriers: web presence, training tokens, tokenization, and connectivity.
-
Multi-granular Training Strategies for Robust Multi-hop Reasoning Over Noisy and Heterogeneous Knowledge Sources
AMKOR is described as a state-of-the-art multi-hop QA system, but the paper provides no reproducible evidence and the reported numbers appear unverifiable.
-
Generalization of Medical Large Language Models through Cross-Domain Weak Supervision
A claimed curriculum-based fine-tuning framework for medical LLMs reports better question answering and response generation, but lacks reproducible evidence.
-
Weak Supervision Dynamic KL-Weighted Diffusion Models Guided by Large Language Models
A vague proposal for LLM-guided diffusion with dynamic KL weighting, backed by unsupported FID/IS tables.
-
Cross-Cultural Fashion Design via Interactive Large Language Models and Diffusion Models
The authors claim that LLM prompt refinement plus a CLIP-based weak supervision filter improves diffusion-based fashion image generation, but the evidence is unverifiable and internally inconsistent.
Discussion (0). Continue with ORCID to comment.