Pith. sign in

REVIEW 13 cited by

Goldfish: Monolingual Language Models for 350 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.10441 v3 pith:Q654GWMP submitted 2024-08-19 cs.CL

classification cs.CL
keywords modelslanguageslanguagemanymonolingualmultilinguallargesmall
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite state-of-the-art performance on reasoning tasks, we find that these models still struggle with basic grammatical text generation in many languages. First, large multilingual models perform worse than bigrams for many languages (e.g. 24% of languages in XGLM 4.5B; 43% in BLOOM 7.1B) using FLORES perplexity as an evaluation metric. Second, when we train small monolingual models with only 125M parameters on 1GB or less data for 350 languages, these small models outperform large multilingual models both in perplexity and on a massively multilingual grammaticality benchmark. To facilitate future work on low-resource language modeling, we release Goldfish, a suite of over 1,000 small monolingual language models trained comparably for 350 languages. These models represent the first publicly-available monolingual language models for 215 of the languages included.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service

    cs.CL 2024-11 conditional novelty 7.0 of 10

    A study of 100,000 tetun.org logs shows users translate high-resource to Tetun mostly for education and science, a domain mix almost opposite to available Tetun corpora.

  2. Syntactic Belief Update as the Driver of Garden Path Processing Difficulty

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Syntactic belief update via generalized Rényi divergence on syntactic trees predicts garden path reading times better than lexical surprisal.

  3. Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An empirical study showing that smaller monolingual language models are more vulnerable to adversarial attacks than larger multilingual models across 70 languages, though multilinguality alone does not guarantee security.

  4. Ask a Local: Detecting Hallucinations With Specialized Model Divergence

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A multilingual hallucination detector that flags words where a language-specialized model's perplexity diverges from the rest, achieving IoU around 0.3.

  5. BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A Unicode script and category based character encoding with constrained merging achieves compression competitive with byte-level BPE while removing the byte-premium penalty for non-Latin scripts.

  6. Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Language models fine-tuned to predict speech reductions and prosodic prominences from text perform above random baselines, and models pretrained on conversational data outperform those pretrained on written data in En...

  7. Why do language models perform worse for morphologically complex languages?

    cs.CL 2024-11 conditional novelty 6.0 of 10

    A language-modeling performance gap between agglutinative and fusional languages largely disappears when training data is measured and scaled in bytes rather than tokens.

  8. On the Limits of Model Merging for Multilinguality in Pre-Training

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    Merging any combination of monolingual pre-trained models leads to performance collapse due to interference, indicating that merging flexibility from fine-tuning does not extend to pre-training.

  9. Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

    cs.CL 2026-08 conditional novelty 3.0 of 10

    AI infrastructure systematically disadvantages Bengali speakers through four compounding structural barriers: web presence, training tokens, tokenization, and connectivity.

  10. Multi-granular Training Strategies for Robust Multi-hop Reasoning Over Noisy and Heterogeneous Knowledge Sources

    cs.CL 2025-02 reject novelty 2.0 of 10

    AMKOR is described as a state-of-the-art multi-hop QA system, but the paper provides no reproducible evidence and the reported numbers appear unverifiable.

  11. Generalization of Medical Large Language Models through Cross-Domain Weak Supervision

    cs.CL 2025-02 reject novelty 2.0 of 10

    A claimed curriculum-based fine-tuning framework for medical LLMs reports better question answering and response generation, but lacks reproducible evidence.

  12. Weak Supervision Dynamic KL-Weighted Diffusion Models Guided by Large Language Models

    cs.CL 2025-02 reject novelty 2.0 of 10

    A vague proposal for LLM-guided diffusion with dynamic KL weighting, backed by unsupported FID/IS tables.

  13. Cross-Cultural Fashion Design via Interactive Large Language Models and Diffusion Models

    cs.CL 2025-01 reject novelty 2.0 of 10

    The authors claim that LLM prompt refinement plus a CLIP-based weak supervision filter improves diffusion-based fashion image generation, but the evidence is unverifiable and internally inconsistent.

Pith tools