Pith. sign in

REVIEW 46 cited by

The State and Fate of Linguistic Diversity and Inclusion in the NLP World

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.09095 v3 pith:ZOBZA6VH submitted 2020-04-20 cs.CL

The State and Fate of Linguistic Diversity and Inclusion in the NLP World

classification cs.CL
keywords languagelanguagesworlddiversitylinguisticresourcestechnologiesagnostic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Language technologies contribute to promoting multilingualism and linguistic diversity around the world. However, only a very small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications. In this paper we look at the relation between the types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time. Our quantitative investigation underlines the disparity between languages, especially in terms of their resources, and calls into question the "language agnostic" status of current models and systems. Through this paper, we attempt to convince the ACL community to prioritise the resolution of the predicaments highlighted here, so that no language is left behind.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 46 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph

    cs.CL 2026-06 unverdicted novelty 7.0

    Cortex uses an Ontological Corpus Graph to structure web-scale corpora, creating a refined 24.14B-token corpus and a new benchmark validated on eight LLMs.

  2. OpenBibleTTS: Large-Scale Speech Resources and TTS Models for Low-Resource Languages

    cs.CL 2026-06 accept novelty 7.0

    OpenBibleTTS supplies speech data and alignments for 37 underrepresented languages and shows that no single TTS system leads on all metrics, with Gemini-TTS highest in listener ratings but monolingual EveryVoice model...

  3. Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models

    cs.CL 2026-06 unverdicted novelty 7.0

    Evaluation across 1.1 million instances shows sycophancy rates spike in low-resource languages, remain topic-agnostic, and correlate with tokenizer fertility.

  4. Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation

    cs.CL 2026-06 unverdicted novelty 7.0

    RL with chrF reward trains LLMs to better utilize in-context linguistic knowledge for zero-shot translation of unseen languages, outperforming ICL and SFT.

  5. Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages

    cs.CL 2026-06 unverdicted novelty 7.0

    MIDI is a new multilingual idiom dataset with sentence and conversational contexts; benchmarking reveals worse performance in low-resource languages and on literal vs. figurative uses.

  6. Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them

    cs.LG 2026-05 conditional novelty 7.0

    Repetition rate mismatch between small-scale proxies and target budgets is the main reason data mixture experiments do not scale; a subsampling procedure that equalizes repetition rates recovers optimal mixtures from ...

  7. Scaling Laws for Mixture Pretraining Under Data Constraints

    cs.LG 2026-05 conditional novelty 7.0

    Repetition-aware scaling laws show scarce target data in pretraining mixtures can be repeated 15-20 times optimally, with the best count depending on data size, compute, and model scale.

  8. Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion

    cs.CL 2026-04 unverdicted novelty 7.0

    TriMix dynamically fuses logits from three model sources to outperform baselines and Proxy Tuning on eight low-resource languages across four model families.

  9. Computational Lesions in Multilingual Language Models Separate Shared and Language-specific Brain Alignment

    cs.CL 2026-04 unverdicted novelty 7.0

    Lesioning a shared core in multilingual LLMs drops whole-brain fMRI encoding correlation by 60.32%, while language-specific lesions selectively weaken predictions only for the matched native language.

  10. Towards Measuring the Representation of Subjective Global Opinions in Language Models

    cs.CL 2023-06 conditional novelty 7.0

    LLMs default to responses more similar to opinions from the USA and some European and South American countries; prompting for a country shifts alignment but can introduce stereotypes, while translation does not reliab...

  11. Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung

    cs.CL 2026-07 conditional novelty 6.5

    CKTN is the first Cham–Khmer–Tay-Nung corpus; vocabulary augmentation plus script-calibrated replaced-token pretraining yields the strongest classification encoder and exposes misleading adaptation metrics.

  12. DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

    cs.CL 2026-07 conditional novelty 6.0

    Dialectal robustness and generation are dissociated in LLMs: benchmarks are driven by pretraining and SFT while alignment reshapes generation invisibly to benchmarks, and the method maximizing dialectal reward is leas...

  13. PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

    cs.CL 2026-07 conditional novelty 6.0

    PluraMath extends PolyMath with human-validated math problems in 18 mid-to-extreme low-resource languages and benchmarks 27 reasoning LLMs, finding a persistent high- vs low-resource performance gap.

  14. The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

    cs.CL 2026-06 unverdicted novelty 6.0

    Benign multilingual fine-tuning causes language-specific safety drifts with adversarial compliance rates rising up to four-fold, decoupled from capability gains.

  15. Same question, different history: language, national identity, and credit in large language models

    cs.CL 2026-06 unverdicted novelty 6.0

    Analysis of 11 LLMs on 21 disputed inventions across 12 languages and 75,896 responses finds query language systematically shifts credit toward lower-status claimants in their associated language while Anglophone figu...

  16. When Does Mixing Help? Analyzing Query Embedding Interpolation in Multilingual Dense Retrieval

    cs.CL 2026-06 conditional novelty 6.0

    Optimal interpolation of query embeddings from parallel translations outperforms the best monolingual query in 88/105 cases on mMARCO, showing English-driven asymmetry and negative correlation with typological distance.

  17. MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models

    cs.CL 2026-06 unverdicted novelty 6.0

    MUDIDI introduces a two-stage LLM pipeline for multilingual dictionary digitization, releases a human-annotated dataset from 30 dictionaries, and shows LLMs outperforming OCR and VLMs on character recognition, markup,...

  18. The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

    cs.CL 2026-06 unverdicted novelty 6.0

    Using a 1PL IRT model on real cultural questions across 13 locales, the study identifies a local-language knowledge-access advantage masked by lower proficiency in raw accuracy.

  19. "Chi nas dal soch el sent de legn" -- Auditing Text Corpora for Lombard

    cs.CL 2026-06 unverdicted novelty 6.0

    Manual audit shows web-scraped Lombard corpora are largely noisy and biased toward Western varieties over Eastern ones.

  20. CRAFT: Cost-aware Refinement And Front-aware Tuning of Prompts

    cs.CL 2026-06 unverdicted novelty 6.0

    CRAFT is a Pareto-front prompt optimizer that allocates scarce LLM validation calls to candidates near the current front using accuracy- and cost-oriented generators plus NSGA-II retention.

  21. The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

    cs.CL 2026-06 unverdicted novelty 6.0

    Geometric analysis shows LLM judges use collapsed score ranges and nearly orthogonal axes to humans on subjective rubrics across Indic languages, with inter-LLM agreement not equaling human alignment except on factual tasks.

  22. Learning When to Translate for Multilingual Reasoning

    cs.CL 2026-06 unverdicted novelty 6.0

    Luar is a reinforcement learning method enabling reasoning language models to decide when to invoke English translation for improved multilingual reasoning.

  23. Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization

    cs.CL 2026-05 unverdicted novelty 6.0

    Introduces the MEA benchmark for multi-target cross-lingual summarization across 24 languages and demonstrates that activation steering from English summarization representations improves performance.

  24. Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs

    cs.CL 2026-05 conditional novelty 6.0

    LLMs default to U.S. frameworks for English prompts and China frameworks for Chinese prompts on jurisdiction-underspecified legal-administrative queries, with the pattern holding across all seven tested models.

  25. DEPART: DEcomposing PARiTy across Multilingual LLMs

    cs.CL 2026-05 unverdicted novelty 6.0

    A Bayesian framework decomposes mLLM variance, showing language features explain 79-92% of language identity variance and that model identity vs. benchmark-model interactions dominate differently for understanding ver...

  26. Beyond Catalogue Counts: the Dataset Visibility Asymmetry in Low-Resource Multilingual NLP

    cs.CL 2026-05 unverdicted novelty 6.0

    Catalogue records show 141 languages as data-poor, but citation mining reveals 609 datasets across 53 languages, exposing a visibility gap in multilingual NLP resources.

  27. Scaling Laws for Mixture Pretraining Under Data Constraints

    cs.LG 2026-05 unverdicted novelty 6.0

    Empirical study shows mixture pretraining tolerates higher target data repetition than single-source training, with a new repetition-aware scaling law enabling principled mixture selection based on data size, compute,...

  28. COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling

    cs.LG 2026-04 unverdicted novelty 6.0

    COMPASS uses semantic clustering on multilingual embeddings to select auxiliary data for PEFT adapters, outperforming linguistic-similarity baselines on multilingual benchmarks while supporting continual adaptation.

  29. SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models

    cs.CL 2025-07 conditional novelty 6.0

    SLoW selects low-frequency word dictionaries to boost LLM translation quality and efficiency across 100 languages from FLORES.

  30. How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

    cs.CL 2024-11 unverdicted novelty 6.0

    The study filters non-English Wikipedia, reveals quality problems, proposes a 4-level ranking, and shows filtered data matches or beats raw data in language modeling with largest gains for lower-quality editions.

  31. Lessons from the Trenches on Reproducible Evaluation of Language Models

    cs.CL 2024-05 accept novelty 6.0

    The paper compiles practical lessons on reproducible LM evaluation and introduces the lm-eval library to mitigate common methodological problems in NLP.

  32. Ethical and social risks of harm from Language Models

    cs.CL 2021-12 accept novelty 6.0

    The authors provide a detailed taxonomy of 21 risks associated with language models, covering discrimination, information leaks, misinformation, malicious applications, interaction harms, and societal impacts like job...

  33. Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

    cs.CL 2026-07 unverdicted novelty 5.0

    Meta-analysis of 33 ACL papers shows inconsistent LLM-as-a-Judge results, overtrust, and single-model reliance in multilingual/low-resource settings, with recommendations for better practice.

  34. LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance

    cs.CL 2026-05 unverdicted novelty 5.0

    LANG combines language-adaptive hint guidance, progressive decay, and difficulty-tailored learning horizons in RL to boost non-English reasoning performance while preserving language consistency.

  35. High-Volume Plaintiff-Side Counsel and Single-Appearance Eviction Cases in Philadelphia

    stat.AP 2026-05 unverdicted novelty 5.0

    High-volume plaintiff-side counsel in Philadelphia eviction cases scales up filing volume and procedural steps but does not produce a broad premium on adverse tenant outcomes such as default or judgment.

  36. Which Are the Low-Resource Languages of the Semantic Web?

    cs.AI 2026-05 unverdicted novelty 5.0

    A multi-level categorization from language distributions in DBpedia, BabelNet, and Wikidata defines low-resource languages for Semantic Web knowledge graphs.

  37. Lost in the Tower of Babel: The Adverse Effects of Incidental Multilingualism in LLMs

    cs.CL 2026-05 unverdicted novelty 5.0

    Incidental multilingualism from uneven web training makes LLMs unequal, brittle, and opaque across languages.

  38. Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling

    cs.CL 2026-04 unverdicted novelty 5.0

    Marco-MoE delivers open multilingual MoE models with 5% activation sparsity that outperform similarly sized dense models on English and multilingual benchmarks through efficient upcycling.

  39. Adam's Law: Textual Frequency Law on Large Language Models

    cs.CL 2026-04 conditional novelty 5.0

    For meaning-matched paraphrases, higher estimated sentence frequency improves LLM prompting and fine-tuning across math, translation, commonsense, and tool-calling tasks.

  40. How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language

    cs.CL 2025-06 conditional novelty 5.0

    Bengali sentiment analysis models exhibit persistent identity-based biases across datasets and developer backgrounds despite similar semantic content.

  41. Bridging the Usability Gap: Lessons from Interpreting Studies for Machine Interpreting Design

    cs.CL 2026-06 unverdicted novelty 4.0

    Machine interpreting should shift from fidelity metrics to three design priorities—agency, grounding, and experience—drawn from interpreting studies to close the usability gap with human-mediated communication.

  42. Model-Based Quality Assessment for Massively Multilingual Parallel Data

    cs.CL 2026-05 unverdicted novelty 4.0

    Large-scale benchmarks of multilingual embeddings and QE models show no universal performer; direction-aware routing and calibration recommended for parallel data assessment.

  43. In Data or Invisible: Toward a Better Digital Representation of Low-Resource Languages with Knowledge Graphs

    cs.AI 2026-05 unverdicted novelty 4.0

    A research plan to analyze language distribution in LOD knowledge graphs and explore cross-lingual transfer plus analogical reasoning to improve coverage for low-resource languages.

  44. CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation

    cs.CL 2026-06 unverdicted novelty 3.0

    Compact 0.8B-7B models for bidirectional Japanese-English translation outperform large multilingual models on real-world domain benchmarks.

  45. Adam's Law: Textual Frequency Law on Large Language Models

    cs.CL 2026-04 unverdicted novelty 3.0

    Frequent sentence-level text improves LLM prompting and fine-tuning performance across math, translation, commonsense, and tool-use tasks via a proposed frequency law and curriculum ordering.

  46. Multilingual Vision-Language Models, A Survey

    cs.CL 2025-09 accept novelty 3.0

    The survey identifies a key tension in multilingual vision-language models between language neutrality via contrastive learning and cultural awareness via diverse data, with most benchmarks relying on translation-base...