Pith. sign in

REVIEW 3 cited by

Systematic Inequalities in Language Technology Performance across the World's Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.06733 v1 pith:P4247YSP submitted 2021-10-13 cs.CL

classification cs.CL
keywords languagetechnologiesgloballanguagesperformanceresearchtechnologyworld
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Natural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development. While the performance of NLP methods has grown enormously over the last decade, this progress has been restricted to a minuscule subset of the world's 6,500 languages. We introduce a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP. Our analyses involve the field at large, but also more in-depth studies on both user-facing technologies (machine translation, language understanding, question answering, text-to-speech synthesis) as well as more linguistic NLP tasks (dependency parsing, morphological inflection). In the process, we (1) quantify disparities in the current state of NLP research, (2) explore some of its associated societal and academic factors, and (3) produce tailored recommendations for evidence-based policy making aimed at promoting more global and equitable language technologies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

    eess.AS 2025-09 conditional novelty 6.0 of 10

    Pre-training HuBERT on 13,164 hours of multilingual child-centered audio and fine-tuning for voice type classification yields 64.6% average F1, beating English-only and adult-speech baselines by 5.9 and 13.2 points.

  2. Bridging Cultural Distance Between Models Default and Local Classroom Demands: How Global Teachers Adopt GenAI to Support Everyday Teaching Practices

    cs.HC 2025-09 conditional novelty 6.0 of 10

    Teachers' experiences with generative AI cluster into low, mid, and high cultural distance, defined by how much adaptation work teachers must do.

  3. Bridging the Gap with Retrieval-Augmented Generation: Making Prosthetic Device User Manuals Available in Marginalised Languages

    cs.LG 2025-06 reject novelty 2.0 of 10

    A proposed RAG-based framework for translating prosthetic device manuals into marginalised languages is described, but no results are presented.

Pith tools