{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:DEJPA3Z57HRBELNYKPSFPHSMQD","short_pith_number":"pith:DEJPA3Z5","schema_version":"1.0","canonical_sha256":"1912f06f3df9e2122db853e4579e4c80e77f0c8c251724fb6bec603e3d111512","source":{"kind":"arxiv","id":"2207.04672","version":3},"attestation_state":"computed","paper":{"title":"No Language Left Behind: Scaling Human-Centered Machine Translation","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"A sparsely gated mixture of experts model trained on mined low-resource data achieves 44% relative BLEU improvement in translating 200 languages.","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Alexandre Mourachko, Al Youngblood, Angela Fan, Anna Sun, Bapi Akula, Chau Tran, Christophe Ropers, Cynthia Gao, Daniel Licht, Dirk Rowe, Elahe Kalbassi, Francisco Guzm\\'an, Gabriel Mejia Gonzalez, Guillaume Wenzek, Holger Schwenk, James Cross, Janice Lam, Jean Maillard, Jeff Wang (NLLB Team), John Hoffman, Kaushik Ram Sadagopan, Kenneth Heafield, Kevin Heffernan, Loic Barrault, Maha Elbayad, Marta R. Costa-juss\\`a, Necip Fazil Ayan, NLLB Team, Onur \\c{C}elebi, Philipp Koehn, Pierre Andrews, Prangthip Hansanti, Safiyyah Saleem, Semarley Jarrett, Sergey Edunov, Shannon Spruit, Shruti Bhosale, Skyler Wang, Vedanuj Goswami","submitted_at":"2022-07-11T07:33:36Z","abstract_excerpt":"Driven by the goal of eradicating language barriers on a global scale, machine translation has solidified itself as a key focus of artificial intelligence research today. However, such efforts have coalesced around a small subset of languages, leaving behind the vast majority of mostly low-resource languages. What does it take to break the 200 language barrier while ensuring safe, high quality results, all while keeping ethical considerations in mind? In No Language Left Behind, we took on this challenge by first contextualizing the need for low-resource language translation support through ex"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2207.04672","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by-sa/4.0/","primary_cat":"cs.CL","submitted_at":"2022-07-11T07:33:36Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"01ed8c3f5073036312b0a2e1fc34f8dfd220e1d0182d756cd9061ffa8c81e502","abstract_canon_sha256":"f38df21833913da5e913b5c418119c400fbbc9f7ebbf02bbc3229f86d01400be"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T04:51:35.136817Z","signature_b64":"0/gRapn12wYSkuGmBTJnx9UtG1Vx+9+1J9HMEk/qhITpHJiUqCMa87ahy/9SKMDrZDtaiuclxSTgYJCrqt3gBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"1912f06f3df9e2122db853e4579e4c80e77f0c8c251724fb6bec603e3d111512","last_reissued_at":"2026-07-05T04:51:35.136376Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T04:51:35.136376Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"No Language Left Behind: Scaling Human-Centered Machine Translation","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"A sparsely gated mixture of experts model trained on mined low-resource data achieves 44% relative BLEU improvement in translating 200 languages.","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Alexandre Mourachko, Al Youngblood, Angela Fan, Anna Sun, Bapi Akula, Chau Tran, Christophe Ropers, Cynthia Gao, Daniel Licht, Dirk Rowe, Elahe Kalbassi, Francisco Guzm\\'an, Gabriel Mejia Gonzalez, Guillaume Wenzek, Holger Schwenk, James Cross, Janice Lam, Jean Maillard, Jeff Wang (NLLB Team), John Hoffman, Kaushik Ram Sadagopan, Kenneth Heafield, Kevin Heffernan, Loic Barrault, Maha Elbayad, Marta R. Costa-juss\\`a, Necip Fazil Ayan, NLLB Team, Onur \\c{C}elebi, Philipp Koehn, Pierre Andrews, Prangthip Hansanti, Safiyyah Saleem, Semarley Jarrett, Sergey Edunov, Shannon Spruit, Shruti Bhosale, Skyler Wang, Vedanuj Goswami","submitted_at":"2022-07-11T07:33:36Z","abstract_excerpt":"Driven by the goal of eradicating language barriers on a global scale, machine translation has solidified itself as a key focus of artificial intelligence research today. However, such efforts have coalesced around a small subset of languages, leaving behind the vast majority of mostly low-resource languages. What does it take to break the 200 language barrier while ensuring safe, high quality results, all while keeping ethical considerations in mind? In No Language Left Behind, we took on this challenge by first contextualizing the need for low-resource language translation support through ex"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Our model achieves an improvement of 44% BLEU relative to the previous state-of-the-art, laying important groundwork towards realizing a universal translation system.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The novel data mining techniques and architectural/training improvements produce genuinely higher-quality and safer translations for low-resource languages rather than merely fitting the new benchmark or human raters.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"A sparsely gated mixture-of-experts model trained on newly mined low-resource data achieves 44% relative BLEU improvement across 200 languages while adding human safety evaluation.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"A sparsely gated mixture of experts model trained on mined low-resource data achieves 44% relative BLEU improvement in translating 200 languages.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"44ba93e32fb0cbbebd1ae1d0318e1be53e526ca60ee6687cd65907d47672a8e9"},"source":{"id":"2207.04672","kind":"arxiv","version":3},"verdict":{"id":"9bda6551-a1a6-41fa-92bf-b973a7f99f4e","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-12T17:48:09.914595Z","strongest_claim":"Our model achieves an improvement of 44% BLEU relative to the previous state-of-the-art, laying important groundwork towards realizing a universal translation system.","one_line_summary":"A sparsely gated mixture-of-experts model trained on newly mined low-resource data achieves 44% relative BLEU improvement across 200 languages while adding human safety evaluation.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The novel data mining techniques and architectural/training improvements produce genuinely higher-quality and safer translations for low-resource languages rather than merely fitting the new benchmark or human raters.","pith_extraction_headline":"A sparsely gated mixture of experts model trained on mined low-resource data achieves 44% relative BLEU improvement in translating 200 languages."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2207.04672/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":16,"sample":[{"doi":"","year":null,"title":"URL https://arxiv.org/abs/2110.03036. Benjamin Akera, Jonathan Mukiibi, Lydia Sanyu Naggayi, Claire Babirye, Isaac Owomugisha, Solomon Nsumba, Joyce Nakatumba-Nabende, Engineer Bainomugisha, Ernest Mw","work_id":"414b85bf-6e31-46c3-acb3-2b154373a894","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2021,"title":"Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ondřej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R","work_id":"e45c1c77-a3b8-4489-996b-3a894895dde2","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.18653/v1/","year":2016,"title":"In: Zong, C., Xia, F., Li, W., Navigli, R","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.18653/v1/2021.iwslt-1.1","year":2021,"title":"doi: 10.18653/v1/2021.iwslt-1.1","work_id":"9da1e387-a255-4516-93af-7bb3533e17d3","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.18653/v1/2020.acl-main.485","year":2022,"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","work_id":"1fe8c7c8-aff7-4b94-9096-e549d7e60789","ref_index":5,"cited_arxiv_id":"1308.3432","is_internal_anchor":true}],"resolved_work":16,"snapshot_sha256":"6b48e14dea6f5c9e7f65d1d65fb16e21b55a0768b0977e6deb0327b11c4c8bd2","internal_anchors":2},"formal_canon":{"evidence_count":2,"snapshot_sha256":"072a74064a96f3c39e78b2e023745ea6bda1bd9326bb5ecac5425ff7571026d4"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2207.04672","created_at":"2026-07-05T04:51:35.136440+00:00"},{"alias_kind":"arxiv_version","alias_value":"2207.04672v3","created_at":"2026-07-05T04:51:35.136440+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2207.04672","created_at":"2026-07-05T04:51:35.136440+00:00"},{"alias_kind":"pith_short_12","alias_value":"DEJPA3Z57HRB","created_at":"2026-07-05T04:51:35.136440+00:00"},{"alias_kind":"pith_short_16","alias_value":"DEJPA3Z57HRBELNY","created_at":"2026-07-05T04:51:35.136440+00:00"},{"alias_kind":"pith_short_8","alias_value":"DEJPA3Z5","created_at":"2026-07-05T04:51:35.136440+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":100,"internal_anchor_count":100,"sample":[{"citing_arxiv_id":"2607.06611","citing_title":"Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts","ref_index":35,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06457","citing_title":"Andha-Dhun: A First Look at Audio Descriptions in Hindi","ref_index":37,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26015","citing_title":"The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar","ref_index":13,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25821","citing_title":"SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment","ref_index":7,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25246","citing_title":"Multilingual Hematology Visual Question Answering Dataset","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24200","citing_title":"MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval","ref_index":7,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26466","citing_title":"Soft Token Alignment for Cross-Lingual Reasoning","ref_index":58,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22494","citing_title":"Deep Learning-Based Sign Language Recognition from Videos and Cross-Lingual Translation to Indian Vernaculars","ref_index":8,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22269","citing_title":"Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2606.18453","citing_title":"LLM Parameters for Math Across Languages: Shared or Separate?","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11875","citing_title":"I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System","ref_index":70,"is_internal_anchor":true},{"citing_arxiv_id":"2606.08748","citing_title":"HydraQE: OSU's Submission for the IWSLT 2026 Speech Translation Metrics Shared Task","ref_index":21,"is_internal_anchor":true},{"citing_arxiv_id":"2606.28551","citing_title":"DataComp-VLM: Improved Open Datasets for Vision-Language Models","ref_index":50,"is_internal_anchor":true},{"citing_arxiv_id":"2607.00171","citing_title":"ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs","ref_index":28,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07020","citing_title":"MADE: Beyond Scoring via a Multilingual Agentic Diagnosing Engine for Fine-Grained Evaluation Insights","ref_index":63,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05444","citing_title":"Multilingual Coreference Resolution via Cycle-Consistent Machine Translation","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03793","citing_title":"Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03259","citing_title":"Beyond \"To whom it may concern\": Tailoring Machine Translation to Audience and Intent","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03241","citing_title":"Benchmarking Speech-to-Speech Translation Models","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03345","citing_title":"Beyond Semantics: Modeling Factual and Affective Perceptual Experiences from Vision-Language Data","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2606.02806","citing_title":"Translating Classical Poetry into Modern Prose","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2606.02147","citing_title":"Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resource Languages","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2606.00285","citing_title":"Model-Based Quality Assessment for Massively Multilingual Parallel Data","ref_index":65,"is_internal_anchor":true},{"citing_arxiv_id":"2606.28551","citing_title":"DataComp-VLM: Improved Open Datasets for Vision-Language Models","ref_index":50,"is_internal_anchor":true},{"citing_arxiv_id":"2606.28325","citing_title":"The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing","ref_index":4,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD","json":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD.json","graph_json":"https://pith.science/api/pith-number/DEJPA3Z57HRBELNYKPSFPHSMQD/graph.json","events_json":"https://pith.science/api/pith-number/DEJPA3Z57HRBELNYKPSFPHSMQD/events.json","paper":"https://pith.science/paper/DEJPA3Z5"},"agent_actions":{"view_html":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD","download_json":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD.json","view_paper":"https://pith.science/paper/DEJPA3Z5","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2207.04672&json=true","fetch_graph":"https://pith.science/api/pith-number/DEJPA3Z57HRBELNYKPSFPHSMQD/graph.json","fetch_events":"https://pith.science/api/pith-number/DEJPA3Z57HRBELNYKPSFPHSMQD/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD/action/timestamp_anchor","attest_storage":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD/action/storage_attestation","attest_author":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD/action/author_attestation","sign_citation":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD/action/citation_signature","submit_replication":"https://pith.science/pith/DEJPA3Z57HRBELNYKPSFPHSMQD/action/replication_record"}},"created_at":"2026-07-05T04:51:35.136440+00:00","updated_at":"2026-07-05T04:51:35.136440+00:00"}