{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2021:CZSH7UJRSFFYL254MGLUIOGZ4C","short_pith_number":"pith:CZSH7UJR","schema_version":"1.0","canonical_sha256":"16647fd131914b85ebbc61974438d9e08bf89fa4b5898c207487aa1a9ac01d69","source":{"kind":"arxiv","id":"2112.10668","version":3},"attestation_state":"computed","paper":{"title":"Few-shot Learning with Multilingual Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Brian O'Horo, Daniel Simig, Jeff Wang, Jingfei Du, Luke Zettlemoyer, Mikel Artetxe, Mona Diab, Myle Ott, Naman Goyal, Punit Singh Koura, Ramakanth Pasunuru, Sam Shleifer, Shruti Bhosale, Shuohui Chen, Tianlu Wang, Todor Mihaylov, Veselin Stoyanov, Vishrav Chaudhary, Xian Li, Xi Victoria Lin, Zornitsa Kozareva","submitted_at":"2021-12-20T16:52:35Z","abstract_excerpt":"Large-scale generative language models such as GPT-3 are competitive few-shot learners. While these models are known to be able to jointly represent many different languages, their training data is dominated by English, potentially limiting their cross-lingual generalization. In this work, we train multilingual generative language models on a corpus covering a diverse set of languages, and study their few- and zero-shot learning capabilities in a wide range of tasks. Our largest model with 7.5 billion parameters sets new state of the art in few-shot learning in more than 20 representative lang"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2112.10668","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2021-12-20T16:52:35Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"16c88dd1fc3bdec83215c46b3267ecf11fe161a16d83dace750d5e388b18b0c7","abstract_canon_sha256":"902754f03849b81ac4c3a56f53a107f99a645214f2cf1a899269493b0b0a4bfe"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T05:14:56.035420Z","signature_b64":"LA3gmr7kWXXMX4ox+6YJdJwDWY8dv/j6Me1ByTK6OyQa3tFgVx6oEr8I4pr0STuHiRfNBZRSq8d3QERhBWthBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"16647fd131914b85ebbc61974438d9e08bf89fa4b5898c207487aa1a9ac01d69","last_reissued_at":"2026-07-05T05:14:56.035009Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T05:14:56.035009Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Few-shot Learning with Multilingual Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Brian O'Horo, Daniel Simig, Jeff Wang, Jingfei Du, Luke Zettlemoyer, Mikel Artetxe, Mona Diab, Myle Ott, Naman Goyal, Punit Singh Koura, Ramakanth Pasunuru, Sam Shleifer, Shruti Bhosale, Shuohui Chen, Tianlu Wang, Todor Mihaylov, Veselin Stoyanov, Vishrav Chaudhary, Xian Li, Xi Victoria Lin, Zornitsa Kozareva","submitted_at":"2021-12-20T16:52:35Z","abstract_excerpt":"Large-scale generative language models such as GPT-3 are competitive few-shot learners. While these models are known to be able to jointly represent many different languages, their training data is dominated by English, potentially limiting their cross-lingual generalization. In this work, we train multilingual generative language models on a corpus covering a diverse set of languages, and study their few- and zero-shot learning capabilities in a wide range of tasks. Our largest model with 7.5 billion parameters sets new state of the art in few-shot learning in more than 20 representative lang"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2112.10668","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2112.10668/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2112.10668","created_at":"2026-07-05T05:14:56.035064+00:00"},{"alias_kind":"arxiv_version","alias_value":"2112.10668v3","created_at":"2026-07-05T05:14:56.035064+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2112.10668","created_at":"2026-07-05T05:14:56.035064+00:00"},{"alias_kind":"pith_short_12","alias_value":"CZSH7UJRSFFY","created_at":"2026-07-05T05:14:56.035064+00:00"},{"alias_kind":"pith_short_16","alias_value":"CZSH7UJRSFFYL254","created_at":"2026-07-05T05:14:56.035064+00:00"},{"alias_kind":"pith_short_8","alias_value":"CZSH7UJR","created_at":"2026-07-05T05:14:56.035064+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":13,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2607.00890","citing_title":"MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages","ref_index":112,"is_internal_anchor":false},{"citing_arxiv_id":"2412.12686","citing_title":"Exploring Cross-lingual Latent Transplantation: Mutual Opportunities and Open Challenges","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22567","citing_title":"LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2507.09205","citing_title":"From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2305.16264","citing_title":"Scaling Data-Constrained Language Models","ref_index":65,"is_internal_anchor":false},{"citing_arxiv_id":"2405.14782","citing_title":"Lessons from the Trenches on Reproducible Evaluation of Language Models","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13225","citing_title":"Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13538","citing_title":"Locale-Conditioned Few-Shot Prompting Mitigates Demonstration Regurgitation in On-Device PII Substitution with Small Language Models","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2211.05100","citing_title":"BLOOM: A 176B-Parameter Open-Access Multilingual Language Model","ref_index":268,"is_internal_anchor":false},{"citing_arxiv_id":"2211.05100","citing_title":"BLOOM: A 176B-Parameter Open-Access Multilingual Language Model","ref_index":176,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06154","citing_title":"Graphlets as Building Blocks for Structural Vocabulary in Knowledge Graph Foundation Models","ref_index":51,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20549","citing_title":"Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2205.01068","citing_title":"OPT: Open Pre-trained Transformer Language Models","ref_index":291,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C","json":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C.json","graph_json":"https://pith.science/api/pith-number/CZSH7UJRSFFYL254MGLUIOGZ4C/graph.json","events_json":"https://pith.science/api/pith-number/CZSH7UJRSFFYL254MGLUIOGZ4C/events.json","paper":"https://pith.science/paper/CZSH7UJR"},"agent_actions":{"view_html":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C","download_json":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C.json","view_paper":"https://pith.science/paper/CZSH7UJR","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2112.10668&json=true","fetch_graph":"https://pith.science/api/pith-number/CZSH7UJRSFFYL254MGLUIOGZ4C/graph.json","fetch_events":"https://pith.science/api/pith-number/CZSH7UJRSFFYL254MGLUIOGZ4C/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C/action/timestamp_anchor","attest_storage":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C/action/storage_attestation","attest_author":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C/action/author_attestation","sign_citation":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C/action/citation_signature","submit_replication":"https://pith.science/pith/CZSH7UJRSFFYL254MGLUIOGZ4C/action/replication_record"}},"created_at":"2026-07-05T05:14:56.035064+00:00","updated_at":"2026-07-05T05:14:56.035064+00:00"}