{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:RE4PAEWTQPJRHNHPGEMPPEDBM5","short_pith_number":"pith:RE4PAEWT","schema_version":"1.0","canonical_sha256":"8938f012d383d313b4ef3118f79061674ca9468b0fcb17841e440088b20a2375","source":{"kind":"arxiv","id":"2407.15390","version":1},"attestation_state":"computed","paper":{"title":"ALLaM: Large Language Models for Arabic and English","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Abdalghani Abujabal, Abdulmohsen Al-Thubaity, Ahmed Abdelali, Ali Alammari, Areeb Alowisheq, Faisal A. Mirza, Ghadah Alabduljabbar, Haidar Khan, Hassan A. Alahmed, Hisham A. Alyahya, Jeril Kuriakose, Majed Alrubaian, Maryam Al Mansour, M Saiful Bari, Nora Al-Twairesh, Norah A. Alzahrani, Nouf M. Alotaibi, Raghad Alkhathran, Raneem Alnajim, Salman Alsubaihi, Shaykhah Z. Alsubaie, Sultan Alrashed, Yazeed Alnumay, Yousef Almushayqih, Zaki Alawami","submitted_at":"2024-07-22T05:35:17Z","abstract_excerpt":"We present ALLaM: Arabic Large Language Model, a series of large language models to support the ecosystem of Arabic Language Technologies (ALT). ALLaM is carefully trained considering the values of language alignment and knowledge transfer at scale. Our autoregressive decoder-only architecture models demonstrate how second-language acquisition via vocabulary expansion and pretraining on a mixture of Arabic and English text can steer a model towards a new language (Arabic) without any catastrophic forgetting in the original language (English). Furthermore, we highlight the effectiveness of usin"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2407.15390","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-07-22T05:35:17Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"df0694bc83a62424d221d325a49550101064115ea37981e2ccade7821d944396","abstract_canon_sha256":"9a78d44dfbe6bbd23bb62582cb7719b585b719d5fc224aad96de02e56cfb0795"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:46:43.120121Z","signature_b64":"rHVwqdjvJ7QyACf2wyc2p2Gxpfo8pjR6LDxzQqJowroir1P6CsYND5/BDcXKAUHEZKvVlYSy4qvpjoMi7PyaAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"8938f012d383d313b4ef3118f79061674ca9468b0fcb17841e440088b20a2375","last_reissued_at":"2026-07-05T08:46:43.119685Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:46:43.119685Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"ALLaM: Large Language Models for Arabic and English","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Abdalghani Abujabal, Abdulmohsen Al-Thubaity, Ahmed Abdelali, Ali Alammari, Areeb Alowisheq, Faisal A. Mirza, Ghadah Alabduljabbar, Haidar Khan, Hassan A. Alahmed, Hisham A. Alyahya, Jeril Kuriakose, Majed Alrubaian, Maryam Al Mansour, M Saiful Bari, Nora Al-Twairesh, Norah A. Alzahrani, Nouf M. Alotaibi, Raghad Alkhathran, Raneem Alnajim, Salman Alsubaihi, Shaykhah Z. Alsubaie, Sultan Alrashed, Yazeed Alnumay, Yousef Almushayqih, Zaki Alawami","submitted_at":"2024-07-22T05:35:17Z","abstract_excerpt":"We present ALLaM: Arabic Large Language Model, a series of large language models to support the ecosystem of Arabic Language Technologies (ALT). ALLaM is carefully trained considering the values of language alignment and knowledge transfer at scale. Our autoregressive decoder-only architecture models demonstrate how second-language acquisition via vocabulary expansion and pretraining on a mixture of Arabic and English text can steer a model towards a new language (Arabic) without any catastrophic forgetting in the original language (English). Furthermore, we highlight the effectiveness of usin"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2407.15390","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2407.15390/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2407.15390","created_at":"2026-07-05T08:46:43.119740+00:00"},{"alias_kind":"arxiv_version","alias_value":"2407.15390v1","created_at":"2026-07-05T08:46:43.119740+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2407.15390","created_at":"2026-07-05T08:46:43.119740+00:00"},{"alias_kind":"pith_short_12","alias_value":"RE4PAEWTQPJR","created_at":"2026-07-05T08:46:43.119740+00:00"},{"alias_kind":"pith_short_16","alias_value":"RE4PAEWTQPJRHNHP","created_at":"2026-07-05T08:46:43.119740+00:00"},{"alias_kind":"pith_short_8","alias_value":"RE4PAEWT","created_at":"2026-07-05T08:46:43.119740+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25476","citing_title":"A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19714","citing_title":"LLM-Based Financial Sentiment Analysis in Arabic: Evidence from Saudi Markets","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17007","citing_title":"HalluScore: Large Language Model Hallucination Question Answering Benchmark","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2604.03380","citing_title":"Noise Steering for Controlled Text Generation: Improving Diversity and Reading-Level Fidelity in Arabic Educational Story Generation","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18490","citing_title":"LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation","ref_index":92,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5","json":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5.json","graph_json":"https://pith.science/api/pith-number/RE4PAEWTQPJRHNHPGEMPPEDBM5/graph.json","events_json":"https://pith.science/api/pith-number/RE4PAEWTQPJRHNHPGEMPPEDBM5/events.json","paper":"https://pith.science/paper/RE4PAEWT"},"agent_actions":{"view_html":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5","download_json":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5.json","view_paper":"https://pith.science/paper/RE4PAEWT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2407.15390&json=true","fetch_graph":"https://pith.science/api/pith-number/RE4PAEWTQPJRHNHPGEMPPEDBM5/graph.json","fetch_events":"https://pith.science/api/pith-number/RE4PAEWTQPJRHNHPGEMPPEDBM5/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5/action/timestamp_anchor","attest_storage":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5/action/storage_attestation","attest_author":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5/action/author_attestation","sign_citation":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5/action/citation_signature","submit_replication":"https://pith.science/pith/RE4PAEWTQPJRHNHPGEMPPEDBM5/action/replication_record"}},"created_at":"2026-07-05T08:46:43.119740+00:00","updated_at":"2026-07-05T08:46:43.119740+00:00"}