Pith. sign in

REVIEW 22 cited by

ALLaM: Large Language Models for Arabic and English

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15390 v1 pith:RE4PAEWT submitted 2024-07-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords arabiclanguagemodelsalignmentallamenglishlargemodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present ALLaM: Arabic Large Language Model, a series of large language models to support the ecosystem of Arabic Language Technologies (ALT). ALLaM is carefully trained considering the values of language alignment and knowledge transfer at scale. Our autoregressive decoder-only architecture models demonstrate how second-language acquisition via vocabulary expansion and pretraining on a mixture of Arabic and English text can steer a model towards a new language (Arabic) without any catastrophic forgetting in the original language (English). Furthermore, we highlight the effectiveness of using parallel/translated data to aid the process of knowledge alignment between languages. Finally, we show that extensive alignment with human preferences can significantly enhance the performance of a language model compared to models of a larger scale with lower quality alignment. ALLaM achieves state-of-the-art performance in various Arabic benchmarks, including MMLU Arabic, ACVA, and Arabic Exams. Our aligned models improve both in Arabic and English from their base aligned models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Instruction-hierarchy compliance in LLMs is asymmetric by language and position, and cross-language conflicts yield systematically higher compliance than same-language ones (Language Boundary Effect).

  2. CrossHallu: Do Hallucination Signals Generalize Across Languages and Domains in Large Language Model's Internals?

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Hallucination signals from LLM internals transfer across English–Arabic and Arabic domains for most models, depending on class separability and feature-space language alignment.

  3. Large Concept Models: Language Modeling in a Sentence Representation Space

    cs.CL 2024-12 conditional novelty 7.0 of 10

    A sentence-level language model trained to autoregressively predict SONAR sentence embeddings can summarize, expand, and generate text in unseen languages.

  4. IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

    cs.CL 2026-08 conditional novelty 6.0 of 10

    IslamicTurathBench is a new expert-reviewed Arabic benchmark that tests LLMs on classical Islamic scholarship across seven disciplines, three difficulty tiers, and three task formats.

  5. Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Arabizi spelling varies systematically across five Arabic dialects, and speakers can often recognize their own dialect's Arabizi, but the recognition result is partly confounded by authors judging their own transcriptions.

  6. PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture

    cs.CL 2025-09 accept novelty 6.0 of 10

    PalmX 2025 introduces a two-subtask MCQA benchmark for Arabic and Islamic cultural knowledge and shows that task-specific fine-tuning, especially LoRA, improves LLM accuracy.

  7. Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Code-switching hurts LLM comprehension when non-English tokens enter English text, but inserting English into other languages often improves accuracy; fine-tuning mitigates losses more reliably than prompting.

  8. The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Arabic text written by LLMs carries detectable stylometric signatures, and fine-tuned XLM-RoBERTa detectors reach near-perfect F1 on academic abstracts but degrade on social media.

  9. Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The new Fann or Flop benchmark measures LLM comprehension of Arabic poetry through expert-written verse explanations and shows current LLMs perform poorly on interpretive depth.

  10. How well can LLMs Grade Essays in Arabic?

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Generative LLMs, including Arabic-specific ones, underperform a fine-tuned BERT model on Arabic essay scoring, with bilingual prompting giving the best LLM results.

  11. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.

  12. RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

    cs.CL 2026-04 accept novelty 5.0 of 10

    A fully open 518M Arabic-specialized LLM, built by vocabulary injection and standard post-training on Qwen2.5-0.5B, beats same-class multilingual baselines and ships at 398 MB quantized.

  13. AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    AraHalluEval introduces a 12-indicator Arabic hallucination taxonomy and finds factual errors dominate, with Allam competitive against reasoning models.

  14. Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Nile-Chat models for dual-script Egyptian Arabic beat strong baselines on newly translated benchmarks, but the evaluation may be inflated by training/eval data overlap and Claude-generated script data.

  15. Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model

    cs.CL 2025-05 reject novelty 5.0 of 10

    A compact 1.5B Arabic-English model beats GPT-4o mini only on the authors' own Tarjama-25 benchmark, while trailing large models on standard WMT24++ and IWSLT2017 tests.

  16. Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A 1.6B Arabic language model trained with synthetic multiple-choice instruction data outperforms 7B-13B models on several Arabic multiple-choice benchmarks.

  17. UI-Level Evaluation of ALLaM 34B: Measuring an Arabic-Centric LLM via HUMAIN Chat

    cs.CL 2025-08 reject novelty 4.0 of 10

    A 115-response UI test of ALLaM 34B claims strong Arabic abilities, but the dialect score in the main table conflicts with the paper's own heat map and the underlying data is missing.

  18. From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    On a new 490-question Arabic depth dataset, Claude 3.5 Sonnet answered about 30 percent correctly, while GPT-4 answered about 9 percent, showing current models are weak on culturally specialized Arabic knowledge.

  19. Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Arabic instructions on the R2R navigation task preserve much of GPT-4o mini's success rate but push Phi-3 and Jais to zero, indicating model capability rather than language is the decisive factor.

  20. Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks

    cs.CL 2025-07 reject novelty 3.0 of 10

    Across Arabic, English, and Kannada benchmarks, 4-bit and 8-bit quantization preserves most accuracy while aggressive pruning degrades larger multilingual models more than smaller ones.

  21. Sovereign Large Language Models: Advantages, Strategy and Regulations

    cs.CY 2025-02 unverdicted novelty 2.0 of 10

    A policy survey identifies common strategies, funding models, and regulations for national large language models across more than 18 world regions.

  22. Large Language Models and Arabic Content: A Review

    cs.CL 2025-05 reject

    A survey of Arabic LLMs and Arabic NLP tasks that reports no new experiments, results, or datasets.

Pith tools