Pith. sign in

REVIEW 10 cited by

ALLaM: Large Language Models for Arabic and English

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15390 v1 pith:RE4PAEWT submitted 2024-07-22 cs.CL cs.AI

ALLaM: Large Language Models for Arabic and English

classification cs.CL cs.AI
keywords arabiclanguagemodelsalignmentallamenglishlargemodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present ALLaM: Arabic Large Language Model, a series of large language models to support the ecosystem of Arabic Language Technologies (ALT). ALLaM is carefully trained considering the values of language alignment and knowledge transfer at scale. Our autoregressive decoder-only architecture models demonstrate how second-language acquisition via vocabulary expansion and pretraining on a mixture of Arabic and English text can steer a model towards a new language (Arabic) without any catastrophic forgetting in the original language (English). Furthermore, we highlight the effectiveness of using parallel/translated data to aid the process of knowledge alignment between languages. Finally, we show that extensive alignment with human preferences can significantly enhance the performance of a language model compared to models of a larger scale with lower quality alignment. ALLaM achieves state-of-the-art performance in various Arabic benchmarks, including MMLU Arabic, ACVA, and Arabic Exams. Our aligned models improve both in Arabic and English from their base aligned models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

    cs.CL 2026-07 conditional novelty 7.0

    Instruction-hierarchy compliance in LLMs is asymmetric by language and position, and cross-language conflicts yield systematically higher compliance than same-language ones (Language Boundary Effect).

  2. CrossHallu: Do Hallucination Signals Generalize Across Languages and Domains in Large Language Model's Internals?

    cs.CL 2026-07 conditional novelty 7.0

    Hallucination signals from LLM internals transfer across English–Arabic and Arabic domains for most models, depending on class separability and feature-space language alignment.

  3. HalluScore: Large Language Model Hallucination Question Answering Benchmark

    cs.CL 2026-05 unverdicted novelty 7.0

    HalluScore is a curated Arabic QA dataset with 827 questions, ground-truth evidence, and human annotations used to measure hallucination rates across 17 LLMs.

  4. LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation

    cs.CL 2026-04 unverdicted novelty 7.0

    LQM introduces a six-level linguistically motivated error taxonomy for MT evaluation and applies it via expert annotation to LLM outputs on a new 3,850-sentence multi-dialect Arabic corpus.

  5. Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

    cs.CL 2026-08 conditional novelty 6.0

    Arabizi spelling varies systematically across five Arabic dialects, and speakers can often recognize their own dialect's Arabizi, but the recognition result is partly confounded by authors judging their own transcriptions.

  6. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 conditional novelty 5.0

    A multi-role red-teaming framework with attacker, target, and jury LLMs measures faithfulness in English and Arabic, finding false-premise prompts and length limits change unfaithfulness rates.

  7. LLM-Based Financial Sentiment Analysis in Arabic: Evidence from Saudi Markets

    cs.CL 2026-05 unverdicted novelty 5.0

    Introduces a multi-stage Arabic financial sentiment pipeline that produces an 84K-sample corpus for company-level analysis tied to Saudi stock market behavior.

  8. RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

    cs.CL 2026-04 accept novelty 5.0

    A fully open 518M Arabic-specialized LLM, built by vocabulary injection and standard post-training on Qwen2.5-0.5B, beats same-class multilingual baselines and ships at 398 MB quantized.

  9. Noise Steering for Controlled Text Generation: Improving Diversity and Reading-Level Fidelity in Arabic Educational Story Generation

    cs.CL 2026-04 unverdicted novelty 5.0

    Residual-stream noise injection raises narrative diversity in Arabic educational stories while preserving reading-grade level, outperforming high-temperature sampling across five 7-9B models.

  10. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 unverdicted novelty 4.0

    Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.