Pith. sign in

REVIEW 6 cited by

AceGPT, Localizing Large Language Models in Arabic

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.12053 v5 pith:S2SGQZ2B submitted 2023-09-21 cs.CL

classification cs.CL
keywords arabicacegptlanguagemodelmodelscomprehensiveculturallarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper is devoted to the development of a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models. Significant concerns emerge when addressing cultural sensitivity and local values. To address this, the paper proposes a comprehensive solution that includes further pre-training with Arabic texts, Supervised Fine-Tuning (SFT) utilizing native Arabic instructions, and GPT-4 responses in Arabic, alongside Reinforcement Learning with AI Feedback (RLAIF) employing a reward model attuned to local culture and values. The goal is to cultivate culturally cognizant and value-aligned Arabic LLMs capable of accommodating the diverse, application-specific needs of Arabic-speaking communities. Comprehensive evaluations reveal that the resulting model, dubbed `AceGPT', sets the state-of-the-art standard for open Arabic LLMs across various benchmarks. Codes, data, and models are in https://github.com/FreedomIntelligence/AceGPT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multiple LLM Agents Debate for Equitable Cultural Alignment

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multi-agent debate framework where two LLMs discuss cultural scenarios and a judge resolves disagreements improves both accuracy and cultural-group parity on NormAd-ETI, letting 7-9B models match a 27B model.

  2. The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Arabic text written by LLMs carries detectable stylometric signatures, and fine-tuned XLM-RoBERTa detectors reach near-perfect F1 on academic abstracts but degrade on social media.

  3. Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The new Fann or Flop benchmark measures LLM comprehension of Arabic poetry through expert-written verse explanations and shows current LLMs perform poorly on interpretive depth.

  4. RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

    cs.CL 2026-04 accept novelty 5.0 of 10

    A fully open 518M Arabic-specialized LLM, built by vocabulary injection and standard post-training on Qwen2.5-0.5B, beats same-class multilingual baselines and ships at 398 MB quantized.

  5. Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model

    cs.CL 2025-05 reject novelty 5.0 of 10

    A compact 1.5B Arabic-English model beats GPT-4o mini only on the authors' own Tarjama-25 benchmark, while trailing large models on standard WMT24++ and IWSLT2017 tests.

  6. Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks

    cs.CL 2025-07 reject novelty 3.0 of 10

    Across Arabic, English, and Kannada benchmarks, 4-bit and 8-bit quantization preserves most accuracy while aggressive pruning degrades larger multilingual models more than smaller ones.

Pith tools