Pith. sign in

REVIEW 11 cited by

Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.16149 v2 pith:IL64DHSM submitted 2023-08-30 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords modelsopenarabicenglishfoundationinstruction-tunedjaisjais-chat
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Jais and Jais-chat, new state-of-the-art Arabic-centric foundation and instruction-tuned open generative large language models (LLMs). The models are based on the GPT-3 decoder-only architecture and are pretrained on a mixture of Arabic and English texts, including source code in various programming languages. With 13 billion parameters, they demonstrate better knowledge and reasoning capabilities in Arabic than any existing open Arabic and multilingual models by a sizable margin, based on extensive evaluation. Moreover, the models are competitive in English compared to English-centric open models of similar size, despite being trained on much less English data. We provide a detailed description of the training, the tuning, the safety alignment, and the evaluation of the models. We release two open versions of the model -- the foundation Jais model, and an instruction-tuned Jais-chat variant -- with the aim of promoting research on Arabic LLMs. Available at https://huggingface.co/inception-mbzuai/jais-13b-chat

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

    cs.CL 2026-08 conditional novelty 6.0 of 10

    ALiBi's linearly growing positional bias underflows floating-point attention in long contexts, zeroing out distant attention weights, with measurable but task-dependent effects on retrieval.

  2. Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Arabizi spelling varies systematically across five Arabic dialects, and speakers can often recognize their own dialect's Arabizi, but the recognition result is partly confounded by authors judging their own transcriptions.

  3. ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

    cs.CL 2026-08 conditional novelty 6.0 of 10

    ArabicDialectSafety, a 25,071-prompt, six-dialect Arabic safety benchmark, shows fine-tuned MARBERTv2 reaches 0.95 binary and 0.90 granular Macro-F1, outperforming prompted frontier LLMs.

  4. Toward Culturally Aligned LLMs through Ontology-Guided Multi-Agent Reasoning

    cs.CL 2026-01 conditional novelty 6.0 of 10

    OG-MAR improves LLM prediction of survey responses by retrieving demographically matched World Values Survey profiles and ontology-derived value relations, then aggregating persona-agent answers with a judge agent.

  5. Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Llama-GENBA-10B is a 10B-parameter trilingual model that reports top Bavarian scores among sub-10B models on a machine-translated benchmark the authors built.

  6. BALSAM: A Platform for Benchmarking Arabic Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    BALSAM is a new Arabic LLM benchmark with blind test sets, and the paper argues that LLM-based judging should replace n-gram and embedding metrics for scoring it.

  7. AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data

    cs.CL 2025-07 conditional novelty 6.0 of 10

    AraTable is the first Arabic tabular QA benchmark; its experiments show LLMs are much weaker at reasoning over Arabic tables than at direct lookup.

  8. SpeLLM: Character-Level Multi-Head Decoding

    cs.CL 2025-07 conditional novelty 6.0 of 10

    SpeLLM converts a standard token-based LLM into a character-spelling model with multiple parallel output heads, achieving competitive downstream performance with a 5.1% average decoding speedup.

  9. RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

    cs.CL 2026-04 accept novelty 5.0 of 10

    A fully open 518M Arabic-specialized LLM, built by vocabulary injection and standard post-training on Qwen2.5-0.5B, beats same-class multilingual baselines and ships at 398 MB quantized.

  10. AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    AraHalluEval introduces a 12-indicator Arabic hallucination taxonomy and finds factual errors dominate, with Allam competitive against reasoning models.

  11. Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks

    cs.CL 2025-07 reject novelty 3.0 of 10

    Across Arabic, English, and Kannada benchmarks, 4-bit and 8-bit quantization preserves most accuracy while aggressive pruning degrades larger multilingual models more than smaller ones.

Pith tools