REVIEW 9 cited by
Fanar: An Arabic-Centric Multimodal Generative AI Platform
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Fanar, a platform for Arabic-centric multimodal generative AI systems, that supports language, speech and image generation tasks. At the heart of Fanar are Fanar Star and Fanar Prime, two highly capable Arabic Large Language Models (LLMs) that are best in the class on well established benchmarks for similar sized models. Fanar Star is a 7B (billion) parameter model that was trained from scratch on nearly 1 trillion clean and deduplicated Arabic, English and Code tokens. Fanar Prime is a 9B parameter model continually trained on the Gemma-2 9B base model on the same 1 trillion token set. Both models are concurrently deployed and designed to address different types of prompts transparently routed through a custom-built orchestrator. The Fanar platform provides many other capabilities including a customized Islamic Retrieval Augmented Generation (RAG) system for handling religious prompts, a Recency RAG for summarizing information about current or recent events that have occurred after the pre-training data cut-off date. The platform provides additional cognitive capabilities including in-house bilingual speech recognition that supports multiple Arabic dialects, voice and image generation that is fine-tuned to better reflect regional characteristics. Finally, Fanar provides an attribution service that can be used to verify the authenticity of fact based generated content. The design, development, and implementation of Fanar was entirely undertaken at Hamad Bin Khalifa University's Qatar Computing Research Institute (QCRI) and was sponsored by Qatar's Ministry of Communications and Information Technology to enable sovereign AI technology development.
Forward citations
Cited by 9 Pith papers
-
Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection
Arabic medical failure in LLMs is a knowledge-routing breakdown, not a knowledge deficit, and TLoRA, a low-rank adapter on the mechanistically identified layer window, improves Arabic medical MCQA over full-network LoRA.
-
HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering
HalluTruthQA provides 2,400 expert-annotated Arabic QA examples with character-level hallucination spans, explanations, and verification candidates, and shows no single LLM excels at all four evaluation tasks.
-
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs
Arabic dialects are causally steerable in LLMs via sparse LAPE neurons and distributed activation vectors, with vector steering giving more reliable dialect control than neuron rescaling.
-
The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—...
-
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
The new Fann or Flop benchmark measures LLM comprehension of Arabic poetry through expert-written verse explanations and shows current LLMs perform poorly on interpretive depth.
-
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
ARB provides 1,356 Arabic multimodal questions with 5,119 human-reviewed reasoning steps and shows leading models score much higher on reasoning fluency than on correct answers.
-
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
AraHalluEval introduces a 12-indicator Arabic hallucination taxonomy and finds factual errors dominate, with Allam competitive against reasoning models.
-
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
SpokenNativQA is a human-recorded Arabic and English spoken question-answering benchmark built from MultiNativQA text pairs, with ASR and LLM baselines.
-
From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation
On a new 490-question Arabic depth dataset, Claude 3.5 Sonnet answered about 30 percent correctly, while GPT-4 answered about 9 percent, showing current models are weak on culturally specialized Arabic knowledge.
Discussion (0). Sign in to comment.