Pith. sign in

REVIEW 3 cited by

AIN: The Arabic INclusive Large Multimodal Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.00094 v2 pith:CAFRVPEC submitted 2025-01-31 cs.CV cs.AIcs.CLcs.HCcs.LG

classification cs.CVcs.AIcs.CLcs.HCcs.LG
keywords arabicmultimodalunderstandinglargevisualacrosscapabilitiesdemonstrates
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Amid the swift progress of large language models (LLMs) and their evolution into large multimodal models (LMMs), significant strides have been made in high-resource languages such as English and Chinese. While Arabic LLMs have seen notable progress, Arabic LMMs remain largely unexplored, often narrowly focusing on a few specific aspects of the language and visual understanding. To bridge this gap, we introduce AIN-the Arabic Inclusive Multimodal Model-designed to excel across diverse domains. AIN is an English-Arabic bilingual LMM designed to excel in English and Arabic, leveraging carefully constructed 3.6 million high-quality Arabic-English multimodal data samples. AIN demonstrates state-of-the-art Arabic performance, while also possessing strong English-language visual capabilities. On the recent CAMEL-Bench benchmark comprising 38 sub-domains including, multi-image understanding, complex visual perception, handwritten document understanding, video understanding, medical imaging, plant diseases, and remote sensing-based land use understanding, our AIN demonstrates strong performance with the 7B model outperforming GPT-4o by an absolute gain of 3.4% averaged over eight domains and 38 sub-domains. AIN's superior capabilities position it as a significant step toward empowering Arabic speakers with advanced multimodal generative AI tools across diverse applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition

    cs.CV 2025-05 conditional novelty 6.0 of 10

    SARD is a new synthetic dataset of 843,622 Arabic book pages with ten fonts, plus baseline OCR benchmarks showing large differences between modern vision-language models and traditional engines.

  2. ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ARB provides 1,356 Arabic multimodal questions with 5,119 human-reviewed reasoning steps and shows leading models score much higher on reasoning fluency than on correct answers.

  3. QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation

    cs.CV 2025-06 reject novelty 4.0 of 10

    Fine-tuning Qwen2-VL on synthetic Arabic data yields QARI v0.2 with CER 0.061 and WER 0.160 on the authors' private test set, but public SARD results show Mistral OCR is more accurate.

Pith tools