HalluScore is a curated Arabic QA dataset with 827 questions, ground-truth evidence, and human annotations used to measure hallucination rates across 17 LLMs.
Fanar: An arabic-centric multimodal generative ai platform
7 Pith papers cite this work. Polarity classification is still indexing.
years
2026 7verdicts
UNVERDICTED 7representative citing papers
A new instruction dataset for Arabic poetry tasks enables fine-tuned LLMs to generate poems matching user-specified styles, rhymes, and criteria, as shown by automated metrics and native-speaker evaluations.
LQM introduces a six-level linguistically motivated error taxonomy for MT evaluation and applies it via expert annotation to LLM outputs on a new 3,850-sentence multi-dialect Arabic corpus.
GLU is a single-pass unsupervised uncertainty score for LLMs formed by multiplying global hidden-state geometric entropy with local token entropy, shown to match or beat baselines on three model families and six benchmarks while catching failure modes local signals miss.
WASIL is a released dataset of Arabic spoken interactions with LLMs that includes audio, ASR outputs, responses, user feedback, and answerability labels to isolate ASR effects.
A metadata-conditioned mT5 model trained on rule-augmented dialectal Arabic data produces translations that better match intended regional varieties than high-resource baselines, despite lower BLEU scores.
Residual-stream noise injection raises narrative diversity in Arabic educational stories while preserving reading-grade level, outperforming high-temperature sampling across five 7-9B models.
citing papers explorer
-
HalluScore: Large Language Model Hallucination Question Answering Benchmark
HalluScore is a curated Arabic QA dataset with 827 questions, ground-truth evidence, and human annotations used to measure hallucination rates across 17 LLMs.
-
Instruction-Guided Poetry Generation in Arabic and Its Dialects
A new instruction dataset for Arabic poetry tasks enables fine-tuned LLMs to generate poems matching user-specified styles, rhymes, and criteria, as shown by automated metrics and native-speaker evaluations.
-
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
LQM introduces a six-level linguistically motivated error taxonomy for MT evaluation and applies it via expert annotation to LLM outputs on a new 3,850-sentence multi-dialect Arabic corpus.
-
Integrating Local and Global Entropy for Uncertainty Quantification in LLMs
GLU is a single-pass unsupervised uncertainty score for LLMs formed by multiplying global hidden-state geometric entropy with local token entropy, shown to match or beat baselines on three model families and six benchmarks while catching failure modes local signals miss.
-
WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
WASIL is a released dataset of Arabic spoken interactions with LLMs that includes audio, ASR outputs, responses, user feedback, and answerability labels to isolate ASR effects.
-
Context-Aware Dialectal Arabic Machine Translation with Interactive Region and Register Selection
A metadata-conditioned mT5 model trained on rule-augmented dialectal Arabic data produces translations that better match intended regional varieties than high-resource baselines, despite lower BLEU scores.
-
Noise Steering for Controlled Text Generation: Improving Diversity and Reading-Level Fidelity in Arabic Educational Story Generation
Residual-stream noise injection raises narrative diversity in Arabic educational stories while preserving reading-grade level, outperforming high-temperature sampling across five 7-9B models.