REVIEW 3 major objections 3 minor 2 cited by
IslamicMMLU offers 10,013 multiple-choice questions to measure how well large language models know Quran, Hadith, and Islamic law.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 20:11 UTC pith:BLVYHSOV
load-bearing objection Useful domain MMLU for Islamic knowledge with a madhab-bias probe; the resource is real if the full paper documents curation and contamination—right now we only have the abstract. the 3 major comments →
IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A 10,013-question multiple-choice suite spanning Quran, Hadith, and Fiqh can expose large, model-to-model differences in Islamic knowledge (average accuracies 39.8–93.8 percent across 26 models) and can surface school-of-thought preferences that prior benchmarks could not measure.
What carries the argument
IslamicMMLU itself: three-track multiple-choice benchmark (2,013 Quran + 4,000 Hadith + 4,000 Fiqh items) plus a madhab-bias detection sub-task that scores preferential alignment with particular schools of Islamic jurisprudence.
Load-bearing premise
That the multiple-choice items are correctly sourced, labeled, and balanced across authentic Islamic texts, difficulty levels, and schools of law so that accuracy and bias scores truly reflect Islamic knowledge rather than artifacts of how the questions were written.
What would settle it
Independent expert re-annotation of a stratified sample of questions that finds systematic source or label errors large enough to reverse the reported accuracy rankings or madhab-bias patterns.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces IslamicMMLU, a 10,013-item multiple-choice benchmark spanning three tracks—Quran (2,013 questions), Hadith (4,000), and Fiqh (4,000)—intended to evaluate LLMs on core Islamic knowledge. Each track is described as containing multiple question types. The authors report initial evaluation of 26 LLMs with average accuracies from 39.8% to 93.8% (highest: Gemini 3 Flash), note a wide span on the Quran track (32.4%–99.3%), and introduce a madhab (school-of-jurisprudence) bias detection task within Fiqh. Arabic-specific models are reported to underperform frontier models. Evaluation code and a public leaderboard are stated to be released.
Significance. If the items are correctly sourced, expert-validated, balanced, and free of contamination, IslamicMMLU would fill a clear gap in culturally grounded LLM evaluation and provide a reusable public resource. The madhab-bias task is a novel and potentially useful probe of school-of-thought preferences. Public code and a leaderboard would support reproducibility. These strengths, however, are conditional on methodological transparency that cannot be verified from the abstract alone; the contribution’s value therefore remains provisional pending full documentation of curation and validation.
major comments (3)
- [Abstract] The central claim that reported accuracies and madhab-bias scores measure Islamic knowledge rests on the unstated premise that the 10,013 items are correctly sourced from authentic materials, free of construction artifacts, balanced across difficulty and madhabs, and free of training-data leakage. The abstract supplies no source lists, item-construction rules, expert-validation protocol, inter-annotator agreement, or contamination checks. Without those details the resource contribution and the 39.8–93.8% accuracy range cannot be trusted to measure the intended construct rather than dataset artifacts or memorization.
- [Abstract (Fiqh / madhab-bias task)] The madhab-bias detection task is presented as a novel contribution, yet the abstract does not define how bias is operationalized (item design per school, scoring metric, madhab coverage, or controls for confounds such as language or difficulty). This methodology is load-bearing for the novelty claim and for interpreting the reported “variable school-of-thought preferences.”
- [Abstract (results / leaderboard)] Comparative claims—including the accuracy span, the underperformance of Arabic-specific models relative to frontier models, and the ranking of 26 systems—are given without evaluation protocol (prompting, few-shot setting, decoding), confidence intervals, or significance tests. These details are required to interpret the public-leaderboard results as meaningful rather than descriptive.
minor comments (3)
- [Abstract] The abstract states that each track comprises “multiple types of questions” but does not enumerate those types; a brief typology would help readers assess coverage.
- [Abstract] “Averaged accuracy across the three tracks” should clarify whether tracks are weighted equally or by item count (Quran is half the size of Hadith/Fiqh).
- [Abstract] Model name “Gemini 3 Flash” should be checked for consistency with publicly documented model identifiers at the time of submission.
Circularity Check
No logical circularity: empirical benchmark paper with no derivation chain, fitted parameters, or self-citation load-bearing claims.
full rationale
This is an abstract-only empirical resource paper introducing IslamicMMLU (10,013 MCQs across Quran/Hadith/Fiqh tracks) and reporting accuracies of 26 LLMs (39.8–93.8%) plus a madhab-bias task. There is no claimed first-principles derivation, no equations, no fitted parameters renamed as predictions, no uniqueness theorems, and no ansatz smuggled via self-citation. The abstract simply defines the benchmark by construction (question counts and tracks) and reports observed model accuracies; that is ordinary benchmark construction, not circular reasoning. Self-citation is absent from the available text. Usual benchmark risks (possible contamination, curation quality, construct validity of accuracy as 'Islamic knowledge') are correctness/validity concerns, not circularity of the kind enumerated in the analyzer rules. Per the hard rules, an honest non-finding of score 0 is required when the paper is self-contained as an empirical report and no specific reduction (Eq. X = Eq. Y by construction, or fitted input called prediction) can be exhibited. Full-text absence precludes deeper inspection but does not manufacture circularity from the abstract.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Multiple-choice accuracy on curated Quran/Hadith/Fiqh items is a valid proxy for LLM competence in Islamic knowledge.
- domain assumption Answer distributions on Fiqh items can reveal madhab (school-of-jurisprudence) bias in models.
- ad hoc to paper The 10,013 items are correctly sourced and labeled from authentic Islamic materials.
read the original abstract
Large language models are increasingly consulted for Islamic knowledge, yet no comprehensive benchmark evaluates their performance across core Islamic disciplines. We introduce IslamicMMLU, a benchmark of 10,013 multiple-choice questions spanning three tracks: Quran (2,013 questions), Hadith (4,000 questions), and Fiqh (jurisprudence, 4,000 questions). Each track is formed of multiple types of questions to examine LLMs capabilities handling different aspects of Islamic knowledge. The benchmark is used to create the IslamicMMLU public leaderboard for evaluating LLMs, and we initially evaluate 26 LLMs, where their averaged accuracy across the three tracks varied between 39.8% to 93.8% (by Gemini 3 Flash). The Quran track shows the widest span (99.3% to 32.4%), while the Fiqh track includes a novel madhab (Islamic school of jurisprudence) bias detection task revealing variable school-of-thought preferences across models. Arabic-specific models show mixed results, but they all underperform compared to frontier models. The evaluation code and leaderboard are made publicly available.
Forward citations
Cited by 2 Pith papers
-
HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering
HalluTruthQA contributes 2,400 expert-annotated Arabic QA examples with hallucination labels, error spans, human explanations, and candidate answers; evaluations show no open LLM leads across all four tasks.
-
AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge
Leading generative AIs are useful for introductory Islamic learning but unreliable as authorities on Fiqh, citations, and madhhab-sensitive rulings without human verification.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.