Pith. sign in

REVIEW 3 major objections 3 minor 2 cited by

IslamicMMLU offers 10,013 multiple-choice questions to measure how well large language models know Quran, Hadith, and Islamic law.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 20:11 UTC pith:BLVYHSOV

load-bearing objection Useful domain MMLU for Islamic knowledge with a madhab-bias probe; the resource is real if the full paper documents curation and contamination—right now we only have the abstract. the 3 major comments →

arxiv 2603.23750 v3 pith:BLVYHSOV submitted 2026-03-24 cs.CL

IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge

classification cs.CL
keywords IslamicMMLUlarge language modelsbenchmarkQuranHadithFiqhmadhab biasArabic NLP
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

People already ask large language models for guidance on Islamic matters, yet there has been no large, systematic test of whether those answers are accurate. This paper introduces IslamicMMLU, a public benchmark of 10,013 multiple-choice questions covering the Quran, Hadith, and Fiqh (Islamic jurisprudence). The questions are designed to probe different facets of each domain, and the Fiqh track includes a new task that checks whether models systematically prefer one school of jurisprudence over others. Running 26 models on the suite produces average accuracies ranging from roughly 40 percent to nearly 94 percent, with the Quran track showing the largest spread. Arabic-specialized models do not automatically outperform general frontier systems. The authors release both the evaluation code and a public leaderboard so that future models can be scored the same way.

Core claim

A 10,013-question multiple-choice suite spanning Quran, Hadith, and Fiqh can expose large, model-to-model differences in Islamic knowledge (average accuracies 39.8–93.8 percent across 26 models) and can surface school-of-thought preferences that prior benchmarks could not measure.

What carries the argument

IslamicMMLU itself: three-track multiple-choice benchmark (2,013 Quran + 4,000 Hadith + 4,000 Fiqh items) plus a madhab-bias detection sub-task that scores preferential alignment with particular schools of Islamic jurisprudence.

Load-bearing premise

That the multiple-choice items are correctly sourced, labeled, and balanced across authentic Islamic texts, difficulty levels, and schools of law so that accuracy and bias scores truly reflect Islamic knowledge rather than artifacts of how the questions were written.

What would settle it

Independent expert re-annotation of a stratified sample of questions that finds systematic source or label errors large enough to reverse the reported accuracy rankings or madhab-bias patterns.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The manuscript introduces IslamicMMLU, a 10,013-item multiple-choice benchmark spanning three tracks—Quran (2,013 questions), Hadith (4,000), and Fiqh (4,000)—intended to evaluate LLMs on core Islamic knowledge. Each track is described as containing multiple question types. The authors report initial evaluation of 26 LLMs with average accuracies from 39.8% to 93.8% (highest: Gemini 3 Flash), note a wide span on the Quran track (32.4%–99.3%), and introduce a madhab (school-of-jurisprudence) bias detection task within Fiqh. Arabic-specific models are reported to underperform frontier models. Evaluation code and a public leaderboard are stated to be released.

Significance. If the items are correctly sourced, expert-validated, balanced, and free of contamination, IslamicMMLU would fill a clear gap in culturally grounded LLM evaluation and provide a reusable public resource. The madhab-bias task is a novel and potentially useful probe of school-of-thought preferences. Public code and a leaderboard would support reproducibility. These strengths, however, are conditional on methodological transparency that cannot be verified from the abstract alone; the contribution’s value therefore remains provisional pending full documentation of curation and validation.

major comments (3)
  1. [Abstract] The central claim that reported accuracies and madhab-bias scores measure Islamic knowledge rests on the unstated premise that the 10,013 items are correctly sourced from authentic materials, free of construction artifacts, balanced across difficulty and madhabs, and free of training-data leakage. The abstract supplies no source lists, item-construction rules, expert-validation protocol, inter-annotator agreement, or contamination checks. Without those details the resource contribution and the 39.8–93.8% accuracy range cannot be trusted to measure the intended construct rather than dataset artifacts or memorization.
  2. [Abstract (Fiqh / madhab-bias task)] The madhab-bias detection task is presented as a novel contribution, yet the abstract does not define how bias is operationalized (item design per school, scoring metric, madhab coverage, or controls for confounds such as language or difficulty). This methodology is load-bearing for the novelty claim and for interpreting the reported “variable school-of-thought preferences.”
  3. [Abstract (results / leaderboard)] Comparative claims—including the accuracy span, the underperformance of Arabic-specific models relative to frontier models, and the ranking of 26 systems—are given without evaluation protocol (prompting, few-shot setting, decoding), confidence intervals, or significance tests. These details are required to interpret the public-leaderboard results as meaningful rather than descriptive.
minor comments (3)
  1. [Abstract] The abstract states that each track comprises “multiple types of questions” but does not enumerate those types; a brief typology would help readers assess coverage.
  2. [Abstract] “Averaged accuracy across the three tracks” should clarify whether tracks are weighted equally or by item count (Quran is half the size of Hadith/Fiqh).
  3. [Abstract] Model name “Gemini 3 Flash” should be checked for consistency with publicly documented model identifiers at the time of submission.

Circularity Check

0 steps flagged

No logical circularity: empirical benchmark paper with no derivation chain, fitted parameters, or self-citation load-bearing claims.

full rationale

This is an abstract-only empirical resource paper introducing IslamicMMLU (10,013 MCQs across Quran/Hadith/Fiqh tracks) and reporting accuracies of 26 LLMs (39.8–93.8%) plus a madhab-bias task. There is no claimed first-principles derivation, no equations, no fitted parameters renamed as predictions, no uniqueness theorems, and no ansatz smuggled via self-citation. The abstract simply defines the benchmark by construction (question counts and tracks) and reports observed model accuracies; that is ordinary benchmark construction, not circular reasoning. Self-citation is absent from the available text. Usual benchmark risks (possible contamination, curation quality, construct validity of accuracy as 'Islamic knowledge') are correctness/validity concerns, not circularity of the kind enumerated in the analyzer rules. Per the hard rules, an honest non-finding of score 0 is required when the paper is self-contained as an empirical report and no specific reduction (Eq. X = Eq. Y by construction, or fitted input called prediction) can be exhibited. Full-text absence precludes deeper inspection but does not manufacture circularity from the abstract.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

Abstract-only review. No free parameters in the modeling sense; the work is a benchmark construction. Implicit domain assumptions are that multiple-choice accuracy is a valid proxy for Islamic knowledge competence and that madhab preference can be read off from answer distributions. No invented physical or mathematical entities. Full item-construction axioms (source corpora, expert panels, exclusion rules) are not stated in the abstract.

axioms (3)
  • domain assumption Multiple-choice accuracy on curated Quran/Hadith/Fiqh items is a valid proxy for LLM competence in Islamic knowledge.
    Standard MMLU-style assumption; not justified in the abstract beyond the existence of the tracks.
  • domain assumption Answer distributions on Fiqh items can reveal madhab (school-of-jurisprudence) bias in models.
    The novel bias task rests on this; abstract does not detail how school labels are assigned or controlled.
  • ad hoc to paper The 10,013 items are correctly sourced and labeled from authentic Islamic materials.
    Load-bearing for any claim that scores measure Islamic knowledge; construction process not described in abstract.

pith-pipeline@v1.1.0-grok45 · 6111 in / 2431 out tokens · 19406 ms · 2026-07-14T20:11:49.629333+00:00 · methodology

0 comments
read the original abstract

Large language models are increasingly consulted for Islamic knowledge, yet no comprehensive benchmark evaluates their performance across core Islamic disciplines. We introduce IslamicMMLU, a benchmark of 10,013 multiple-choice questions spanning three tracks: Quran (2,013 questions), Hadith (4,000 questions), and Fiqh (jurisprudence, 4,000 questions). Each track is formed of multiple types of questions to examine LLMs capabilities handling different aspects of Islamic knowledge. The benchmark is used to create the IslamicMMLU public leaderboard for evaluating LLMs, and we initially evaluate 26 LLMs, where their averaged accuracy across the three tracks varied between 39.8% to 93.8% (by Gemini 3 Flash). The Quran track shows the widest span (99.3% to 32.4%), while the Fiqh track includes a novel madhab (Islamic school of jurisprudence) bias detection task revealing variable school-of-thought preferences across models. Arabic-specific models show mixed results, but they all underperform compared to frontier models. The evaluation code and leaderboard are made publicly available.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

    cs.CL 2026-07 conditional novelty 6.0

    HalluTruthQA contributes 2,400 expert-annotated Arabic QA examples with hallucination labels, error spans, human explanations, and candidate answers; evaluations show no open LLM leads across all four tasks.

  2. AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge

    cs.AI 2026-07 conditional novelty 5.0

    Leading generative AIs are useful for introductory Islamic learning but unreliable as authorities on Fiqh, citations, and madhhab-sensitive rulings without human verification.