Pith. sign in

REVIEW 4 major objections 5 minor 11 references

Generative AI is useful for basic Islamic learning but not authoritative for rulings or research without primary-source checks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 13:41 UTC pith:2R3CIV7I

load-bearing objection Useful open-ended, dual-country caution study on Islamic AI; the qualitative citation/provenance failures land, but the 1–5 domain scores that underwrite the hierarchy are not independently checkable. the 4 major comments →

arxiv 2607.28237 v1 pith:2R3CIV7I submitted 2026-07-30 cs.AI

AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge

classification cs.AI
keywords Islamic educationgenerative AIhallucinationcitation fidelityQur'anHadithFiqhMadhhab
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Muslims increasingly ask chatbots for Qur'anic meaning, Hadith explanation, fiqh rulings, and pastoral advice, yet there has been little open-ended evidence on whether those answers are authentic and verifiable. This study put fifty realistic religious questions to six leading generative systems under ordinary use in Australia and the United Kingdom, then scored domain accuracy, citations, hallucinations, madhhab handling, uncertainty, and source provenance. Performance is strongest where scholarly consensus is broad—especially Qur'anic interpretation and ethical guidance—and weakest on fiqh reasoning and school-of-thought-sensitive topics. Incomplete Hadith numbers, unnamed scholars, unverifiable claims, and occasional overconfident rulings recur across tools; answers can also differ in references by geography or device. The authors conclude that current systems are assistive aids for introductory learning and discovery, not standalone authorities for religious rulings, education, or Islamic research without verification against authenticated sources and qualified scholars.

Core claim

Across six major generative AI systems and fifty open-ended Islamic questions, reliability is high on consensus domains (Qur'anic interpretation ~4.4/5, ethics above 4) and substantially lower on Fiqh and Madhhab-sensitive topics (~2.1–2.3), with widespread citation incompleteness and source centralisation; therefore the systems are valuable as assistive tools but not sufficiently reliable as authoritative sources for high-trust religious guidance or research without human and primary-source verification.

What carries the argument

A mixed-method open-ended evaluation framework that maps four research questions to domain accuracy/risk scores, citation and hallucination tallies, fiqh consistency and uncertainty handling, source-ecosystem mapping, and AU/UK geographical comparison on real user-collected responses.

Load-bearing premise

The authors' domain scores, risk ratings, and citation-issue counts validly measure authenticity even though participants freely chose tools, only five complete survey sets were analysed, some UK material was incomplete, and no gold-standard scholarly adjudication protocol or inter-rater reliability is reported.

What would settle it

Have qualified scholars independently re-score the same response corpus (or a larger released set) with a fixed gold-standard rubric for exact Hadith numbers, madhhab attribution, and ruling accuracy; if Fiqh/Madhhab scores rise to Qur'an-level reliability and incomplete citations become rare, the core claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • AI outputs on Islamic topics should be treated as starting points for verification, not as citable religious authority.
  • Educators can use AI for introductory themes but should require students to trace Qur'an, Hadith numbers, grading, and madhhab context.
  • Developers of Islamic AI should expose exact citations, authentication status, scholar/madhhab labels, and abstention on disputed issues.
  • Identical prompts may yield different supporting sources by region or device, so reproducibility and trust need geographic and retrieval testing.
  • Larger multilingual, madhhab-balanced, and retrieval-augmented benchmarks are needed to measure whether curated Islamic knowledge bases reduce these failures.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If source ecosystems stay centralised on a few highly indexed sites, minority and less-digitised scholarly traditions will keep being under-represented in everyday AI answers.
  • Product interfaces that force exact Hadith identifiers and named madhhab labels before displaying a ruling could cut the dominant failure mode faster than waiting for base-model gains.
  • Pastoral and self-help style fluency may increase user trust precisely where evidential quality is weakest, raising a design tension between helpful tone and appropriate uncertainty.
  • Device- and region-dependent retrieval implies that "the model said" is incomplete; audits should log the retrieval path, not only the final text.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper empirically evaluates six generative AI systems on fifty open-ended Islamic questions spanning Qur’an, Hadith, Fiqh, ethics, pastoral advice, and Madhhab-sensitive topics, with responses collected under unrestricted real-world use in Australia and the UK. Using a mixed-method framework (domain accuracy, citation verification, hallucination analysis, jurisprudential consistency, uncertainty handling, source provenance, and geographic comparison), it addresses four RQs on authenticity, hallucinations/citations, madhhab/uncertainty handling, and suitability for high-trust religious use. The central claim is that current systems are comparatively reliable on broad-consensus domains (especially Qur’anic interpretation and ethical guidance) but substantially weaker on Fiqh and Madhhab-sensitive reasoning, with recurring incomplete or unverifiable citations; they are therefore useful as assistive tools for introductory learning but should not be treated as authoritative without verification against primary sources and qualified scholars.

Significance. The topic is timely and under-studied: Muslims increasingly consult LLMs for religious guidance, yet open-ended, citation- and madhhab-aware evaluation remains scarce relative to multiple-choice Islamic benchmarks (IslamicMMLU, FiqhQA, etc.). The paper’s strengths include a realistic question bank, explicit attention to citation fidelity and ikhtilaf, qualitative documentation of failure modes (missing Hadith numbers, device-dependent sources, a Tajweed error), and a clear assistive-vs-authority recommendation that is practically useful for educators and developers. If the domain hierarchy and citation findings hold under tighter methods, this would be a useful early empirical baseline for Islamic AI reliability. Significance is currently limited by incomplete reporting of scoring protocols and partial data use, which weaken the quantitative claims that underwrite the strongest conclusions.

major comments (4)
  1. [§6.1, Figures 7–8] §6.1 and Figures 7–8 present domain performance (~4.4 Qur’an, ~3.0 Hadith, ~2.1 Fiqh, ~2.3 Madhhab) and risk scores that directly underwrite RQ1/RQ4 and the assistive-only conclusion, but the manuscript does not report a gold-standard answer key, named scholarly adjudicators, a coded accuracy/authenticity rubric with decision rules, or inter-rater reliability. Without these, the numeric hierarchy cannot be independently checked as a measurement of authenticity. Either release the coding protocol and reliability statistics, or demote Figures 7–8 to illustrative summaries and rest the central claim primarily on the qualitative citation/hallucination evidence (Table 5, Fig. 10).
  2. [§4, Table 1] §4 states that conclusions are based on only five of ten survey sets “where we received complete responses,” and Table 1 notes UK material as incomplete/short for some sets (e.g., Set 4). Fifty questions are advertised in the abstract and introduction, but the analysed corpus is substantially smaller and geographically imbalanced. This is load-bearing for general claims about “current generative AI” across domains. The paper should report exact n per set/region/tool, justify sufficiency for each RQ, and avoid implying a full 50-question balanced design unless the missing sets are completed or clearly scoped out.
  3. [§4.2, §6.2, Figure 9] §4.2 allowed unrestricted participant choice of platform, subscription, and interaction style, yet §6.2/Figure 9 reports comparative thematic scores across ChatGPT, Gemini, Claude, Copilot, DeepSeek, and Dola AI (fluency, citation, fiqh consistency, uncertainty). Free tool choice confounds model identity with prompting depth, user skill, and retrieval settings, so model-to-model rankings are not causally interpretable. Either reframe Figure 9 as descriptive “fingerprint” patterns under naturalistic use (consistent with Table 2) and drop ranked model superiority language, or add a controlled same-prompt re-query study that isolates model family.
  4. [§6.3, Table 5] Table 5 and related text label issues as “Common,” “Frequent,” “Occasional,” or “Rare” without counts, denominators, or per-model rates. RQ2 asks “to what extent” systems produce hallucinations and unverifiable references; ordinal labels alone cannot answer extent. Provide tallies (e.g., fraction of Hadith-bearing answers missing numbers; fraction of fiqh answers naming madhhab/scholar) over the analysed responses so the citation-reliability claim is quantifiable and reproducible.
minor comments (5)
  1. [Keywords] Keywords (“Islamic education; curriculum; pedagogy; higher education; educational technology”) do not match the paper’s actual focus on AI reliability, hallucination, and source fidelity in Qur’an/Hadith/Fiqh. Replace with topic-appropriate terms.
  2. [§5.1–5.5] Figures 2–6 are described as “combined” multi-panel analyses but are not reproducible from the text (axes, coding units, and how similarity levels were computed are unclear). Add brief methods for how structural/thematic similarity was scored.
  3. [§5; Acknowledgements] Minor prose/typo issues: e.g., “Hallucination and references issued found” (§5); awkward Acknowledgements list formatting; inconsistent capitalization of madhhab/Madhhab.
  4. [§1] Related-work benchmarks (IslamicMMLU, FiqhQA, IslamTrust, inheritance) are well cited in the introduction; ensure in-text claims about their scores match the cited preprints and note multiple-choice vs open-ended difficulty when contrasting with this study.
  5. [§6.5, Figure 12] Geographic AU vs UK variation is an interesting contribution (§6.5, Fig. 12) but remains exploratory; avoid over-generalizing “retrieval environment” effects from two countries and small device examples without controlled IP/region tests.

Circularity Check

0 steps flagged

No significant circularity: empirical evaluation whose conclusions are not forced by construction from fitted inputs or self-citation chains.

full rationale

This manuscript is a mixed-method empirical evaluation of generative AI on open-ended Islamic questions (50 prompts; analysis restricted to five complete sets; AU/UK collection). It does not present a first-principles derivation, uniqueness theorem, or fitted-parameter-to-prediction pipeline. Domain scores (e.g., Fig. 7 ~4.4 Qur’an vs ~2.1 Fiqh), risk ratings (Fig. 8), thematic bars (Fig. 9), and hallucination tallies (Table 5) are author-coded judgments of collected responses, not algebraic identities with the inputs. References (IslamicMMLU, FiqhQA, IslamTrust, FACTS, SimpleQA, etc.) are external benchmarks used for motivation, not load-bearing self-citations that force the central assistive-vs-authority claim. Methodological concerns—free tool choice, incomplete UK material, lack of reported gold-standard key or inter-rater protocol—affect reproducibility and correctness risk, not circularity. No step reduces a claimed prediction to its defining inputs by construction. Score 0; steps empty.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 2 invented entities

The central claim rests on ordinary empirical-social-science assumptions plus author-defined notions of Islamic authenticity and risk, not on free physical parameters or new ontological entities. Load-bearing premises are that the collected real-world answers represent current system behavior, that incomplete citation and fiqh simplification are the right reliability constructs, and that author thematic coding tracks scholarly authenticity.

free parameters (3)
  • Domain performance scores (approx. 1–5 radar values) = ~4.4 Qur’an; ~4.0 ethics; ~3.8 pastoral; ~3.0 Hadith; ~2.3 Madhhab; ~2.1 Fiqh
    Numeric domain scores (e.g., Qur’an ~4.4, Fiqh ~2.1) function as fitted/assigned summary statistics from author evaluation rather than instrumented measurements with a published scale calibration.
  • Risk scores (1–5) per Islamic domain = Qur’an ~1.5 through Fiqh ~4.7
    Risk ratings (~1.5 to ~4.7) are author-assigned aggregates without a stated probabilistic loss model or external calibration set.
  • Thematic model scores for fluency/citation/fiqh/uncertainty = 1–5 ordinal judgments per system/dimension
    Figure 9 scores (e.g., ChatGPT fluency 4.7, Gemini citation 3.8) are judgment-derived comparators, not automatic metrics.
axioms (5)
  • domain assumption Authentic Islamic answers require traceable Qur’an/Hadith attribution, recognition of ikhtilaf/madhhabs, and responsible uncertainty/abstention—not fluency alone.
    Stated throughout Motivation, Problem Statement, and evaluation criteria mapping (Table 4); defines the success standard.
  • ad hoc to paper Real-world unrestricted user–tool interaction is a valid evaluation setting for high-trust religious use, even if it confounds controlled model comparison.
    §4.2 explicitly allows any public AI system and subscription level to simulate ordinary use.
  • ad hoc to paper Analyzing five complete survey sets is sufficient to support the four research questions and general conclusions about current generative AI.
    §4 states conclusions will be based on 5 sets with complete responses; Sets 6–10 are listed but not comparably analyzed.
  • domain assumption Overlap with a small ecosystem of indexed Islamic websites (Quran.com, Sunnah.com, IslamQA, etc.) substantially shapes AI outputs and thus authenticity risk.
    §6.4 provenance analysis; plausible but inferred from repeated motifs rather than logged retrieval traces.
  • standard math Standard mixed-method thematic comparison plus citation checking can ground claims about hallucination and jurisprudential consistency.
    Method framing in §5; ordinary qualitative/quantitative social-science practice, not a formal proof system.
invented entities (2)
  • AI tool fingerprint patterns (Table 2) no independent evidence
    purpose: Attribute stylistic differences (e.g., Claude academic, DeepSeek categorical) independently of factual correctness.
    Heuristic labels from observed prose; useful descriptively but not a validated latent construct with external tests.
  • Source ecosystem centralisation (as a named phenomenon in §6.4) no independent evidence
    purpose: Explain repeated interpretations and underrepresentation of minority/less-digitized scholarship.
    Interpretive synthesis of common URLs/themes; no independent crawl ranking or retrieval-log causal identification in the paper.

pith-pipeline@v1.2.0-daily-grok45 · 24601 in / 3928 out tokens · 80559 ms · 2026-07-31T13:41:17.833163+00:00 · methodology

0 comments
read the original abstract

Generative Artificial Intelligence (AI) is increasingly used by Muslims for religious guidance, Qur'anic interpretation, Hadith explanation, jurisprudential rulings, and Islamic education. Despite its growing adoption, there is limited empirical evidence on whether current AI systems provide authentic, verifiable, and trustworthy Islamic knowledge suitable for high-trust religious contexts. This study evaluates six leading generative AI systems using fifty realistic open-ended Islamic questions covering Qur'anic interpretation, Hadith, Fiqh, ethics, pastoral advice, and Madhhab-sensitive topics. Responses were collected under real-world conditions from participants in Australia and the United Kingdom and analysed using a mixed-method framework examining domain accuracy, citation verification, hallucinations, jurisprudential consistency, uncertainty handling, source provenance, and geographical variation. The study addresses four research questions: (1) How accurate and authentic are AI-generated responses across major Islamic knowledge domains? (2) To what extent do AI systems produce hallucinations, incomplete citations, or unverifiable religious references? (3) How consistently do models handle jurisprudential disagreement, Madhhab diversity, and uncertainty? (4) Are current AI systems sufficiently reliable for religious guidance, Islamic education, and scholarly research? Overall, current generative AI systems are valuable as assistive tools for introductory Islamic learning but should not be treated as authoritative sources for religious rulings or Islamic research without verification against authenticated primary sources and qualified scholarly expertise. This study provides one of the first comprehensive empirical evaluations of AI reliability within Islamic knowledge, offering practical guidance for researchers, educators, AI developers, and the wider Muslim community.

Figures

Figures reproduced from arXiv: 2607.28237 by Muhammad Sajjad Akbar.

Figure 1
Figure 1. Figure 1: Generative AI adoption and reliability challenges across general factuality bench [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Combined analysis of Set 1 responses, showing structural similarity, common Islamic [PITH_FULL_IMAGE:figures/full_fig_p020_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Combined analysis of Set 2 responses showing music-related jurisprudential patterns, [PITH_FULL_IMAGE:figures/full_fig_p023_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Combined analysis of Set 3 responses showing jurisprudential complexity, compara [PITH_FULL_IMAGE:figures/full_fig_p026_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Combined analysis of Set 4 responses showing similarity levels, comparative Fajr [PITH_FULL_IMAGE:figures/full_fig_p029_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Combined analysis of Set 5 responses showing narrative and emotional characteris [PITH_FULL_IMAGE:figures/full_fig_p032_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Performance of the evaluated AI systems across six Islamic knowledge domains. [PITH_FULL_IMAGE:figures/full_fig_p037_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Risk assessment of AI-generated responses across Islamic knowledge domains. [PITH_FULL_IMAGE:figures/full_fig_p038_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Comparative thematic evaluation of AI systems across fluency, citation completeness, [PITH_FULL_IMAGE:figures/full_fig_p039_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Examples of AI-generated Islamic responses exhibiting generic explanations and [PITH_FULL_IMAGE:figures/full_fig_p044_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Example illustrating a hallucination in a generative AI explanation of Tajweed rules. [PITH_FULL_IMAGE:figures/full_fig_p049_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Generative AI provides different sources for the same query across different devices, [PITH_FULL_IMAGE:figures/full_fig_p050_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

11 extracted references · 3 linked inside Pith

  1. [1]

    Global AI Adoption in 2025: A Widening Digital Divide , year =

  2. [2]

    2025 , institution =

    Freeman, James and Hall, Rachel and Pownall, Madeleine , title =. 2025 , institution =

  3. [3]

    2025 Microsoft AI in Education Report , year =

  4. [4]

    arXiv preprint arXiv:2501.03200 , year=

    The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input , author=. arXiv preprint arXiv:2501.03200 , year=

  5. [5]

    arXiv preprint arXiv:2411.04368 , year=

    Measuring short-form factuality in large language models , author=. arXiv preprint arXiv:2411.04368 , year=

  6. [6]

    JMIR Mental Health , volume=

    Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models: experimental study , author=. JMIR Mental Health , volume=. 2025 , publisher=

  7. [7]

    Journal of empirical legal studies , volume=

    Hallucination-free? Assessing the reliability of leading AI legal research tools , author=. Journal of empirical legal studies , volume=. 2025 , publisher=

  8. [8]

    arXiv preprint arXiv:2603.23750 , year=

    IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge , author=. arXiv preprint arXiv:2603.23750 , year=

  9. [9]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume=

    Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions , author=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume=

  10. [10]

    5th Muslims in ML Workshop co-located with NeurIPS 2025 , year=

    IslamTrust: A Benchmark for LLMs Alignment with Islamic Values , author=. 5th Muslims in ML Workshop co-located with NeurIPS 2025 , year=

  11. [11]

    Proceedings of The Third Arabic Natural Language Processing Conference , pages=

    Assessing large language models on islamic legal reasoning: Evidence from inheritance law evaluation , author=. Proceedings of The Third Arabic Natural Language Processing Conference , pages=