{"id":"fd9f3261-514e-4248-9d72-cf79c186c8ec","arxiv_id":"2608.04703","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"IslamicTurathBench is a new expert-reviewed Arabic benchmark that tests LLMs on classical Islamic scholarship across seven disciplines, three difficulty tiers, and three task formats.","lead":"This paper releases IslamicTurathBench, a benchmark dataset of 3,465 expert-authored Arabic questions for testing large language models on classical Islamic scholarship. It matters because it gives researchers a shared instrument for measuring AI performance across seven Islamic disciplines, three difficulty levels, and three question formats.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gold-answer faithfulness is unverified: the benchmark's validity rests on a small authoring team's bounded reference answers, with controlled plurality limited to 32 KNOW items; an independent audit could reveal systematic omission of legitimate scholarly answers.","rationale":"The reader's weakest assumption identifies gold-answer faithfulness and the erasure of plurality as the load-bearing risk. I agree that this is the central validity condition: the benchmark's entire evaluation layer assumes the released reference answers are both correct and exhaustive. The paper provides strong process evidence—expert authoring, a documented refinement audit, 112 corrections, Krippendorff's alpha of 0.81—but all of it comes from the same small team, and the 148-item human panel does not break the circle because its open-ended responses were scored by the same LLM-judge pipeline used for models (Section 5.7 and S6). The paper deserves credit for disclosing AI-assisted drafting, for bounding questions through source/school specification, and for releasing machine-readable artifacts with checksums and metadata. Those features make the benchmark reproducible, but reproducibility is not the same as gold-standard validity. The concern is concrete rather than speculative: only 32 KNOW items are permitted multiple acceptable answers, and the fiqh source list is Shafi'i-heavy, so the risk of systematic omission is localized and checkable. If an independent audit finds that legitimate alternative answers are absent from the gold lists, the headline scores and human-model comparisons would need re-interpretation. That is enough to keep the reader's CONDITIONAL verdict unchanged, but not enough to reject the dataset, which remains a substantial, well-documented resource.","tokens_in":23398,"tokens_out":4997,"duration_ms":63385,"concrete_test":"Randomly sample 60 KNOW items (20 per scholarly-demand tier, oversampling Fiqh), redact the gold answers, and ask three independent Islamic-studies scholars—not among the authors—to provide complete answers using only the question text. Have a separate expert panel compare each response against the released possible_item lists and judge whether any legitimate, source-grounded answer is missing or incorrectly marked wrong. If more than 10% of sampled items have an omitted valid answer, the bounded-gold design requires revision. Additionally, compute the school distribution of all Fiqh items from source_links.csv; if more than 80% trace to Shafi'i sources, the discipline label should be narrowed to 'Shafi'i jurisprudence' or supplemented with non-Shafi'i works.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ISTB is a valid evaluation instrument for classical Islamic scholarship rests on Section 2.4 and Table 6: gold answers are source-faithful, determinate, and pluralism-controlled. The supporting evidence is entirely internal: three professors and one MA-level researcher authored and reviewed all items, and the refinement audit (Section 4.3) corrected 112 items using flags generated by the authors' own LLM pipeline. No independent, external expert audit verifies that the bounded reference answers—especially the KNOW possible_item lists—do not exclude legitimate answers. The paper explicitly limits controlled plurality to 32 KNOW questions (Section 2.3); every other KNOW item is required to have 'one correct answer, or one complete set of required answer elements.' In a scholarly tradition where recognized schools legitimately differ, forcing a single gold list may systematically penalize valid alternative positions unless each question explicitly names one school, author, or source. The fiqh source list is almost entirely Shafi'i (five of six listed works are Shafi'i or Shafi'i-oriented), so broad 'Fiqh' scores conflate school-specific knowledge with general jurisprudence. Because every downstream metric—MCQ exact match, LLM-judge COMP/KNOW scoring, and the human-panel comparison—treats these gold answers as ground truth, an unrepresentative gold set would undermine the benchmark's central utility regardless of the sophistication of the scoring pipeline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents IslamicTurathBench (ISTB), an Arabic question-answering benchmark consisting of 3,465 expert-authored items linked to 35 classical and contemporary Islamic scholarly works across seven disciplines (Quran sciences, Hadith sciences, theology, jurisprudence, principles of jurisprudence, Sufism, and prophetic biography). Items are organized along a scholarly-demand axis (Beginner, Intermediate, Advanced) and a task-format axis (MCQ, passage-based comprehension, open-ended knowledge questions). The release includes metadata, validation artifacts, aggregated scores from a 148-item human reference panel, and zero-shot baselines from ten LLM-based systems, with an LLM-judge pipeline for open-ended scoring. The paper claims the dataset supports reproducible, multi-granular evaluation of LLM behavior in the Islamic scholarly tradition.","tokens_in":23705,"tokens_out":5790,"duration_ms":72338,"significance":"If the dataset's validity holds, ISTB is a valuable resource: it is source-linked, discipline-balanced, metadata-rich, and publicly released with integrity manifests, schemas, a loader, and evaluation code. The authors are transparent about construction choices, including AI-assisted drafting as a bounded aid, a documented refinement audit, and explicit source-bounding of the Fiqh items. The inclusion of a human reference panel and a judge-selection screen with expert review are additional strengths. The main risks are external: the gold answers have not been independently audited, and the human reference scores are generated by the same LLM-judge pipeline used for model evaluation. These issues are fixable and do not invalidate the dataset's potential, but they need to be addressed before the benchmark can support fine-grained capability claims.","major_comments":[{"comment":"The central validity claim—that gold answers are source-faithful and that legitimate scholarly plurality has been adequately controlled—rests entirely on internal review. The refinement audit in Section 4.3 corrected 112 items using flags generated by the authors' own LLM pipeline (e.g., absolute judge disagreement ≥0.5 between GPT-5.2 and Gemini-2.5-Flash, low-score consensus, and passage mismatches), and no independent external expert audit verifies that the KNOW reference lists (possible_item_* and no_requested_items) do not omit legitimate answers from recognized schools or authors. Because every downstream metric treats these gold answers as ground truth, I recommend an external, non-author audit of a stratified sample of items, with the audit protocol and any disagreements reported in the paper.","section":"Section 2.4, Section 4.3, Table 6"},{"comment":"The human reference scores are generated by the same LLM-judge protocol used to score model responses, and raw human responses are excluded from the release. This makes the headline human–model comparison (Table 13) dependent on the judge models' scoring preferences rather than on an independent measure of scholarly quality. If the judge pipeline has systematic biases, those biases can affect human and model responses differently, so the claim that only Gemini-3-Pro exceeds the human panel on KNOW may be an artifact of judge behavior. Please re-score a sample of human and model KNOW/COMP responses with blinded external experts and report agreement, or release de-identified raw human responses so the community can independently re-score them.","section":"Section 2.5, Section 5.3, Table 13"},{"comment":"The Fiqh discipline is almost entirely Shafi'i-school material: five of the six listed Fiqh sources are Shafi'i or Shafi'i-oriented (Minhaj al-Talibin, Ans al-Matalib, Hashiyat al-Ramli, Mughni al-Muhtaj, al-Fiqh al-Manhaji), and the sixth (Ihya' 'Ulum al-Din) is not a Fiqh manual. Yet results are reported under the unqualified label 'Fiqh' in Table 10 and Figure 5. The caveat in Section 2.2 that Fiqh items are strictly bound by the selected sources is not carried through to the results, so users may read Shafi'i-specific knowledge as general jurisprudence. Please rename the axis (e.g., 'Fiqh (Shafi'i school)') or add a prominent note to all Fiqh results and to the dataset README.","section":"Table 2, Figure 5, Table 10"}],"minor_comments":[{"comment":"Please state how the 95% confidence intervals in Table 10 were computed (e.g., bootstrap over items, over cells, or over design cells) and whether the cell-balanced macro-average variance was estimated accordingly.","section":"Section 5.4, Table 10"},{"comment":"Please clarify the unit of the reported interval Krippendorff's α (0.81): was it computed over flagged items, over correction-fix categories, or over expert judgments of the 462 candidate questions?","section":"Section 4.3"},{"comment":"Please define the sample sizes n=3,966 (COMP) and n=7,662 (KNOW) explicitly; if these are numbers of judged responses, it would also be helpful to report per-judge-pair agreement so readers can interpret the disagreement-escalation threshold |j1−j2|>0.2.","section":"Section 5.3"},{"comment":"The evaluation code and aggregate result files are released, but raw model responses and intermediate LLM-judge outputs are not; releasing at least per-response judge scores would substantially strengthen the reproducibility of Tables 10–13.","section":"Section 7"},{"comment":"The transliteration of the tradition term is inconsistent ('turath' in the title and once in the abstract, 'turāth' elsewhere); please standardize it.","section":"Title and Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a credible, well-documented benchmark release, and the dataset construction is the strongest part. The evaluation layer has a real but contained circularity, and the fiqh slice is Shafi'i-heavy in a way the paper discloses but users should take seriously.\n\nWhat's new is scope, not method. ISTB is the first Arabic benchmark I know that combines seven classical disciplines, three pedagogically-grounded demand tiers, and three task formats—MCQ, passage-grounded COMP, and closed-book KNOW—with items tied to 35 named source texts. That is a genuinely useful instrument for Arabic NLP and Islamic Studies. The release is properly engineered: JSON/CSV, SHA-256 manifest, schemas, loader, aggregate statistics. The construction pipeline is transparent: expert-authored items, manual chunk selection, AI-assisted drafting only under human filtering, a refinement audit that flagged 462 candidates and corrected 112 items, with an inter-rater alpha of 0.81. That is real work and it shows.\n\nThe soft spots are mostly in the evaluation layer. The human reference panel's open-ended scores come from the same LLM-judge pipeline used on the models. That makes the panel a matched comparison point rather than an independent gold standard—fine for model-to-human comparison, but it does not validate the judge. The paper should say this more plainly. The gold answers themselves rest on internal expert review; no independent audit verifies that the bounded reference lists do not exclude legitimate alternatives. Given the tradition, that is a genuine open question, though not fatal: the authors constrain questions by source, school, or author, and only 32 KNOW items use controlled plurality. Accepting that means the benchmark measures performance under a specific determinacy policy, not all of fiqh.\n\nThe fiqh selection is the point I'd push in revision. Five of six fiqh sources are Shafi'i-oriented, so the 'Fiqh' discipline marginal is really 'Shafi'i fiqh.' The paper says coverage is bounded by selected works, but the label will mislead casual users. Either rename the dimension or add a prominent warning in the data card.\n\nWho should read this: anyone building or evaluating Arabic religious-domain models, and anyone designing benchmarks for pluralistic scholarly traditions. The dataset is a solid start and an honest one. It deserves rigorous refereeing, with requests for clearer framing on the human-panel scoring and the fiqh coverage.","headline":"Credible, well-documented benchmark release; the dataset construction is the strong part, and the evaluation layer has a contained circularity plus a Shafi'i-heavy fiqh slice that the paper discloses but should frame more carefully.","tokens_in":24223,"tokens_out":3189,"would_cite":true,"duration_ms":34061,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"IslamicTurathBench is an expert-built Arabic benchmark of 3,465 questions that ties language-model scores to specific classical source works, disciplines, difficulty tiers, and answer formats in Islamic scholarship.","keywords":["LLM evaluation benchmark","Islamic scholarly tradition (turath)","Arabic NLP","multiple-choice QA","reading comprehension","open-ended knowledge QA","religious and cultural domains","human reference panel"],"falsifier":"Take a stratified random sample of about two hundred items and have an independent panel of Islamic Studies scholars, chosen to cover multiple jurisprudential schools and not involved in the dataset's creation, answer them directly from the cited source works without seeing the gold answers. If a substantial fraction of gold answers (say more than five percent) are judged unsupported by the named source text, or if the accepted-answer sets of KNOW items prove incomplete against the sources, then the ground-truth layer is not stable enough to support the reported comparisons.","tokens_in":1901,"feed_emoji":"📚","tokens_out":3329,"duration_ms":129673,"temperature":0.7,"pith_summary":"IslamicTurathBench (ISTB) is a new Arabic benchmark for measuring how well large language models handle the classical Islamic scholarly tradition (turath). It contains 3,465 expert-written question-answer items drawn from 35 recognized source works spanning more than twelve centuries and seven disciplines: Quran sciences, Hadith sciences, theology, jurisprudence, principles of jurisprudence, Sufism, and Prophetic biography. The benchmark's central wager is that a 3 x 3 design -- three scholarly-demand tiers (Beginner, Intermediate, Advanced) crossed with three task formats (multiple choice, passage-based comprehension, and closed-book open-ended questions) -- can separate what a model knows from how the question is asked. A 148-item scholarly human reference panel and zero-shot baselines from ten systems accompany the release. If the benchmark works as intended, it gives researchers a reproducible instrument for diagnosing where models fail on source-grounded religious scholarship rather than on generic web-style religious talk.","feed_headline":"Benchmark pits LLMs against 1,200 years of Islamic scholarship","feed_subtitle":"Expert-reviewed Arabic questions show models answer well with passages but poorly from memory.","key_machinery":"The load-bearing object is the bounded gold-answer schema. Multiple-choice items carry a single keyed option; comprehension (COMP) items tether the answer to a supplied source passage; and closed-book knowledge (KNOW) items encode reference answers as finite accepted sets (possible_item_1 through possible_item_8 plus a no_requested_items threshold demanding one, several, or all elements). This determinacy constraint is what makes open-ended scoring tractable in a tradition with legitimate scholarly plurality: when disagreement is real, the question must name its source, author, or school. On top of this, the aggregate metric $$\\mathrm{ISTB}(m)=\\frac{1}{|C|}\\sum_{(d,t,k)\\in C} S_{d,t,k}(m)$$ averages per-discipline, per-format, per-tier cell means so that no large item group (MCQ holds 2,276 of the 3,465 items) dominates the headline score. Open-ended answers are scored by a reference-guided LLM-as-a-judge pipeline with two judges and a head judge for disagreements above 0.2, with inter-judge agreement near 0.94.","core_discovery":"The central claim is that language-model performance on the classical Islamic scholarly tradition is measurable, source-attributable, and systematically task-dependent, and that the field has lacked a resource able to show this. The paper argues that ISTB supplies that resource: every item is traceable to a named source work, each question is constrained to a determinate reference answer so that scoring stays stable despite legitimate scholarly plurality, and the aggregate ISTB score is a cell-balanced macro-average giving equal weight to every discipline-format-demand cell. On the paper's own reading of its baselines, the data show three things: passage-grounded comprehension scores highest (mean 0.843 across systems), closed-book open-ended knowledge questions score lowest (0.574), and performance falls monotonically from Beginner (0.819) to Advanced (0.672). Against the matched human subset, models exceed the scholarly panel on passage-grounded questions, while on closed-book knowledge questions only the strongest reported system tops the panel, and the paper takes this as evidence that single-number leaderboards would hide the real behaviour.","pith_inferences":["The consistent KNOW deficit across all ten systems suggests that parametric memory of the classical tradition -- vintage- and source-specific knowledge -- is the binding constraint; a testable prediction is that retrieval augmentation will narrow the KNOW gap far more than the MCQ or COMP gaps.","The jurisprudence items are built on a Shafi'i-majority source list, so a 'Fiqh' marginal should be read as a school-specific sample; reading it as cross-school coverage would overstate the benchmark's reach.","The per-source author death dates in the metadata make possible a historical-gradient analysis -- plotting model accuracy against author death date to test whether later commentaries are systematically harder for models than foundational works.","The confidence-calibration results imply that deployed Islamic question-answering systems should not treat verbalized confidence as a reliability signal; abstention or source-citation policies could be evaluated with the same instrument."],"forward_implications":["Discipline, tier, format, and source-work filters let researchers draw a performance profile of any system against specific scholarly fields and specific historical texts.","Because format marginals diverge so sharply (COMP 0.843 vs KNOW 0.574 on average), a benchmark or leaderboard that reports a single number without format control will misrepresent where a model's competence actually lies.","The source-linked design turns ISTB into an evaluation harness for retrieval-augmented generation: one can index the documented corpus and compare closed-book against retrieval-supported answering on the same questions.","The matched human subset lets future work report human-versus-model comparisons holding task format fixed, rather than one aggregate ranking.","The 112 audit-corrected items, 32 controlled-plurality KNOW items, and published refinement protocol give users a visible quality-control trail for auditing the gold standard itself."],"supporting_citations":[{"why":"The closest comparable Islamic-domain benchmark; its multi-school jurisprudence scope and complexity rubric anchor the positioning of ISTB's multi-discipline design.","marker":"[4]"},{"why":"MCQ-only Islamic knowledge benchmark; supplies the comparison point demonstrating why a single format under-represents model capability.","marker":"[5]"},{"why":"Shared task whose Beginner/Intermediate/Advanced rubric and MCQ formats directly inform ISTB's demand axis and MCQ design.","marker":"[6]"},{"why":"Classical knowledge taxonomy cited as authority for consolidating the seven-discipline structure of the turath corpus.","marker":"[10]"},{"why":"Source of the pedagogical staging principle (al-tadarruj fi talab al-'ilm) that anchors the three scholarly-demand tiers.","marker":"[13]"},{"why":"Bloom's revised taxonomy, used to map each demand tier to a specific cognitive operation from remembering through creating.","marker":"[14]"},{"why":"Supplies the parametric versus non-parametric memory distinction motivating the closed-book versus passage-supported format axis.","marker":"[17]"},{"why":"The benchmark deposit itself: the released dataset, integrity manifest, source metadata, and aggregated human reference layer.","marker":"[24]"},{"why":"LLM-as-a-judge protocol adapted for reference-guided scoring of open-ended Arabic answers against gold rubrics.","marker":"[34]"}],"fun_headline_variants":["LLMs ace Islamic texts when given passages, stumble from memory","New benchmark shows LLMs lag on Islamic scholarship recall","IslamicTurathBench: 12 centuries of scholarship test LLMs","Expert benchmark: LLMs strong with passages, weak from memory"],"cache_read_input_tokens":26368,"weakest_assumption_plain":"The benchmark's validity rests on the assumption that the expert-authored gold answers and their difficulty labels faithfully represent what the selected turath works actually say, and that bounding every answer to a determinate reference has not erased legitimate scholarly plurality.","fun_headline_variants_meta":{"raw":{"variants":["LLMs ace Islamic texts when given passages, stumble from memory","New benchmark shows LLMs lag on Islamic scholarship recall","IslamicTurathBench: 12 centuries of scholarship test LLMs","Expert benchmark: LLMs strong with passages, weak from memory"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1425,"prompt_tokens":965,"completion_tokens":460,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":390}},"tokens_in":581,"tokens_out":460,"duration_ms":5778,"temperature":1.0,"reasoning_tokens":390,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:35:10.959066+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified random sample of about two hundred items and have an independent panel of Islamic Studies scholars, chosen to cover multiple jurisprudential schools and not involved in the dataset's creation, answer them directly from the cited source works without seeing the gold answers. If a substantial fraction of gold answers (say more than five percent) are judged unsupported by the named source text, or if the accepted-answer sets of KNOW items prove incomplete against the sources, then the ground-truth layer is not stable enough to support the reported comparisons.","supporting_citations":[{"cited_title":"& Iqbal, W","cited_arxiv_id":null,"evidence_quote":"The closest comparable Islamic-domain benchmark; its multi-school jurisprudence scope and complexity rubric anchor the positioning of ISTB's multi-discipline design."},{"cited_title":"IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge","cited_arxiv_id":"2603.23750","evidence_quote":"MCQ-only Islamic knowledge benchmark; supplies the comparison point demonstrating why a single format under-represents model capability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shared task whose Beginner/Intermediate/Advanced rubric and MCQ formats directly inform ISTB's demand axis and MCQ design."},{"cited_title":"Kitāb al-ʿIbar wa-dīwān al-mubtadaʾ wa-l-khabar fī tārīkh al-ʿArab wa-l-Barbar wa-man ʿāṣarahum min dhawī al -shaʾn al-akbar","cited_arxiv_id":null,"evidence_quote":"Classical knowledge taxonomy cited as authority for consolidating the seven-discipline structure of the turath corpus."},{"cited_title":"Taʿlīm al -mutaʿallim ṭarīq al -taʿallum","cited_arxiv_id":null,"evidence_quote":"Source of the pedagogical staging principle (al-tadarruj fi talab al-'ilm) that anchors the three scholarly-demand tiers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Bloom's revised taxonomy, used to map each demand tier to a specific cognitive operation from remembering through creating."},{"cited_title":"& Hajishirzi , H","cited_arxiv_id":null,"evidence_quote":"Supplies the parametric versus non-parametric memory distinction motivating the closed-book versus passage-supported format axis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The benchmark deposit itself: the released dataset, integrity manifest, source metadata, and aggregated human reference layer."}],"review_version":1}