{"id":"377b362c-86ea-4384-881b-3ec0b29f6354","arxiv_id":"2608.03966","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"HalluTruthQA-4K is an expert-curated Arabic QA corpus of 4,000 instances that combines response-level hallucination labels, exact error spans, explanations, taxonomy, and multiple-choice factual verification.","lead":"This paper presents HalluTruthQA-4K, a 4,000-question Arabic QA dataset for detecting, locating, and explaining hallucinations in language model answers. It pairs each question with expert-verified answers, character-level error spans, error types, human explanations, and a six-option verification task across four knowledge domains.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4 reports 1,626 of 1,643 hallucinated responses have localized error spans, but the annotation protocol requires a span for every hallucinated response; the missing 17 conflict with the stated data format.","rationale":"The reader correctly identified the provisional status of the reported statistics and the non-independent agreement metric as weaknesses. I add a more specific, internally checkable inconsistency: Section 4's claim that only 1,626 of 1,643 hallucinated responses contain at least one localized span conflicts with the mandatory-span rule stated in Section 3.1 and the data-format description in Section 3.4. This is load-bearing because the corpus's unique value is span-level annotation; even a small number of hallucinated examples without spans contradicts the schema. The concern does not overturn the conditional verdict: the mismatch affects about 1% of the corpus and may be a counting typo or an annotation gap that the final release could fix. Requiring the span-count validation and confirmation of the final statistics is a proportionate condition for acceptance, so the reader's CONDITIONAL verdict should stand.","tokens_in":12866,"tokens_out":7711,"duration_ms":69540,"concrete_test":"Download a pinned/versioned release of HalluTruthQA-4K from the Hugging Face repository and run a validation script that counts records with label equal to hallucination whose hallucinations list is empty, printing the IDs and the corresponding span_start, span_end, and hallucinated_span fields. If the count is nonzero (in particular 17), then Section 4's 1,626/1,643 figure and the Section 3.1/3.4 guarantee that every hallucinated response carries a localized span cannot both be true. If the count is zero, the discrepancy is a manuscript typo and the corpus itself is consistent with the stated protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that HalluTruthQA-4K provides fine-grained, character-level span annotations for every hallucinated response. Under the protocol in Section 3.1, each hallucinated response must have at least one erroneous span, and for entirely unusable responses the complete response is annotated as the span. Section 3.4 likewise states that the hallucinations field is non-empty whenever label is hallucination. However, Section 4 reports: \"Of the 1,643 hallucinated responses, 1,626 (99.0%) contain at least one localized error.\" This leaves exactly 17 hallucinated responses with no localized span, which the protocol and schema make impossible. The discrepancy is either an annotation-rule violation, a counting error in the manuscript, or an undocumented exception. Because the resource's distinctive contribution is co-locating response-level detection with exact character-level localization, a 17-instance gap means the headline statistics are internally inconsistent. The issue is compounded by the explicit caveats in Section 4 and Table 4 that values marked with must be recomputed and confirmed using the final corpus release and complete annotation logs; until those checks are run, the reported 1,643 hallucinated responses, 1,843 spans, and kappa 0.92 are provisional. This is not a disagreement with external consensus but a direct conflict between the stated annotation scheme and the reported corpus statistics.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces HalluTruthQA-4K, a 4,000-instance Arabic question-answering corpus across Islamic knowledge, history, science, and geography, extending the authors' earlier HalluTruthQA benchmark. Each instance contains a question, a Fanar-1-9B-Instruct generated response, a verified reference answer, a binary hallucination label, six candidate answers, and, for hallucinated responses, character-level erroneous spans, human-written explanations, and a two-level hallucination taxonomy. The construction pipeline involves expert question and reference preparation, controlled generation, expert annotation, an independent review pass, and adjudication. The paper additionally reports corpus statistics, annotation agreement, taxonomy distributions, and the corpus's role as the official dataset for Track 2 of the HalluScoring 2026 shared task. The headline claims are 1,643 hallucinated responses, 2,357 non-hallucinated responses, and 1,843 annotated erroneous spans.","tokens_in":13056,"tokens_out":5850,"duration_ms":48408,"significance":"If the reported statistics hold after verification, HalluTruthQA-4K would be a substantial and reusable resource for Arabic hallucination research, uniquely aligning response-level detection, span-level localization, explanation generation, and candidate-based factual verification on the same instances. The paper is transparent about its construction methodology, including the controlled generation setting, the decision rules for boundary cases, and the explicit caveat that several values must be confirmed with the final release. The public release and the shared-task adoption increase the potential impact. However, the current manuscript contains an internal inconsistency between the annotation protocol and the reported span counts that must be resolved before the resource's central claim can be trusted.","major_comments":[{"comment":"Section 3.4 states that the hallucinations field is empty only when label is no_hallucination, and Section 3.1 with Table 3 requires a span for every hallucinated response, including entirely unusable responses where the complete response is annotated as the span. Yet Section 4 reports that only 1,626 of the 1,643 hallucinated responses (99.0%) contain at least one localized error, leaving 17 hallucinated responses without any span. This directly contradicts the data schema and the paper's central claim that every hallucinated response carries exact character-level localization. Please reconcile the protocol with the reported counts: either correct the number, document a legitimate exception for the 17 cases, or clarify that the 1,626 figure is preliminary and state the status of the remaining 17 instances in the final release.","section":"Section 3.4 and Section 4"},{"comment":"The agreement figures are computed between the initial domain-expert annotations and the independent verification pass, not between two independent expert annotators on the same data. Because each of the four experts annotated a disjoint domain subset, Cohen's κ = 0.92 for the binary label and the span F1/IoU values measure expert–reviewer consistency rather than standard inter-annotator reliability, and the expert retains final adjudication. The abstract's unqualified use of the term 'inter-annotator agreement' therefore overstates the strength of the reliability evidence. Please either report a small independent double-annotation sample, or consistently describe these values as expert–reviewer agreement in the abstract and throughout the paper.","section":"Section 3.1 and Table 4"}],"minor_comments":[{"comment":"The text says 'Values marked with must be confirmed' and 'All values marked with must be recomputed,' but no symbol appears to mark any particular value in the table or the body. Please insert the actual marker (e.g., an asterisk) and apply it to the specific provisional values.","section":"Table 4 and Section 4"},{"comment":"The macro-level distribution is first stated to be over the 2,400 training and development instances inherited from HalluTruthQA, but the subsequent domain-level percentages (e.g., 83.0% for geography) do not explicitly state that they are also restricted to that subset. Please clarify whether these percentages refer to the 1,037 hallucinated responses in the train/dev subset or to all 1,643 hallucinated responses in the full 4,000 corpus.","section":"Section 3.3"},{"comment":"The reference list contains two very similar entries for Huang et al. (2025a, ACM TOIS, and 2025b, with the same title). Please consolidate them or clearly differentiate the two works, and ensure that in-text citations match the intended sources.","section":"References"},{"comment":"The abstract refers to 'inter-annotator agreement' while Table 4 and the surrounding text use 'Expert–reviewer agreement.' Please align the terminology across the paper to avoid implying a different evaluation design.","section":"Abstract and Table 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a resource paper whose central value depends on the accuracy of its annotation statistics and the completeness of its span annotations. The reported 1,626/1,643 span-bearing hallucinated responses conflict with the protocol and schema, and this must be fixed before publication. The provisional nature of the headline numbers is disclosed, but the abstract presents them as final. The agreement metric is also weaker than the abstract suggests. These issues are local and fixable, so I view major revision as appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid resource paper. The 4K corpus with response-level labels, character-level spans, explanations, a two-level taxonomy, and six-option factual verification is a real step up from the binary-label Arabic datasets, and the authors are honest that the headline statistics are provisional. But the internal numbers do not fully hang together yet, and the statistics section needs a careful pass.\n\nThe genuinely new piece is the expansion to 4,000 instances with the joint annotation of exact spans, explanations, and candidate verification for every hallucinated response. That combination is not present in AraHalluEval or HalluScore, as far as the comparison table shows. The construction pipeline is standard but well described: expert-prepared questions, controlled generation with Fanar-1-9B-Instruct, expert annotation, independent verification, and adjudication. That is credible.\n\nThe soft spots are real but not fatal. The paper itself marks the distribution numbers and agreement values with \"must be recomputed and verified using the final corpus release.\" That caveat is responsible, but it means the abstract's 1,643 hallucinated responses, 1,843 spans, and kappa 0.92 are not yet established. The agreement metric is also expert-vs-reviewer, not independent expert-expert agreement, so the quality evidence is weaker than the kappa alone suggests.\n\nMore concretely, Section 3.4 says the hallucinations field is non-empty whenever label is hallucination, and Section 3.1 says entirely unusable responses get the whole response as the erroneous span. But Section 4 reports that 1,626 of the 1,643 hallucinated responses contain a localized error, leaving 17 with no span. That is a direct contradiction with the stated schema. It could be a counting error or an undocumented exception, but it needs to be fixed. It does not sink the resource, but it means the reported statistics are internally inconsistent.\n\nThe other analyses, like the chi-square test and length distributions, look fine and are appropriately descriptive. The citation pattern is reasonable; self-citation refers to the prior HalluTruthQA benchmark, which is disclosed and is a normal incremental step.\n\nWho is this for? Researchers working on Arabic hallucination detection, span-level localization, or factual verification, and anyone using the HalluScoring 2026 shared task. If I needed a fine-grained Arabic QA hallucination corpus, I would use it after the release is versioned and the numbers are confirmed.\n\nRecommendation: send to peer review. The resource deserves referee time, but the authors should be asked to reconcile the 17-instance discrepancy, replace provisional numbers with final ones (with a version or commit identifier), and clarify the agreement protocol.","headline":"A genuinely useful Arabic hallucination corpus, but the provisional statistics and a 17-instance span-count inconsistency need to be fixed before the numbers can be trusted.","tokens_in":13657,"tokens_out":2171,"would_cite":true,"duration_ms":19396,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HalluTruthQA-4K claims that Arabic hallucination evaluation should not stop at a response-level label, and provides 4,000 expert-curated QA instances that align detection with character-precise error spans, human explanations…","keywords":["Arabic hallucination detection","question answering","span-level error localization","factual verification","hallucination taxonomy","corpus annotation","LLM evaluation"],"falsifier":"Take a random sample of roughly 200 released instances, have independent Arabic-speaking experts re-annotate hallucination labels, error spans, and factual reference answers against the same source materials, and compare their annotations with the gold labels; if agreement falls well below the reported Cohen's $\\kappa = 0.92$ or the span-level character F1 of 0.88, the claim of a reliable fine-grained resource is not supported.","tokens_in":12637,"feed_emoji":"🔍","tokens_out":7216,"duration_ms":61246,"temperature":0.7,"pith_summary":"This paper argues that existing Arabic hallucination resources are limited because they mostly stop at a binary label for the whole response, leaving unanswered where the error is, why it is wrong, and what the correct fact is. It introduces HalluTruthQA-4K, a corpus of 4,000 expert-curated Arabic question-answering instances across Islamic knowledge, history, science, and geography, in which every generated response carries a hallucination label, exact character-level erroneous spans, human explanations, hierarchical hallucination types, and a verified reference answer with five plausible distractors. The corpus reports 1,643 hallucinated responses and 1,843 annotated error spans, and it serves as the official dataset for Track 2 of the HalluScoring 2026 shared task. The intended payoff is that the same instances support complementary evaluations, detection, localization, explanation, and factual verification, so a system that says an answer is unreliable can also be judged on whether it can say where, why, and what the truth is. This matters because fluent but factually unreliable Arabic model answers are hard to catch without source-sensitive evidence.","feed_headline":"Arabic QA corpus pinpoints hallucinated spans in 4,000 answers","feed_subtitle":"One benchmark ties labels, exact error spans, explanations, and verified answers for 4,000 Arabic QA pairs.","key_machinery":"The load-bearing mechanism is the aligned annotation record, especially the character-level span tuple with zero-based, end-exclusive offsets, so that slicing the generated answer from the start offset to the end offset reproduces the exact erroneous text. Around that tuple the instance binds a response-level hallucination label, a five-macro-type and 22-micro-type taxonomy, a human-written explanation, the verified gold answer, and six candidate options. This design turns one corpus into a multi-task testbed, because each annotation layer can be scored separately on identical inputs.","core_discovery":"The central claim is that fine-grained hallucination evaluation for Arabic question answering is achievable in one aligned benchmark, and HalluTruthQA-4K is that benchmark. Each of its 4,000 instances links a domain-expert-verified question and reference answer to a free-form response generated by a single Arabic-centric model, a binary hallucination label, and, when the response is hallucinated, one or more character-offset error spans with human-written explanations and macro- and micro-type hallucination categories. The corpus also adds six manually constructed answer options, one correct and five plausible distractors, so factual verification is tested by selecting the verified answer rather than by surface stylistic cues. The authors argue that the annotation layers capture distinct capabilities: a system may detect a hallucination yet fail to locate the erroneous span, explain it, or select the correct answer. The resource therefore defines evaluation tasks that measure not only whether a model rejects unreliable output but whether it can repair it.","pith_inferences":["If the annotation pipeline holds up on the final release, the same construction process could be adapted to other Arabic varieties and dialects, since the current corpus covers only one register and four knowledge domains.","The corpus makes it possible to test whether span-level supervision improves detection accuracy, for example by training one system on response-level labels and another on the same labels plus exact error spans, and comparing them on the held-out test split.","The taxonomy's distinction between error types, such as citation mismatch versus wrong source attribution, could support failure-pattern analysis of specific model families, showing whether hallucination types are model-specific or domain-driven.","The reported lexical overlap statistics suggest a ceiling on multiple-choice-only verification, so a useful extension would be measuring how often the correct option is selected when the model's own free-form answer is hallucinated, which the paper does not analyze."],"forward_implications":["Models can be evaluated on four linked tasks over the same 4,000 instances: response-level hallucination detection, span-level error localization, explanation generation, and multiple-choice factual verification.","Because hallucination rates differ by domain, from 49.4% in Islamic knowledge to 34.0% in history, domain-aware reporting becomes necessary for fair Arabic hallucination benchmarks.","Since 63.5% of annotated spans begin in the first quarter of a generated answer and most hallucinations occur inside otherwise fluent responses, response-level accuracy alone is insufficient for reliability evaluation.","The controlled single-generator design, using Fanar-1-9B-Instruct with fixed decoding settings, makes comparisons reproducible across systems on a common test set.","Because 54.7% of hallucinated answers match a candidate option exactly while 32.3% show no lexical overlap with any option, both multiple-choice verification and free-form span annotations are needed.","The corpus supports the official test data for Track 2 of the HalluScoring 2026 shared task, enabling comparable evaluation of Arabic hallucination-detection systems."],"supporting_citations":[{"why":"The original HalluTruthQA benchmark whose 2,400 instances this corpus extends and whose task dimensions motivate the design.","marker":"[Bouchekif et al., 2026c]"},{"why":"Provides the single Arabic-centric model, Fanar-1-9B-Instruct, used to generate every free-form response under the controlled configuration.","marker":"[Fanar Team et al., 2025]"},{"why":"HalluScore, the closest Arabic QA hallucination benchmark with verified evidence and reviewed explanations, whose partial localization is the gap this corpus fills.","marker":"[Alansari and Luqman, 2026a]"},{"why":"Mu-SHROOM, which frames hallucination detection as a span-labeling task and supplies the localization task formulation adopted here.","marker":"[Vazquez et al., 2025]"},{"why":"IslamicEval, which annotates complete intended Qur'anic and Hadith quotation spans rather than erroneous characters, defining the contrast for exact span annotation.","marker":"[Mubarak et al., 2025]"}],"fun_headline_variants":["Arabic QA corpus tags exact hallucinated spans in 4,000 answers","HalluTruthQA-4K: span-level annotations for 4,000 Arabic QA pairs","From label to location: Arabic QA benchmark localizes hallucinated spans","4,000 Arabic QA instances link error spans to explanations and verified facts","Fine-grained Arabic hallucination detection: exact spans, reasons, and correct answers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the accuracy of the expert annotations and the finality of the reported statistics; the paper itself warns in Section 4 that the values must be recomputed and verified on the final release, and the agreement metric compares initial experts with a reviewer pass rather than independent expert-expert agreement.","fun_headline_variants_meta":{"raw":{"variants":["Arabic QA corpus tags exact hallucinated spans in 4,000 answers","HalluTruthQA-4K: span-level annotations for 4,000 Arabic QA pairs","From label to location: Arabic QA benchmark localizes hallucinated spans","4,000 Arabic QA instances link error spans to explanations and verified facts","Fine-grained Arabic hallucination detection: exact spans, reasons, and correct answers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000631,"raw_usage":{"total_tokens":2966,"prompt_tokens":1050,"completion_tokens":1916,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":666,"completion_tokens_details":{"reasoning_tokens":1815}},"tokens_in":666,"tokens_out":1916,"duration_ms":13403,"temperature":1.0,"reasoning_tokens":1815,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:43:55.166487+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of roughly 200 released instances, have independent Arabic-speaking experts re-annotate hallucination labels, error spans, and factual reference answers against the same source materials, and compare their annotations with the gold labels; if agreement falls well below the reported Cohen's $\\kappa = 0.92$ or the span-level character F1 of 0.88, the claim of a reliable fine-grained resource is not supported.","supporting_citations":[],"review_version":2}