{"id":"41975e91-5e39-4b08-a2cd-de0c442f44b0","arxiv_id":"2608.04591","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs over-certify negative answers from partial evidence, especially when completeness is implied rather than stated, and prompting mainly trades over-closure for under-closure.","lead":"This paper introduces CROWN-QA, a benchmark that tests whether large language models can tell when evidence is complete enough to justify saying \"no\" versus when missing information should remain \"unknown.\" Across three model families, models often treat partial evidence as full coverage, and better prompting shifts the error type without fixing it.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The L3 implicit-partial gold labels rest on an untested semantic assumption; without a human baseline, the headline asymmetry could be a labeling artifact rather than a model failure.","rationale":"We examined whether the L3 label-validity concern is truly load-bearing or whether independent evidence in the paper would survive even if L3 labels were wrong. CROWN-Real's Proceedings component does show that a strict-subset variant C yields more over-closure than a narrower-complete variant B across 12 of 15 model-condition cells (Table 20), which supports a general 'partial evidence is harder than scope-mismatch' claim. However, the paper's headline asymmetry is specifically about implicit-partial evidence, and the CROWN-Real C variant is not the same as the L3 implicit-partial construction; the transfer claim relies on the L3 definition being valid. Moreover, the L3 results drive RQ2 and the Conclusion. We also checked whether the paper's two-label rule could rescue the labels: under the rule, ambiguous coverage should map to Unknown, so only phrases that actually convey complete coverage would invalidate the gold label. The authors' manual rejection is meant to catch that, but it is not an empirical demonstration. We found no other concern as load-bearing: the bootstrap intervals are properly pair-resampled, the paired construction is a genuine control, the certificate analysis is clearly labeled as diagnostic of reported fields, and the prompting conclusions are supported by item-level transitions. The reader's weakest assumption is the same one we identify, so we agree with the reader; the CONDITIONAL verdict remains appropriate and our stress-test does not move it. We recommend the human annotation study as the decisive check.","tokens_in":25210,"tokens_out":7578,"duration_ms":85122,"concrete_test":"Recruit at least 10 human annotators per item (or a representative stratified sample of, e.g., 50 L3 pairs) and present the exact question, listed titles, and coverage sentence from each L3 complete and partial member. Ask annotators to choose: (A) evidence is complete for the exact query scope, (B) evidence is partial/non-exhaustive, or (C) unclear. Map A to Certified-Negative and B/C to Unknown per the paper's rule, and compute agreement with the benchmark gold labels separately for complete and partial members, plus per-family agreement for the four partial families. If partial-member agreement with Unknown falls below a prespecified threshold (e.g., 85%) or any partial family is majority-rated as complete, the L3 asymmetry is not a valid measure of model reasoning.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The dominant failure reported in the abstract and conclusion—models often recognize implicitly complete evidence yet treat implicitly partial evidence as query-covering—is quantified almost entirely in the L3 regime of CROWN-Synth (Tables 5 and 14). The gold label Unknown for L3 partial members is assigned by the authors' controlled metadata: the source-type phrases 'activity feed,' 'update bulletin,' 'highlights digest,' and 'briefing summary' are stipulated to convey partial, non-exhaustive coverage to competent readers, while 'system of record,' 'master index,' 'canonical register,' and 'definitive index' are stipulated to convey completeness. Quality Control reports only that the authors 'manually inspect a stratified sample' and use an LLM judge as a secondary check; no human annotation study, agreement rate, or external norm is provided. Because the gold label is derived from the construction rather than from reader judgments, the L3 asymmetry could be an artifact: if one or more partial families are not reliably interpreted as non-exhaustive (or are interpreted as complete), then the partial member's Unknown label is wrong, and models that answer Certified-Negative are not over-closing. The paper's rejection of items that 'inadvertently establish exhaustive coverage' is an author judgment, not an empirical validation of the source-type semantics. Since the central claim is specifically about implicit-partial evidence, this is the load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper formalizes completeness-sensitive negative reasoning: a negative answer ('Certified-Negative') is justified only when the evidence completely covers the query scope; otherwise the correct answer is 'Unknown.' It introduces CROWN-Synth, a controlled paired benchmark with explicit, paraphrased, implicit, and adversarial coverage regimes plus scope-mismatch pairs, and CROWN-Real, an A/B/C contrast-set evaluation grounded in ACL Anthology proceedings and DailyMed drug labels. Three LLMs are evaluated under seven prompting conditions, with directional error rates, paired completeness sensitivity, structured certificate elicitation, and bootstrap confidence intervals. The principal finding is an asymmetric over-closure pattern: models often answer correctly on implicitly complete evidence but over-close on implicitly partial evidence, and prompting mostly redistributes the two error types rather than resolving them.","tokens_in":25530,"tokens_out":7859,"duration_ms":90412,"significance":"The task formulation is clean, practically motivated, and well positioned against prior abstention, sufficiency, and open-world work. The controlled paired design—same question, same observed facts, varying only query-relative coverage—is a genuine methodological advance, and the paper ships exact prompts, datasets, raw outputs, inference and scoring code, and bootstrap analyses. The main empirical claims are presented with confidence intervals and item-level decompositions, and the CROWN-Real Proceedings result is a useful transfer check. However, the headline asymmetry is only as strong as the L3 implicit-partial gold labels, which currently rest on the authors' stipulation about source-type semantics rather than on human judgments; this is the main risk to the paper's central contribution.","major_comments":[{"comment":"The headline asymmetry—that models recognize implicitly complete evidence but treat implicitly partial evidence as query-covering—is quantified in the L3 regime, where the gold label Unknown for the partial member is assigned by stipulating that the source-type families 'activity feed,' 'update bulletin,' 'highlights digest,' and 'briefing summary' convey partial, non-exhaustive coverage. The Quality Control section reports only manual inspection of a stratified sample and an LLM judge as a secondary check; no human baseline, agreement rate, or independent norm is provided, and an LLM judge could share the very over-closure tendency under study. Because Table 5 and Table 14 compute the complete-versus-partial asymmetry against these labels (e.g., Haiku Def.+CoT: complete 92.9% vs. partial 22.1%; Gemma partial 0.0% in several cells), a mis-calibrated label would make the asymmetry partly a construction artifact rather than a model reasoning failure. I ask for a human annotation study of the L3 items, or at minimum of the eight source-type families, reporting the proportion of competent readers who infer 'partial/non-exhaustive' for each family and reporting agreement; the inspected/rejected item counts should also be reported. If the labels are confirmed, the asymmetry stands; if not, the corresponding claims must be reframed as applying to the stipulated semantics only.","section":"CROWN-Synth coverage regimes; Quality Control; Tables 5 and 14"}],"minor_comments":[{"comment":"The Naive condition still instructs the model to 'classify whether the observed absence licenses a negative answer using only the provided evidence,' so it is a minimally framed baseline rather than a truly default zero-shot response; consider renaming it or adding an even less directive condition to support claims about default behavior.","section":"Models and Evaluation Conditions; Appendix B"},{"comment":"The 'earliest erroneous field' assignment compares model-reported fields against the constructed gold metadata, whose L3 evidence-coverage labels are the same stipulation discussed in the major comment; the text should state explicitly that the gold evidence scope is the authors' metadata rather than an independently validated annotation, and the RQ4 percentages should be interpreted conditionally on that metadata.","section":"Diagnostic Decomposition; Table 7"},{"comment":"The conclusion states that 'partial evidence is at least as error-prone as narrower-scope complete evidence across every model and condition' for CROWN-Real Proceedings, but the main-text Table 8 shows only four conditions per model; the full 15-cell support first appears in Appendix D, Table 16. Please point to the appendix table in the main text when making this claim.","section":"Conclusion; Appendix D Table 16"},{"comment":"The claim that 'explicit task rules improve Acc and CS for all three models' is visually supported by Table 3, but the statement could be sharpened by noting that the bootstrap intervals in Appendix D, Table 17 confirm the aggregate improvements while the directional OCR/UCR changes differ by model; the current wording already conveys this, so only a small clarification is needed.","section":"Overall Closure Profiles; Table 3"},{"comment":"Table 6 reports item-level accuracy changes for two transitions (Definition-aware to Definition-aware+CoT and Certificate to Certificate+CoT) without bootstrap intervals for those item-group contrasts; the appendix provides intervals for aggregate metrics, but it would be useful to state that the item-group changes are descriptive and may be less stable than the aggregate contrasts.","section":"Prompting Effects; Table 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within scope and the artifact release is a definite strength. My recommendation rests on one load-bearing point: the L3 implicit-partial labels need human validation before the headline asymmetry can be accepted as a model failure rather than a labeling stipulation. The requested validation is feasible within the manuscript's scope and would substantially increase the paper's impact."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thanks for sharing this. I read it carefully. The core contribution is real: CROWN-QA's paired construction fixes the question and observed facts and varies only whether the evidence covers the query scope. That is not present in AbsenceBench, SufficientContext, or the abstention benchmarks, and it's a meaningful step beyond showing that LLMs miss missing information. The L1-L4 regime design is thoughtful, especially the Scope-Mismatch cases that separate cue-spotting from containment reasoning. The empirical story is consistent: across three model families and seven prompting conditions, models over-close on partial evidence, and the L3 implicit-partial regime is the worst. The bootstrap intervals are careful, and the CROWN-Real transfer gives some external anchor.\n\nThe soft spot the reader flagged is real: the L3 gold labels for partial members rest on the stipulation that 'highlights digest,' 'activity feed,' 'update bulletin,' and 'briefing summary' read as non-exhaustive to competent readers. The paper does manual inspection plus an LLM judge, but no human annotation study or agreement rate. If those phrases don't license 'Unknown' for people, then the model's Certified-Negative isn't a reasoning failure, it's a disagreement with an unpublished norm. That's a legitimate threat to the headline asymmetry. I don't think it's fatal: the phrases are plausibly partial, and the asymmetry appears across all four partial families, which makes a systematic labeling error less likely. But the authors need to either run a quick human norming study or weaken the 'implicit' claim. Also, the dataset isn't linked in the preprint, which makes replication harder.\n\nOne more thing: the certificate analysis compares model-reported fields against the paper's own gold metadata. That's diagnostic rather than circular, but it means the 'evidence-coverage mischaracterization' conclusion is about consistency with the benchmark's annotation, not about human perception of what the evidence says. Worth stating clearly.\n\nBottom line: this is a solid, useful paper with one load-bearing assumption that needs empirical support. It deserves a serious referee. If the authors add a human baseline for L3 and release the data, the conditional verdict becomes a clear accept.","headline":"A well-designed paired benchmark that cleanly isolates query-relative coverage and finds a consistent over-closure failure, but the load-bearing implicit-partial labels lack a human baseline.","tokens_in":26011,"tokens_out":2707,"would_cite":true,"duration_ms":29286,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs treat missing items as confirmed absences even when the evidence is partial","keywords":["completeness-sensitive negative reasoning","certified negative","query-relative coverage","over-closure","abstention","coverage regimes","CROWN-QA","retrieval-augmented generation"],"falsifier":"A human-subject study on the L3 partial items would settle the central claim: present competent readers with the same question and listed titles under the four partial source-type phrases and ask whether the list could still fail to contain a relevant item. If a large share of readers treat 'highlights digest' or 'activity feed' as potentially exhaustive in context, the gold Unknown label is contested and the asymmetry would be partly a labeling artifact; conversely, if readers reliably read these phrases as partial, the model failure is confirmed.","tokens_in":25062,"feed_emoji":"🔍","tokens_out":7600,"duration_ms":70114,"temperature":0.7,"pith_summary":"The paper argues that a 'no' answer to a question about a record, list, or retrieved context is justified only when the evidence completely covers the question's scope; otherwise the correct answer is 'unknown.' It introduces CROWN-QA, a controlled paired benchmark that fixes the question and the observed facts while varying only query-relative coverage, to test whether large language models make this distinction. Across three model families, the benchmark finds substantial over-closure: models answer 'certified negative' on evidence that is actually partial, and the most persistent failure is asymmetric, concentrated on implicitly partial source-type framings such as 'activity feed' and 'highlights digest' rather than on the matched implicitly complete framings. The paper claims this failure matters because prompting interventions redistribute the errors between over- and under-closure instead of repairing the underlying judgment, and because the same partial-coverage gap appears on real-document content.","feed_headline":"LLMs overtrust partial lists and answer 'no' too often","feed_subtitle":"A paired benchmark shows models usually know a complete index licenses 'no,' but treat digests and feeds as complete too.","key_machinery":"The load-bearing object is CROWN-QA itself, built from two controlled components. CROWN-Synth is a paired core of 5,000 examples (2,500 pairs) in which the same question and identical observed factual content are matched across two members that differ only in whether the evidence closes the query scope, across four coverage-expression regimes: L1 explicit ('this is the complete list'), L2 paraphrased ('the official registry of all X'), L3 implicit (source-type phrases such as 'master index' versus 'highlights digest' with no completeness language), and L4 adversarial (a saliently mentioned complete source that does not license closure), plus Scope-Mismatch cases where evidence is complete only for a narrower scope. CROWN-Real is a 533-set contrast evaluation on ACL Anthology proceedings and DailyMed drug labels, where a fixed question and target phrase are paired with variant A (query-covering complete), variant B (complete for a narrower scope), and variant C (non-exhaustive). The formal criterion is the containment condition $c = 1$ if and only if the evidence establishes complete coverage of its asserted scope and $S(q) \\subseteq S(E)$, and the diagnostic instrument is a structured certificate $(\\hat{S}_q, \\hat{S}_E, \\hat{c})$ that separates query-scope extraction, evidence-coverage characterization, and the final Boolean judgment.","core_discovery":"CROWN-QA defines completeness-sensitive negative reasoning as an absence-conditioned judgment: given that the queried fact is unsupported, the label is Certified-Negative only when the evidence establishes complete coverage $c = \\mathrm{Comp}(E, S(q)) = 1$, meaning the evidence's scope $S(E)$ contains the query scope $S(q)$; otherwise the answer is Unknown. The central empirical discovery is an asymmetric closure failure on the CROWN-Synth L3 regime, where coverage is conveyed implicitly through source-type phrases: models answer the implicitly complete member ('the master index lists these titles') correctly 82.4-100.0% of the time, yet answer the matched implicitly partial member ('the highlights digest lists these titles') as Certified-Negative instead of Unknown, with partial-member correct rates of only 0.0-27.9%. The asymmetry holds across all four complete and four partial source-type families, and the same ordering -- partial evidence at least as error-prone as scope-mismatched complete evidence -- transfers to the real-document Proceedings contrast sets, where the partial-versus-narrower-complete gap is nonnegative in all 15 model-condition cells and excludes zero in 12. Prompting interventions (definition-aware rules, chain of thought, abstention, self-check, certificate elicitation) shift errors between over-closure and under-closure without a consistent repair, and structured certificates trace the earliest error most often to a mischaracterized evidence-coverage field rather than to the final judgment.","pith_inferences":["The asymmetry suggests models may be running a heuristic in which certain source phrases ('official registry,' 'master index') mean exhaustive while digest-style phrases do not; a direct follow-up would swap the source-type phrases across framings and check whether closure rates follow the phrases rather than the semantics.","A human baseline on the L3 partial items would settle whether the gold 'Unknown' labels match competent-reader intuitions; without one, part of the headline asymmetry could be a labeling artifact rather than a reasoning failure.","The certificate design hints at a training intervention the paper does not pursue: supervise the evidence-coverage field directly, since that is the field where errors first appear.","The real-document ordering -- partial lists more error-prone than narrower-but-complete lists -- suggests models have partially learned scope containment but not partiality inference, a distinction that could transfer beyond question answering to tool use and planning."],"forward_implications":["Retrieval-augmented systems that answer 'no' from whatever context was retrieved would produce unjustified negatives whenever retrieval is partial; the results imply that absent support alone never licenses negation.","High-stakes settings -- contraindication lists, exclusion lists, official indexes -- should treat 'not found' as 'unknown' unless the evidence shown is complete for the full query scope.","No tested prompting style (rules, chain of thought, abstention bias, self-check, structured certificates) consistently repairs partial-evidence over-closure across models, so prompt-level gains are not a reliable fix.","Because the earliest certificate error is most often a mischaracterized evidence-coverage field (45-72% of incorrect outputs), the paper points to coverage inference, not the final judgment, as the bottleneck to target.","Since partial evidence is at least as error-prone as narrower-scope complete evidence on real documents, completeness auditing of retrieved context should treat explicit partiality markers, not just scope mismatches, as a trigger for uncertainty."],"supporting_citations":[{"why":"Gives the retrieval-augmented-generation setting whose retrieved context is the evidence that CROWN-QA evaluates for completeness.","marker":"(Lewis et al. 2020)"},{"why":"Supplies the open-world versus closed-world completeness distinction that licenses a negative answer only under complete query-covering evidence.","marker":"(Razniewski et al. 2024)"},{"why":"The closest prior benchmark on omitted information, which CROWN-QA extends by asking whether missing support licenses negation rather than only what is missing.","marker":"(Fu et al. 2025)"},{"why":"The abstention benchmark that motivates the Unknown label; CROWN-QA distinguishes abstention from a licensed negative answer.","marker":"(Kirichenko et al. 2025)"},{"why":"The sufficient-context evaluation whose sufficiency judgment CROWN-QA refines into query-relative scope containment.","marker":"(Joren et al. 2025)"},{"why":"The negation benchmark documenting LLM difficulty with explicit negation, which CROWN-QA moves beyond to negative licensing.","marker":"(García-Ferrero et al. 2023)"}],"fun_headline_variants":["LLMs say 'no' when evidence is partial, not complete","Models can't tell 'absent' from 'not found' in partial lists","When absence is weak evidence, LLMs still answer definitively","LLMs overtrust lists that only show part of the story","Completeness-blind: LLMs treat missing data as confirmed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the L3 implicit-partial gold labels are valid, meaning source-type phrases such as 'activity feed,' 'update bulletin,' 'highlights digest,' and 'briefing summary' convey partial, non-exhaustive coverage to competent readers, so the gold label Unknown is correct independent of the model, yet the paper validates these labels only by manual inspection plus an LLM judge, with no human baseline or agreement rate reported.","fun_headline_variants_meta":{"raw":{"variants":["LLMs say 'no' when evidence is partial, not complete","Models can't tell 'absent' from 'not found' in partial lists","When absence is weak evidence, LLMs still answer definitively","LLMs overtrust lists that only show part of the story","Completeness-blind: LLMs treat missing data as confirmed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00029,"raw_usage":{"total_tokens":1766,"prompt_tokens":1082,"completion_tokens":684,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":698,"completion_tokens_details":{"reasoning_tokens":608}},"tokens_in":698,"tokens_out":684,"duration_ms":7218,"temperature":1.0,"reasoning_tokens":608,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:17:27.114865+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human-subject study on the L3 partial items would settle the central claim: present competent readers with the same question and listed titles under the four partial source-type phrases and ask whether the list could still fail to contain a relevant item. If a large share of readers treat 'highlights digest' or 'activity feed' as potentially exhaustive in context, the gold Unknown label is contested and the asymmetry would be partly a labeling artifact; conversely, if readers reliably read these phrases as partial, the model failure is confirmed.","supporting_citations":[],"review_version":1}