{"id":"dcb3f500-3cac-4d3e-ab91-6da6f5cd79c9","arxiv_id":"2507.16199","paper_version":6,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Adding an extra 'Unknown' option to True/False prompts causes LLMs to abstain on questions they can answer, and random words reproduce the effect, indicating abstention is partly a prompt artifact.","lead":"This paper shows that adding an 'Unknown' option to true/false questions makes large language models choose it on roughly a third of questions they otherwise answer correctly, even when the option is a random word. It argues that LLM abstention is often a prompt artifact rather than genuine uncertainty, which matters for benchmarks and systems that treat abstention as a confidence signal.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"S5 forced-rerun recovery is confounded by the coercive 'must select' instruction and unreported label balance, so the 52–75% 'can answer' evidence may show compliance, not latent capability.","rationale":"The reader identified the coercive S5 prompt as the weakest assumption, and I agree. This is the single most load-bearing concern because the paper's headline claim—'abstention can be a prompt artifact' in the sense of denying knowledge the model has—requires demonstrating that the model can in fact answer the abstained items. C1 (structural trigger) is solidly supported by S2–S4: the accuracy drop, the Abs Rate jump to 32.9%, and the random-word control are convincing and mutually reinforcing. S9 adds stability, and S10's directional discriminability on truly-Unknown samples is a good sanity check. The weak link is C2's S5: the probe is coercive, the subset's label balance is unreported, and the paper's own Appendix C.4 concedes that only ~28% of abstentions are cleanly 'known' under a simple decomposition. A neutral S1-style rerun would settle whether the recovery is capability or compliance. I do not think this warrants rejection, because the core phenomenon and the structural-trigger claim are well demonstrated, and the paper already frames the phenomenon as 'in addition to genuine uncertainty.' The verdict should remain CONDITIONAL, pending the neutral rerun and code/data release. My concern does not move the verdict; it sharpens the condition.","tokens_in":25693,"tokens_out":8764,"duration_ms":93440,"concrete_test":"Rerun the S5 abstention subsets with (a) the original S1 prompt verbatim (no 'Unknown' option, no coercive follow-up), and (b) a minimal neutral prompt: 'Answer True or False. Format your response exactly as: Reasoning: <reasoning> Final answer: <True or False>.' Also report the majority-class baseline and the True/False label counts for each model×dataset abstention subset. If accuracy on (a) and (b) does not exceed the majority-class baseline by a significant margin, or drops to ~50%, the S5 recovery is compliance/format-driven, not latent capability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that abstention is a prompt artifact—meaning the model denies knowledge it has—rests on C2, and C2 rests entirely on S5 (w/o 'Unknown' Option Rerun). The S5 prompt (Appendix I) is not a neutral probe of capability. It replays the previous S2 abstention and then instructs: 'pay more attention and avoid mistakes. This can be reasoned out based on objective factors. Subjective ability limits should be overcome. You must select one of the original labels: True or False.' These are explicit commands to reverse the abstention, and the 52–75% recovery could be instruction-following behavior rather than evidence that the model 'can' answer in any meaningful default sense. The paper never runs the plain S1 prompt (which also has no 'Unknown' option) on the abstention subset; it only runs this coercive multi-turn follow-up. Additionally, the label distribution of the abstention subset is not reported, so the 'above 50% random baseline' claim is unverified: if the subset is imbalanced toward True or False, always predicting the majority label would yield accuracy above 50% without any latent knowledge. The paper's own Appendix C.4 acknowledges that the forced-rerun accuracy implies only α ≈ 28% of abstentions are 'known' (α = 2·P(correct|forced) − 1), so most abstentions remain consistent with genuine uncertainty. Without a neutral, label-balanced capability probe, the title's stronger reading—abstention as denial of knowledge rather than just a format-induced output shift—is not secured. The random-word and format-ablation results (S3, S4) robustly establish the structural trigger, but the 'denies it can answer' component is the load-bearing bridge to the artifact interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the effect of adding an \"Unknown\" option to True/False and multiple-choice prompts. Across three commercial LLMs and six benchmarks, it reports that on TFQs the extra option produces large abstention rates and accuracy drops (average −15.75% accuracy, 32.9% abstention), that replacing \"Unknown\" with a random word preserves the effect, that rerunning abstained items without the option recovers 52–75% accuracy, that reasoning traces remain largely unchanged while final answers switch to \"Unknown\", and that the bias persists across temperature and emerges with instruction tuning. The authors organize the work around four claims (C1–C4) tracing the phenomenon from prompt structure to representation and training origin.","tokens_in":25977,"tokens_out":7277,"duration_ms":71965,"significance":"If the effect is real, the paper makes a valuable methodological point: abstention rates are not directly interpretable as epistemic uncertainty, and benchmarks that include \"Unknown\" labels need counterfactual controls. The design has notable strengths: the S2 effect is large and consistent across models; the random-word condition (S4) is an elegant control; S9 provides persistence evidence; S10 shows temperature invariance; the trace evaluation includes manual verification of 100 samples; and Appendix E.3 offers falsifiable predictions. The paper is not circular: the key quantities are measured against external datasets and baseline prompts, not fitted to the conclusion. However, the strongest reading (\"the model denies it can answer even when it can\") hinges on a capability probe whose prompt is coercive and whose label balance is unreported, so the central C2 claim needs repair.","major_comments":[{"comment":"The S5 rerun is not a neutral test of whether the model can answer the abstained items. The prompt replays the prior S2 response ending in \"Unknown\" and instructs the model to \"pay more attention\", \"overcome subjective ability limits\", and \"must select one of the original labels: True or False\". That is an explicit command to reverse the abstention, so the 52–75% recovery rate may measure instruction-following under pressure rather than latent capability. The obvious control is to run the plain S1 prompt (no \"Unknown\" option, no follow-up turn) on the abstention subset; the paper does not report this control. Since C2 and the abstract's \"denies it can answer even when it can\" rest on S5, this is load-bearing.","section":"§4.2.1 / Appendix I (S5)"},{"comment":"The \"above 50% random baseline\" interpretation of S5 is not verifiable without the label distribution of the abstention subset. If the subset is imbalanced toward True or False, always predicting the majority label gives accuracy above 50% without any latent knowledge; the paper never reports this balance. Moreover, the paper's own decomposition in Appendix C.4 (alpha = 2*P(correct|forced) − 1) yields a pooled known share of about 28%, meaning most abstentions remain consistent with genuine uncertainty. The text should report the label balance, provide per-label accuracy, and soften claims that S5 proves the model \"can\" answer.","section":"§4.2.1 / Appendix C.4"},{"comment":"C3 overstates what S8 shows. The logit-lens experiment tracks log P(\"Unknown\") across layers and demonstrates that the Unknown logit rises only in later layers, but it does not directly read out a True/False prediction from mid-layer hidden states; the authors concede this in Appendix D.1 (\"does not directly read out a True/False prediction from mid-layer hidden states\"). The \"mid-layer representations preserve the correct answer\" part of C3 is therefore inferred from the absence of a mid-layer Unknown rise plus the behavioral S5 result. A direct mid-layer probe of the gold label is needed before claiming representation-level preservation.","section":"§4.3.2 / Appendix D.1 (S8)"},{"comment":"The self-diagnosis prompt is leading. It asks the model to choose between \"subjective incapability\" and \"the question is objectively unanswerable - the given information is genuinely insufficient\", which offers a face-saving justification for the model's prior \"Unknown\" answer. Unsurprisingly, 95–100% of responses select option B. This does not establish that the model \"sincerely believes\" the abstention is warranted; it may simply be a post-hoc rationalization consistent with its own previous output. The introspective-gap claim would be stronger with a less leading prompt or an open-ended attribution.","section":"§4.2.2 / Appendix I (S6)"}],"minor_comments":[{"comment":"The \"FLD MCQ\" and \"FOLIO MCQ\" columns are S3 conversions, but the header \"Acc (S2)\" may confuse readers into thinking these are separate S2 runs; clarify in the caption that these are S3 runs with MCQ-style letter labels.","section":"Table 1 / §4.1.3"},{"comment":"The S3 prompt includes an additional instruction (\"Select 'C. Unknown' ONLY if the relationship is genuinely undeterminable... Do NOT select it simply because you feel uncertain\") that is absent from S2, so the S3-versus-S2 comparison is not perfectly controlled; the conclusion that question format is not the root cause should acknowledge this confound explicitly.","section":"§4.1.3 / Appendix I (S3)"},{"comment":"The abstract says replacing \"Unknown\" with a random word produces an \"identical effect\"; the data show statistically indistinguishable rates, not numerically identical rates. Use \"statistically indistinguishable\" or report confidence intervals for the difference.","section":"Abstract / §4.1.4"},{"comment":"The paper reports many proportions without confidence intervals (e.g., 32.9% Abs Rate, 52.4% 3/3 persistence, 42.3% NLI-recoverable). Adding bootstrap confidence intervals and releasing the code and data would materially strengthen the quantitative claims.","section":"General / §3.5"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid core result (extra-option abstention inflation) and the S4 control is clever. The main obstacle is that C2's capability claim rests on a coercive prompt and unreported label balance; if the authors cannot run a neutral S1-subset control, they should reframe the title and abstract to the weaker, well-supported claim that abstention can be format-induced rather than that the model demonstrably denies knowledge it has. I would also encourage the editor to require data/code release, since the supplementary quantitative appendix relies on many pooled tests that cannot be checked from the text alone."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline result is real: on True/False tasks, adding an 'Unknown' option cuts accuracy by roughly 16 points and pushes abstention to about 33%, and swapping the option label for a random word like 'Cerulean' leaves the behavior essentially unchanged. That random-word ablation, the TFQ/MCQ asymmetry, and the persistence across re-draws are the genuinely new pieces, and they are well demonstrated across three models and two datasets. The paper is worth engaging.\n\nWhere I part company with the authors' framing is C2. The S5 rerun that supposedly shows 'the model can answer when forced' uses a coercive prompt: it replays the abstention, tells the model to 'pay more attention,' overcome 'subjective ability limits,' and 'must select one of the original labels.' That measures instruction-following at least as much as latent capability. The paper's own Appendix C.4 does the algebra: with 52–75% forced accuracy, the estimated 'known' share is at most roughly 28% pooled, meaning most abstentions remain consistent with genuine uncertainty. The title's stronger reading—abstention as denial of knowledge, not just a format-induced output shift—is not secured by the evidence as presented. Also missing is the label balance of the abstention subset, which matters because a skewed subset could produce above-50% 'recovery' without any latent knowledge. That is a minor fix, but it should be reported.\n\nTwo smaller soft spots. The S3 format ablation includes an extra instruction telling the model not to choose 'Unknown' unless genuinely undetermined, so the MCQ control is not a clean structural comparison. And the C3 'mid-layer preservation' claim is inferred, which the authors admit; a direct probe would be the natural follow-up. No code or versioned model identifiers are provided, which is a real limitation for reproduction.\n\nNone of this sinks the central phenomenon. The structural trigger finding is robust, the random-word result is memorable, and the implications for abstention-based routers and benchmark construction are practical. The paper is for anyone building TFQ-format evals or using abstention as a calibration signal. It deserves a serious referee, but the authors should be pushed to reframe C2, report the subset balance, and release artifacts.","headline":"The random-word ablation is a real and memorable result, but the 'can answer when forced' claim rests on a coercive rerun prompt, and the paper's own algebra caps the known share near 28%.","tokens_in":26577,"tokens_out":2173,"would_cite":true,"duration_ms":22771,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding an 'Unknown' option makes LLMs falsely abstain even on questions they can answer, and renaming the option to a random word like 'Cerulean' changes nothing, showing abstention is partly a prompt artifact rather than genuine…","keywords":["LLM abstention","Abstention Inflation","prompt artifact","uncertainty calibration","instruction tuning","True/False questions","later-layer override","abstention bias"],"falsifier":"Re-run S5 with a neutral instruction that only removes the 'Unknown' option ('Please choose True or False') and measure accuracy on the formerly abstained items; if it falls to the 50% chance level, the recovery was compliance with the pressure wording and the C2 capability claim collapses. A second decisive check is a direct mid-layer probe reading the gold True/False label from hidden states of an abstaining model: if the label is absent before the final layers, the C3 override story fails.","tokens_in":25478,"feed_emoji":"🤷","tokens_out":10653,"duration_ms":100104,"temperature":0.7,"pith_summary":"This paper argues that LLM abstention is not only an expression of genuine uncertainty but also a prompt-induced artifact, a phenomenon it names Abstention Inflation. On True/False questions, adding an 'Unknown' option makes three frontier models abstain on 32.9% of items on average while accuracy drops by 15.75 percentage points; replacing the 'Unknown' label with a random word such as 'Cerulean' leaves the abstention rate unchanged. When the option is removed in a forced rerun, accuracy recovers to 52–75%, yet models attribute 95–100% of their abstentions to the item being objectively unknowable — an introspective gap in which the model denies capability it has. Representation probes locate the override in the later transformer layers, and factor analysis ties the bias to instruction tuning rather than stochastic noise. The stakes: any system or benchmark that consumes 'Unknown' outputs as calibrated uncertainty will inherit a format-dependent bias as though it were genuine doubt.","feed_headline":"Adding an 'Unknown' option makes LLMs fake ignorance","feed_subtitle":"Models abstain even on questions they can answer, and renaming the option 'Cerulean' changes nothing.","key_machinery":"The load-bearing device is the contrast between two versions of the same question: one with a designated extra option and one without. Ten settings build a ladder on this contrast: S1–S2 baseline vs. added 'Unknown'; S3 format conversion; S4 word-content ablation ('Unknown' → 'Indeterminate' → 'Cerulean'); S5 forced rerun without the option; S6 self-diagnosis; S7 reasoning-trace F1 plus a DeBERTa NLI probe; S8 a logit-lens read of $\\log P(\\text{``Unknown''})$ across all 33 layers of an open-weight model in base, instruction-tuned, and RL variants; S9 persistence across three redraws; S10 temperature, difficulty, size, and alignment sweeps. The key comparison is the TFQ-vs-MCQ double dissociation: the same extra option moves abstention by tens of points on binary logic questions and by only small margins on four-option MCQs, which is what separates a structural trigger from a semantic or difficulty effect.","core_discovery":"The central claim is that abstention behavior in LLMs is inflated by the structural presence of an extra response option, regardless of the option's meaning. The paper establishes four progressive propositions: (C1) the trigger is structural — 'Unknown' and 'Cerulean' behave identically, and almost nothing changes when the format is switched from True/False to letter-coded labels; (C2) the effect makes models deny knowledge they demonstrably have, since removing the option recovers 52–75% accuracy on the very items the model abstains on, while self-diagnosis denies any subjective incapability; (C3) the override happens at the output end of the network, with reasoning-trace quality essentially unchanged and the 'Unknown' logit rising only in the last layers; and (C4) the bias is stable across repeated draws and temperatures and is installed by instruction tuning, as base models show far lower abstention inflation than their instruction-tuned counterparts. The net position is that abstention is often a learned surface pattern, not a faithful uncertainty signal, and that benchmarks and routers should not take a single-format 'Unknown' label at face value.","pith_inferences":["This suggests the effect may be a general 'escape-hatch option' phenomenon: any extra, low-commitment option could act as an abstention slot, so a natural next test is whether 'None of the above' or a confidence scale triggers the same inflation on MCQ-style tasks — a comparison the paper does not run.","The C3 claim is indirect: mid-layer preservation is inferred from the absence of a mid-layer Unknown-logit rise, not from reading the gold label out of those layers. A direct hidden-state probe for True/False would distinguish 'override at the output' from 'the answer never formed'.","The abstention tax has a deployment corollary the authors leave implicit: if each point of instruction-tuning accuracy is bought with roughly a point of false abstention, then net delivered accuracy — answers users actually receive — may be roughly flat on TFQ-with-Unknown setups, making such channels costlier than headline Acc suggests.","Because the S5 rerun prompt tells the model to 'pay more attention' and that it 'must select one of the original labels', part of the 52–75% recovery could be compliance. Re-running S5 with a neutral instruction is the cleanest way to separate capability from instruction-following, which would directly test the paper's C2."],"forward_implications":["Downstream systems that route on abstention — abstention-based routers, confidence routers, multi-agent pipelines — will inherit the extra-option bias as if it were genuine epistemic uncertainty.","Benchmark designers who include an 'Unknown' category should report Abs Rate beside accuracy and include a w/o-option rerun as a routine sanity check; otherwise the abstention number mixes inflated and genuine refusals.","Reformatting binary questions away from a True/False-with-extra-option setup removes most of the inflation at no capability cost, according to the format-ablation results.","Instruction tuning raises accuracy and false abstention together — the paper quantifies an 'abstention tax' near one point of extra abstention per point of accuracy gained in the vulnerable format.","Models still distinguish answerable from truly-unknown items by a wide margin, so the bias is a directional over-trigger rather than a collapse of the model's ability to tell the two populations apart."],"supporting_citations":[{"why":"Supplies the AbstentionBench benchmark that reports 24% degradation of abstention under fine-tuning; the paper reads the same behavior on answerable items as prompt-inflated abstention.","marker":"Kirichenko et al., 2025"},{"why":"Provides the FLD synthetic logic corpus on which the main True/False accuracy drops and abstention rates are measured.","marker":"Morishita et al., 2024"},{"why":"Provides the FOLIO human-annotated first-order-logic entailment benchmark, the second TFQ dataset showing the effect.","marker":"Han et al., 2024"},{"why":"Provides the ARC MCQ benchmark whose near-ceiling accuracy serves as the control showing the extra option barely moves letter-coded tasks.","marker":"Clark et al., 2018"},{"why":"Establishes the prior view that LLMs can partially predict their own accuracy, the baseline that C2's capability–self-report dissociation challenges.","marker":"Kadavath et al., 2022"},{"why":"Documents the Neutral bias in NLI probes, which the paper uses to calibrate its reasoning-trace evaluation in S7.","marker":"Pavlick and Kwiatkowski, 2019"},{"why":"Defines the RLHF alignment paradigm whose preference data, the paper argues, conflates genuine and inflated abstention, grounding the C4 instruction-tuning attribution.","marker":"Ouyang et al., 2022"}],"fun_headline_variants":["LLMs fake uncertainty when given an extra answer option","Unknown or 'Cerulean', LLMs abstain either way","Abstention inflation: LLMs skip questions they can answer","LLM ignorance is a prompt artifact, not true uncertainty","Adding 'Unknown' makes LLMs deny what they know"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The S5 forced-rerun prompt urges the model to 'pay more attention', to overcome 'subjective ability limits', and states it 'must select one of the original labels', so the 52–75% recovery could be instruction-following rather than latent capability; the companion claim that mid-layer representations preserve the correct answer is inferred indirectly from a logit-lens contrast, not directly probed.","fun_headline_variants_meta":{"raw":{"variants":["LLMs fake uncertainty when given an extra answer option","Unknown or 'Cerulean', LLMs abstain either way","Abstention inflation: LLMs skip questions they can answer","LLM ignorance is a prompt artifact, not true uncertainty","Adding 'Unknown' makes LLMs deny what they know"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000246,"raw_usage":{"total_tokens":1580,"prompt_tokens":1027,"completion_tokens":553,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":470}},"tokens_in":643,"tokens_out":553,"duration_ms":6062,"temperature":1.0,"reasoning_tokens":470,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:15:51.482015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run S5 with a neutral instruction that only removes the 'Unknown' option ('Please choose True or False') and measure accuracy on the formerly abstained items; if it falls to the 50% chance level, the recovery was compliance with the pressure wording and the C2 capability claim collapses. A second decisive check is a direct mid-layer probe reading the gold True/False label from hidden states of an abstaining model: if the label is absent before the final layers, the C3 override story fails.","supporting_citations":[],"review_version":1}