{"id":"298511e1-e5d2-49db-b7b9-d8ebe6bb7d2c","arxiv_id":"2607.08456","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Answer correctness and question answerability are separate axes: ordinary confidence tracks the first while hidden probes track the second, and a factorized dual-threshold policy certifies both risk budgets at higher correct-answer coverage.","lead":"LLMs must refuse both wrong answers and unanswerable questions, but a single confidence score cannot separate them. Across five models, answer-confidence tracks correctness while hidden-state probes track answerability, enabling dual-risk certified abstention that beats single-threshold policies.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged small-n certification and moderate CREPE AUROC.","rationale":"The reader's weakest_assumption correctly isolates the practical limits of the answerability probe (moderate natural-data AUROC, surface recoverability on SelfAware, tiny W counts). Those limits justify CONDITIONAL rather than unconditional ACCEPT, but they are already measured and do not undermine the crossed geometry or the qualitative superiority of the two-threshold policy under the stated budgets. Because the paper supplies the label-artifact, transfer, elicitation, and replication-audit checks that would have been the natural places for a deeper flaw, no additional load-bearing concern is required. The concrete test simply scales the certification sample—the one place where the dual-risk guarantee is statistically fragile—so that the same claim can be re-checked with tighter wrong-answer budgets or lower variance. Verdict and confidence therefore stay as the reader set them.","tokens_in":18807,"tokens_out":513,"duration_ms":4906,"concrete_test":"Re-run the factorized vs. answer-confidence-only certification pipeline of Table 4 / Appendix B on a 3–5\times larger SelfAware-style sample (or pooled multi-seed draws) so that each certification split contains ≥60 W items; if the factorized policy still certifies both budgets at ≥2\times the single-threshold C-coverage at 8B and remains the only certifying policy at 14B, the dual-risk claim is robust to the current small-n limit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the same-decision C/W/U L-geometry plus the dual-risk factorized policy that certifies higher C-coverage than single-axis baselines. The paper already runs the controls that would most threaten that claim: label-artifact training without W items (Table 8), surface bag-of-words and elicitation bounds (Sections 4–5), SelfAware\to CREPE transfer, and a 100-resplit certificate audit (Appendix D). The remaining soft spots—18–22 W items per certification split, CREPE AUROC only 0.69–0.78, SelfAware surface confound—are explicitly quantified and do not reverse the reported asymmetry or the scale-dependent certification pattern. No hidden inconsistency or untested assumption appears load-bearing enough to overturn the geometry or the dual-budget result.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that LLM abstention conflates two distinct failure modes—emitting a wrong answer to an answerable question (W) versus answering an unanswerable or false-premise question (U)—and that these map to separate score axes on the same items. Across five instruction-tuned models (2B–14B, three families), ordinary answer-confidence separates correct-answerable (C) from W/U but barely separates W from U, while a linear hidden-state answerability probe does the reverse (an L-shaped C/W/U geometry). The asymmetry is worst on naturally occurring false presuppositions (CREPE), where output confidence, P(IK), P(True), and direct premise-check elicitation stay near chance while internal readouts reach 0.69–0.78 AUROC. Instructed premise-checking contests sound and false premises indiscriminately; routing the same instruction with the probe roughly triples challenge precision. The authors formalize three-class selective acceptance with separate risk budgets (α_U=0.15, α_W=0.50) and show a factorized two-threshold policy certifies higher C-coverage than single-axis baselines under leak-free split-conformal certification, with a scale-dependent asymmetry: the unanswerable budget is controllable, while the wrong-answer budget is limited by model accuracy.","tokens_in":19096,"tokens_out":1784,"duration_ms":41287,"significance":"If the same-decision L-geometry and dual-risk certificates hold, the paper cleanly separates two requirements that the dominant single-threshold abstention recipe cannot represent, and turns that separation into an operational policy with per-axis risk attribution. Strengths that should be credited: equal-capacity trained readouts for all four source×axis cells (Table 1); a label-artifact control training the answerability direction on C-vs-U only that still places held-out W with C (Table 8); honest surface and elicitation bounds (bag-of-words 0.87 on SelfAware; P(IK)/premise-check on CREPE); SelfAware→CREPE transfer; leak-free four-way splits with Clopper–Pearson dual-risk certificates; and a 100-resplit end-to-end certificate audit (Appendix D). The premise-routing pilot is a concrete behavioral payoff beyond pure detection. The contribution is not that answerability is internally legible (already shown by Slobodkin et al. and Lavi et al.), but the controlled crossed geometry, its measurement under those bounds, and the certified two-budget composition.","major_comments":[{"comment":"Abstract and §7 claim that at 14B the factorized policy is “the only policy that certifies at all.” Tables 4 and 9 contradict this: on Qwen2.5-14B, answerability-only also certifies (test C-coverage 0.35, UCB(R_U)=0.07, UCB(R_W)=0.45 ≤ 0.50), while answer-confidence-only and the learned joint scalar fail. Factorized is better (0.47) but not unique. The accompanying gloss that “both confidence-thresholding alternatives fail a budget” is true of those two baselines but does not justify the stronger uniqueness claim. Please correct the abstract, §7, and any parallel conclusion language so the 14B comparison matches the tables.","section":"Abstract; §7; Tables 4 and 9"},{"comment":"Certification and test splits are very small for the dual-risk claim: n_C on the test split is 8–17 (Table 4), and certification splits contain only 18–22 W items (Table 9). Under those counts α_W=0.50 is a lenient budget; tighter W budgets never certify (Table 10), partly by sample limit rather than signal quality, as the paper notes. The 100-resplit audit (Appendix D) is the right mitigation and shows issued certificates are rarely violated, but issue rates themselves are low on smaller models (18/100 on Gemma 2B). The scale-dependent story (W budget becomes certifiable as accuracy rises) is plausible but currently rests on thin binomial counts. Either enlarge the SelfAware sample so that W and C denominators support tighter budgets, or move the small-n fragility and the α_W=0.50 choice into the main text of §7 rather than leaving them mostly in the appendix and Limitations.","section":"§7; Tables 4, 9, 10; Appendix B/D"},{"comment":"On CREPE the internal signal is only moderate (hidden readout 0.69–0.73; difference-of-means 0.74–0.78; Table 2), and the certified false-premise gate yields low admissible coverage (0.09–0.25 on models that certify; Table 11). The dual-risk factorized policy of §7 is evaluated only on SelfAware, where answerability is largely surface-recoverable (bag-of-words AUROC 0.87). The paper is careful about this, but the abstract’s policy claim is written as if the two-axis certificate transfers cleanly to the natural-data regime where the output blind spot is sharpest. Either run the factorized (or at least dual-signal) policy on a setting that has both axes without the SelfAware surface confound, or explicitly scope the dual-budget certificates to SelfAware and keep CREPE as an admissibility-only result.","section":"Abstract; §5; §7; Tables 2 and 11"}],"minor_comments":[{"comment":"Table 4’s checkmark formatting is hard to parse in the preprint layout (checkmarks run into adjacent cells). Align with the clearer Table 9 layout or use an explicit “cert. yes/no” column.","section":"Table 4"},{"comment":"SelfAware answerable accuracies are low (0.28–0.42). The paper flags caution on the correctness axis; a one-sentence reminder next to Figure 1 and Table 1 would help readers who skip the setup paragraph.","section":"§3; Figure 1; Table 1"},{"comment":"The neuro-symbolic framing (open/closed-world boundary, Logic Tensor Networks) is interesting but optional; a shorter pointer in the introduction and a fuller paragraph in Related Work or Discussion would reduce the risk that readers treat it as a load-bearing formal result.","section":"§1; §2; §7"},{"comment":"Premise-routing pilot (Table 3) reports strict template detection and NLI validation bounds; releasing the manual audit sample is good. State the NLI model and decision rule in the main text or Appendix E so the 47%/61% “validated correct contest” numbers are reproducible from the paper alone.","section":"§6; Appendix E"},{"comment":"Minor wording: abstract sentence “each clear a separately certifies behave differently” appears truncated/garbled in the provided text; repair before camera-ready.","section":"Abstract"},{"comment":"Quantization validation is only on Qwen2.5-3B (Appendix H). A one-line note that 8-bit/4-bit results for 7B–14B are therefore pipeline-validated by transfer, not re-validated per model, would be clearer.","section":"§3; Appendix H"}],"recommendation":"minor_revision","confidential_remarks":"The factual overclaim that factorized is the only certifying policy at 14B is the main fix; once corrected, the paper is a careful, well-controlled contribution rather than an oversold one. Novelty relative to Slobodkin et al. (2023), Lavi et al. (2026), and concurrent query-level uncertainty work is disclosed honestly; I would not treat that as a rejection ground. Fit for a serious CL/ML journal is good if the dual-risk certificates and CREPE scope are stated precisely. No concerns about ethics or data use."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful bit is the same-item C/W/U geometry: answer-confidence separates correct from wrong/unanswerable but barely separates the latter two, while a hidden answerability probe does the reverse and collapses C with W. That L-shape holds from 2B to 14B, survives a label-artifact control (train C-vs-U only, held-out W still sits with C), and is sharpest on CREPE where every output-side signal they tried (raw confidence, trained readout, P(IK), P(True), direct premise check) sits near chance while the probe reaches 0.69–0.78. The dual-threshold policy then certifies separate risk budgets and, at 8B, roughly doubles certified correct coverage versus confidence-only; at 14B it is the only policy that certifies at all. Probe-routed premise checking is a clean behavioral payoff: the instruction alone contests indiscriminately, the probe triples precision.\n\nWhat is new is not that answerability is internally legible—Slobodkin, Lavi, Kadavath, and concurrent query-level work already established pieces of that—but the controlled same-decision geometry, the CREPE natural-data stress test where surface confounds are weak by construction, and the factorized dual-risk certificates with 100-resplit audits. The paper positions itself accurately against those priors and runs the surface, elicitation, and transfer controls the claim needs. Math is standard split-conformal / Learn-Then-Test; no circularity. Citation pattern is clean.\n\nSoft spots are real but already quantified: SelfAware answerability is largely bag-of-words recoverable (0.87), CREPE AUROC is only moderate, certification splits have 18–22 W items so tighter wrong-answer budgets cannot certify, and code is still “to be released.” None of those reverse the asymmetry or the scale-dependent certification pattern. Small models only; no causal intervention.\n\nThis is for people building selective-prediction or hallucination gates who need per-axis risk statements rather than a single fused score. It deserves a serious referee. I would engage with it and expect to cite the geometry and the dual-budget framing.","headline":"Solid empirical decomposition of abstention into two axes, with honest controls and dual-risk certificates that actually move the needle over single-threshold baselines.","tokens_in":19723,"tokens_out":535,"would_cite":true,"duration_ms":5646,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Answer correctness and question answerability are separate axes of LLM abstention that a single confidence score cannot tell apart.","keywords":["LLM abstention","answerability","answer confidence","hidden-state probes","selective prediction","false presuppositions","dual-risk certification","CREPE"],"falsifier":"On a large held-out sample of natural false-premise questions, if trained answer-confidence or direct premise-check elicitation matched or beat the hidden-state probe, and a two-threshold factorized policy failed to improve certified dual-risk coverage over a single calibrated confidence threshold, the claimed axis separation would not hold.","tokens_in":19701,"feed_emoji":"🛑","tokens_out":930,"duration_ms":19128,"temperature":0.7,"pith_summary":"Selective answering is usually done by thresholding one confidence score. This paper shows that recipe confuses two different failures: emitting a wrong answer to an answerable question, and answering a question that should not be answered at all. Across five instruction-tuned models from 2B to 14B, ordinary answer-confidence tracks whether an attempted answer succeeded but is nearly blind to whether the question was admissible, while a linear probe on hidden states does the reverse. The blind spot does not shrink with scale and is worst on naturally occurring false-premise questions, where confidence, self-assessment, and even direct premise checks stay near chance while the internal probe still works. The authors turn the separation into a factorized policy that answers only when both scores clear independent thresholds, certifying separate risk budgets on unanswerable answers and wrong answers, often at higher coverage of correct answers than any single-score rule.","feed_headline":"Two scores, not one, for safe LLM abstention","feed_subtitle":"Ordinary confidence tracks wrong answers; a hidden probe tracks unanswerable ones—and using both certifies safer refusals.","key_machinery":"Factorized Abstention: answer only when both an answerability score and a correctness score clear independent thresholds, with dual-risk certification of separate budgets on the unanswerable-answer rate and the wrong-answer rate.","core_discovery":"Correct-answerable, wrong-answerable, and unanswerable questions form a crossed L-shaped geometry on the same decisions: ordinary answer-confidence separates successful from unsuccessful execution but barely separates wrong answers from unanswerable questions, while a hidden-state answerability readout separates admissibility from inadmissibility but collapses correct and wrong answerable items. The pattern holds from 2B to 14B and is sharpest on natural false-presupposition data.","pith_inferences":["Production systems that report only one confidence number will be systematically miscalibrated for open-world user questions even when they look well-calibrated on closed-world benchmarks.","The moderate CREPE signal suggests that stronger answerability representations, if they can be trained or steered, would unlock much higher certified coverage without losing per-axis risk attribution.","The same factorization could apply to other selective behaviors such as tool use, retrieval triggering, or multi-step planning, wherever “should I attempt” and “will the attempt succeed” diverge."],"forward_implications":["Single-threshold confidence policies systematically over-abstain or fail to control unanswerable-answer rates because they cannot represent the open region of inadmissible questions.","Certified control of unanswerable answers is possible at every scale tested, while wrong-answer control is limited by model accuracy and only becomes useful as models get more accurate.","Prompting a model to check premises without an external discriminator causes it to challenge sound and false premises alike.","Routing the same premise-check instruction with a hidden-state probe roughly triples challenge precision.","At 14B in the tested range, only the factorized two-axis policy certifies both risk budgets at all."],"fun_headline_variants":["Two axes of LLM abstention: correctness vs answerability","Confidence tracks wrong answers; probes track unanswerables","Hidden-state probes spot unanswerable questions confidence misses","False-premise questions expose a scale-resistant LLM blind spot","Dual scores certify safer refusals than single confidence threshold"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The claim rests on a linear readout of final prompt-token hidden states giving a general enough answerability signal for risk-controlled gating and premise routing, even though that signal is only moderate on natural data and SelfAware answerability is largely available from surface features alone.","fun_headline_variants_meta":{"raw":{"variants":["Two axes of LLM abstention: correctness vs answerability","Confidence tracks wrong answers; probes track unanswerables","Hidden-state probes spot unanswerable questions confidence misses","False-premise questions expose a scale-resistant LLM blind spot","Dual scores certify safer refusals than single confidence threshold"]},"model":"grok-4.5","effort":"low","cost_usd":0.003482,"raw_usage":{"total_tokens":1237,"prompt_tokens":889,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":34820000,"prompt_tokens_details":{"text_tokens":889,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":266,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":889,"tokens_out":82,"duration_ms":12641,"temperature":1.0,"reasoning_tokens":266,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T07:10:42.927479+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a large held-out sample of natural false-premise questions, if trained answer-confidence or direct premise-check elicitation matched or beat the hidden-state probe, and a two-threshold factorized policy failed to improve certified dual-risk coverage over a single calibrated confidence threshold, the claimed axis separation would not hold.","supporting_citations":[],"review_version":1}