{"id":"f283a7a6-d1fd-469b-af6e-2ab59867c6a3","arxiv_id":"2607.03870","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"SALT enables judge-free unit-level uncertainty evaluation on deterministic long-form tasks and shows atomic ranking collapse, path-dependent error drivers, and a reasoning–ranking trade-off across 50+ LLMs.","lead":"SALT is a zero-noise benchmark of six procedurally generated long-form tasks with unique deterministic answers, so unit-level correctness needs no LLM judge. Across 50+ models it shows atomic confidence ranking largely fails, prefix correctness drives later errors more than length, and reasoning boosts accuracy while hurting ranking.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified that overturns the internal claims; the main soft spot remains external representativeness already flagged by the reader.","rationale":"The reader correctly identifies the central empirical package and the weakest assumption (representativeness of synthetic single-answer, strict-index tasks). That assumption qualifies how far the findings should be extrapolated to free-form multi-valid generation, but it does not falsify the zero-noise measurements, the granularity gap, the controlled DNA interventions, or the paired reasoning comparisons on SALT. The paper already discloses incomplete coverage of ARC-AGI/Maze, the non-alignment choice, and conditional external transfer. No stronger load-bearing flaw (e.g., circular use of judges, broken calibration math, or unreported selection of confidence functions) is present. Therefore the ACCEPT verdict with medium correctness risk for real-world generalization stands; no adjustment is warranted.","tokens_in":52597,"tokens_out":547,"duration_ms":5451,"concrete_test":"Re-run the atomic vs. line Macro-AUROC comparison (Figure 5) and the reasoning trade-off (Figure 6) after applying Needleman–Wunsch unit alignment (Appendix H) on all six main tasks, and separately recompute the prefix-intervention curves of Figure 4 on Matrix Multiplication or First-Order Logic (higher dependence) with the same cubic B-spline protocol. If atomic AUROC remains near chance and the reasoning–AUROC drop and global-prefix dominance persist under both changes, the internal claims hold; large reversals would show the results are artifacts of strict indexing or DNA’s independence structure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper’s strongest internal claims (atomic ranking near chance while line-level is informative; separable prefix-correctness vs. bounded length drivers; reasoning improves precision while degrading AUROC) rest on deterministic, index-aligned unit labels over six single-answer structured tasks, with the main causal intervention on DNA (low inter-atom semantic dependence). That design is methodologically clean and the statistics (Wilcoxon, paired interventions, mediation) support the within-SALT conclusions. The load-bearing external condition is that these dynamics transfer to open-ended or multi-valid long-form generation; the paper itself reports only conditional, calibration-regime-dependent agreement with AIME/MMLU-Pro and excludes ARC-AGI/Maze from the main aggregate. That is a genuine scope limit on impact, not an internal inconsistency or circularity. No hidden math error, label-noise circularity, or untested aggregation choice undermines the reported SALT results.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces SALT, a procedurally generated long-form benchmark of six single-answer tasks (math, code, logic, DNA translation, multi-needle) with deterministic unit-level ground truth, enabling zero-noise evaluation of precision, ECE, and AUROC at atomic and line resolutions without external judges. Across 50+ LLMs it reports that (i) logits-based confidence functions dominate verbalized ones, with different functions best for calibration vs ranking; (ii) confidence ranking largely fails at atomic resolution while remaining informative at line level; (iii) controlled prefix interventions (mainly DNA) separate two drivers of future error—propagation from corrupted prefixes, with global correctness dominating local, and a bounded length-related degradation; and (iv) reasoning (CoT or trained reasoners) improves precision while systematically degrading AUROC. Code is released.","tokens_in":52901,"tokens_out":982,"duration_ms":8634,"significance":"If the within-SALT results hold, the paper supplies a clean, contamination-resistant substrate for fine-grained uncertainty evaluation that prior long-form benchmarks lack because of judge noise. The atomic ranking failure, separable prefix vs length drivers, and reasoning–ranking trade-off are concrete, falsifiable findings with direct implications for selective generation and risk-critical deployment. Strengths include the large multi-model sweep, Wilcoxon dominance tests, mediation/ANM-style analysis of precision→ECE, and controlled interventions with mutual adjustment and saturation diagnostics (Appendix B). External representativeness remains the main limit on impact, which the paper itself flags via conditional AIME/MMLU-Pro transfer.","major_comments":[{"comment":"Section 5.1 and Appendix B.2: the causal claim that future correctness has two separable drivers (global prefix correctness dominating local, plus bounded length degradation) rests primarily on DNA interventions chosen for low inter-atom semantic dependence. Appendix K formalizes that other tasks (logic, multi-needle, matrix mult) have stronger input-span or logical-output dependence. Without analogous interventions on at least one higher-dependence task, the generality of the two-driver claim beyond DNA is under-supported for the main-text framing.","section":null},{"comment":"Section 5.2 (reasoning trade-off) and Appendix G.5–G.6: the AUROC degradation under CoT/reasoning is a central claim, but the paired Instruct vs Reasoning comparisons and hybrid CoT+reasoning analysis are reported mainly in aggregate. Task-level and model-pair breakdowns (Figures 39–40) show substantial heterogeneity; the manuscript should state more clearly for which tasks/models the ranking degradation is robust versus precision-driven or task-specific, so the trade-off is not over-generalized from the median effect.","section":null}],"minor_comments":[{"comment":"Section 4.2 vs abstract/intro: the text mentions eight tasks then six with full coverage (ARC-AGI and Maze held out). Align the abstract and contribution list with the six-task main aggregate to avoid confusion.","section":null},{"comment":"Figure 1 and Section 2.3: ECE and AUROC are illustrated at generation/line/atom levels; a short note that generation-level AUROC is often undefined or degenerate when entire generations are all-correct or all-incorrect would help readers interpret the granularity gap.","section":null},{"comment":"Appendix H: the Needleman–Wunsch alignment ablation is useful; a one-sentence pointer in Section 4.3 to why strict indexing is preferred for precision (redundant units) would strengthen the main-text justification.","section":null},{"comment":"Table 1 / task sizes: DNA contributes ~29k of ~55k atoms. Confirm that task-equal averaging (stated in Appendix A) is used for all main figures so DNA does not dominate aggregate AUROC/ECE.","section":null},{"comment":"Typos and polish: e.g., 'Words Collection' prompt figure caption reused for Kronecker in one place; 'Maro-PRR' in a figure caption; minor notation consistency for U_gen vs U_gt.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The internal methodology is solid and the paper is a genuine contribution to uncertainty evaluation infrastructure. The main risk is over-claiming transfer to open-ended long-form generation; the authors are already cautious in places. Minor revision to scope the intervention and reasoning claims more carefully should be sufficient. Fit for a strong ML venue is good if the external-validity language is tightened."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a careful empirical systems paper that actually ships something usable. SALT gives six procedurally generated long-form tasks with a priori unique ground truth and natural unit decomposition, so you can score correctness, ECE, and AUROC at atom and line resolution without LLM judges. That alone is new relative to FActScore-style pipelines and short-form calibration suites. They evaluate 50+ models, release code, and back the main claims with Wilcoxon tests, paired prefix interventions (splines + mutual adjustment), and mediation for the reasoning effect.\n\nWhat holds up inside SALT: (1) ranking largely collapses at atomic resolution (many models near chance on micro/macro AUROC) while line-level remains informative; (2) future correctness has two separable drivers—global prefix correctness dominates local, and length adds a bounded early-saturating degradation; (3) CoT or trained reasoners raise precision but systematically hurt AUROC. Logits-based confidences beat verbalized ones; different aggregators win for ECE vs ranking; binned calibration helps ECE and hurts ranking. The math and stats look standard and honestly applied. Citations cover the right prior work without heavy self-citation games.\n\nSoft spots are real but scoped. The causal intervention is mainly on DNA (low inter-atom dependence). Main aggregates exclude ARC-AGI/Maze. Transfer to AIME/MMLU-Pro is conditional and calibration-regime dependent—the paper says so. Strict index-aligned string match after fencing is clean for these tasks but is not open-ended multi-valid generation. That qualifies impact; it does not break the internal results.\n\nWho should read it: anyone working on selective generation, long-form reliability, or confidence functions who is tired of noisy atomic-fact judges. I would bring it to reading group, cite the benchmark and the three regularities, and send it to peer review. It deserves referee time.","headline":"Clean zero-noise long-form uncertainty benchmark plus three solid empirical regularities; main limit is how far synthetic single-answer tasks travel.","tokens_in":53491,"tokens_out":466,"would_cite":true,"duration_ms":6833,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"With zero-noise unit labels, LLM confidence ranking largely fails at atomic resolution, and reasoning boosts accuracy while harming the ability to rank errors.","keywords":["LLM uncertainty estimation","long-form generation","deterministic ground truth","atomic evaluation","confidence ranking","calibration","reasoning trade-off","prefix error propagation"],"falsifier":"Re-run the same models and confidence functions on open-ended long-form tasks with high-quality multi-annotator atomic labels (or another zero-noise multi-valid setting) and check whether atomic AUROC remains near chance, the global-prefix propagation effect replicates, and the reasoning-induced AUROC drop still appears; if atomic ranking becomes strong or the trade-off vanishes, the central claims do not transfer.","tokens_in":53481,"feed_emoji":"🔬","tokens_out":693,"duration_ms":6179,"temperature":0.7,"pith_summary":"Long-form LLM outputs contain many local pieces, some right and some wrong, so uncertainty tools must flag the bad pieces rather than reject the whole response. Existing long-form benchmarks introduce label noise through judges or multi-valid answers, which can systematically distort ranking and calibration metrics. This paper introduces SALT, a suite of six procedurally generated tasks (math, code, logic, translation, multi-needle retrieval) that have a single deterministic long textual ground truth, so every atomic unit can be scored by exact string match without any external judge. Across 50+ models the authors show three concrete results: raw confidence signals often fail to separate correct from incorrect atoms even when coarser line-level ranking still works; future correctness is driven by two separable factors—error propagation from corrupted prefixes (global prefix quality dominates local) plus a bounded length-related degradation that saturates early; and both CoT prompting and trained reasoners improve precision while systematically degrading confidence ranking. The practical claim is that risk-critical systems need high-resolution, zero-noise evaluation and must treat reasoning’s accuracy gain as potentially paid for by worse self-ranking of errors.","feed_headline":"LLM confidence ranking fails at atom scale","feed_subtitle":"Zero-noise long-form labels show reasoning boosts accuracy while harming error ranking","key_machinery":"SALT (Single-answer Atomic Long-form Target): six procedurally generated tasks with one known long textual ground truth, enabling exact unit-level correctness labels, multi-granularity scoring, and controlled atom-level prefix interventions without judges or noisy decomposition.","core_discovery":"On a deterministic long-form benchmark with exact unit labels, confidence ranking largely collapses at atomic resolution for most models even when line-level ranking remains informative; controlled prefix interventions separate two drivers of later errors—propagation dominated by global context correctness and a bounded length effect—and reasoning (CoT or specialized training) improves accuracy while degrading ranking ability.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLM confidence ranking collapses at atomic resolution","Atomic units break LLM ranking despite line-level success","Deterministic labels expose atom-scale confidence failure","Reasoning lifts accuracy but degrades LLM error ranking","Prefix errors and length bound drive long-form uncertainty"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"That the error dynamics, atomic ranking failure, and reasoning–ranking trade-off measured on these six single-answer structured tasks with strict index-aligned string matching transfer to open-ended long-form settings where multiple answers can be valid.","fun_headline_variants_meta":{"raw":{"variants":["LLM confidence ranking collapses at atomic resolution","Atomic units break LLM ranking despite line-level success","Deterministic labels expose atom-scale confidence failure","Reasoning lifts accuracy but degrades LLM error ranking","Prefix errors and length bound drive long-form uncertainty"]},"model":"grok-4.5","effort":"low","cost_usd":0.006114,"raw_usage":{"total_tokens":1599,"prompt_tokens":774,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":61140000,"prompt_tokens_details":{"text_tokens":774,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":770,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":774,"tokens_out":55,"duration_ms":5804,"temperature":1.0,"reasoning_tokens":770,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T23:22:07.542388+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same models and confidence functions on open-ended long-form tasks with high-quality multi-annotator atomic labels (or another zero-noise multi-valid setting) and check whether atomic AUROC remains near chance, the global-prefix propagation effect replicates, and the reasoning-induced AUROC drop still appears; if atomic ranking becomes strong or the trade-off vanishes, the central claims do not transfer.","supporting_citations":[],"review_version":1}