{"id":"31d2afe7-21dd-4702-818c-69bace12799a","arxiv_id":"2606.05170","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"At matched accuracy, 85 of 210 open-weight LLM pairs have disjoint severity-tail slopes, so error rate alone cannot rank catastrophic-failure risk.","lead":"Open-weight LLMs with nearly identical error rates can still differ sharply in how severe their rare failures are. The paper argues that a Gutenberg–Richter-style tail index should be reported alongside accuracy for safer model selection.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Human-consensus b CIs for the 85-pair headline rest on ~35 items/model, so the matched-accuracy discriminator may be underpowered relative to the judge baseline.","rationale":"The reader correctly flags score reliability, overcall, and mmin as soft spots and lands on CONDITIONAL. Those issues matter, but the single most load-bearing threat to the strongest claim is narrower: the 85-pair human-consensus discriminator is presented as the primary evidence that severity is invisible to ε, yet the human sample that underwrites those CIs is only ~35 items per model. Rank validation (ρ=0.89) and ICC do not automatically transfer to tight per-model b intervals or to a 2.7× jump in disjoint pairs. The paper’s own Resolution Bound and the judge-side exceedance sweep (Appendix P) show that tail CIs need substantial support; the human n is far below that. A direct recompute of the 85 count on the 519-item set would settle whether the headline is measurement-supported or an artifact of thin human tails / unclear consensus expansion. This does not overturn the useful evaluation axis or the non-reducibility intuition, so the verdict stays CONDITIONAL rather than REJECT, with the condition now focused on human-b precision.","tokens_in":26841,"tokens_out":727,"duration_ms":6515,"concrete_test":"Recompute the matched-accuracy disjoint-CI count using only the 519-item human scores with the same Aki/bootstrap pipeline and mmin rule as the main text; report how many of the 85 pairs (or of the 15-model subset) still have non-overlapping 95% CIs, and the median human SE(b). If the count falls near or below the judge baseline of ~31 (or most CIs overlap), the human-consensus 85-pair headline is not supported at the claimed precision.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that 85/210 pairs have disjoint 95% b CIs at |Δε|<0.05 on human-consensus scoring (e.g. DeepSeek-v3.2 vs Ministral-14B, Δb=0.47). That count is the paper’s headline and is 2.7× the 31-pair LLM-judge figure. Yet §8 and the human-validation section state the 519-item study covers only ~35 items per model across 15 models. Per-model b and its bootstrap CI are estimated from those few positive-severity scores after mmin selection (often requiring ≥30 exceedances in the judge pipeline). With n≈35 total items, the number of upper-tail exceedances is typically far smaller, so the human-derived CIs are wide and the 85-pair count is sensitive to sampling noise and to how “human-consensus scoring” is expanded beyond the 519-item set (Appendix B’s 186k figure is not clearly item-level human re-scoring of all responses). The Resolution Bound (Prop. 2) already flags median SE(b)≈0.064 and min detectable Δb≈0.253 under the larger judge n; the human subsample is an order of magnitude thinner. If the 85 figure is driven by under-powered human CIs rather than true tail separation, the strongest claim weakens even though rank correlation ρ=0.89 and ICC=0.85 remain intact.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that open-weight LLMs with matched error rates can still differ substantially in the shape of their error-severity distributions, and that this shape should be reported alongside accuracy. It introduces ERRORQUAKE-10K (10,000 queries, 8 domains, 5 tiers), scores responses on a 0–4 severity grid with a dual LLM-judge pipeline plus a 519-item three-rater human study, and summarizes upper-tail behavior with a Gutenberg–Richter slope b for 21 open-weight models. The headline is that 85 of 210 model pairs have disjoint 95% b CIs at |Δε|<0.05 on human-consensus scoring (31 on the LLM-judge baseline), supported by a Non-Reducibility Theorem, mutual-information estimates, a mechanism taxonomy, and several robustness checks; pre-registered failures (Exp. 3 magnitude calibration, S1 coarsening) are reported.","tokens_in":27281,"tokens_out":1843,"duration_ms":28202,"significance":"If the matched-accuracy discriminator is solid, the paper supplies a practically useful second axis for hallucination evaluation: catastrophic load can diverge by an order of magnitude at fixed ε, which matters for deployment gating in high-stakes factual settings. Strengths include an open 10K benchmark and scoring toolkit, honest pre-registration outcomes, human ICC(2,k=3)=0.85 with human–judge rank ρ=0.89, a non-parametric tail-ratio cross-check, domain jackknife and aggregation robustness on the judge baseline, and an explicit mechanism taxonomy (κ=0.83) linking high severity to fabrication. The Non-Reducibility result and I(b;model|ε)=1.56 bits frame the claim cleanly. These are real contributions to evaluation methodology even if some secondary scaling claims remain sensitivity analyses.","major_comments":[{"comment":"§4.1 headline and §8: The central 85/210 disjoint-CI claim is stated on “human-consensus scoring,” yet the human study is only 519 items (~35 per model × 15 models), which §8 itself calls “limited for per-model b-value precision.” Prop. 2’s Resolution Bound (median SE≈0.064, min detectable Δb≈0.253) is calibrated to the large judge n; with ~35 items the upper-tail exceedance counts after mmin selection are typically far smaller, so human b CIs should be much wider and the 85-pair count more fragile. The manuscript also cites a “full 186,521-item human-consensus scoring (Appendix B)” for higher precision, but Appendix B is the 4K-vs-10K scale-up and does not document item-level human re-scoring of the full catalog. Please define exactly how human-consensus b and its bootstrap CIs are constructed for all 21 models, report per-model n≥mmin and SE(b) under that construction, and recompute th","section":"§4.1, §8, Appendix B, Prop. 2"},{"comment":"§2 dual-judge reliability and Appendix L: Pre-tiebreak ICC(2,1)=0.374 and final averaged ICC(2,k=2)=0.545 are only fair–moderate, and the 340-item audit finds 33.5% overcall at score 2.0 (vs 13.7% human). S2 overcall correction narrowly fails the pre-registered ρ>0.85 threshold (0.847). Because b is an upper-tail slope on a 9-level grid, systematic mid-scale overcall and moderate inter-judge agreement can shift mmin selection and compress or inflate tails differently across models. The paper should either (i) show that the matched-accuracy disjoint-CI count is stable under a human-calibrated overcall correction applied model-wise, or (ii) center the headline on the more conservative judge-baseline 31-pair result plus the non-parametric tail-ratio check, with human data used strictly for ranking/ICC validation.","section":"§2, §4.6 S2, Appendix L"},{"comment":"§4.3 and S5 (Appendix T): The dense scaling claim ρs=−0.562 (human −0.86) is already demoted to a sensitivity observation, but the mmin sweep flips the sign to +0.79/+0.84 under fixed mmin=0.5 or 1.5. That means “larger models have heavier tails” is true only for the KS-selected upper-tail estimator, while bulk decay moves the opposite way. Given that free parameter mmin is load-bearing for both b ranking and the scaling narrative, the main text should state more sharply that b is not a unique summary of the severity distribution, report bulk vs upper-tail slopes side by side in the main results table, and avoid language that equates b with a single model-level “heaviness” without specifying the cutoff regime.","section":"§4.3, §4.6 S5, Appendix T"},{"comment":"§4.2 operational “heavy-tailed” claim: On a bounded discrete grid {0.5,…,4.0} with only eight positive bins, asymptotic heavy-tail language is unavailable, and zero models are pure power-law; 13/21 are stretched exponential and 4 exponential. The operational definition (“slower than exponential on the positive grid, or excess mass at M≥2.5”) is reasonable but should be the primary claim in the abstract/title framing. Gutenberg–Richter b remains a useful slope summary, but the paper should not lean on seismological heavy-tail connotations beyond what the discrete BIC/Vuong evidence supports, especially when four models are BIC-best exponential yet still enter the b catalog.","section":"§4.2, Abstract, title"}],"minor_comments":[{"comment":"Table 1 and §4.4: Exp. 3 is correctly marked FAIL on magnitude calibration; consider moving the rank-only ρs=0.443 result fully into a “partial signal” subsection so readers do not over-read the catastrophe-prediction language in the contributions list (C6).","section":"Table 1, §4.4, C6"},{"comment":"Figure 1 vs Figure 4: b values in the four-panel figure (e.g. deepseek-v3.2 b=0.66) do not always match Table 3 (0.595) or the all-model grid labels; reconcile fitted b across figures and the main table.","section":"Figure 1, Figure 4, Table 3"},{"comment":"§5 Theorem 1: The existence construction is clear, but the empirical I(b;model|ε)=1.56 bits depends on a 5-bin discretisation that is not specified in the main text; state bin edges and sensitivity in Appendix A.","section":"§5, Appendix A"},{"comment":"NeurIPS checklist items 8, 12, 14, 15 are answered No (compute details, upstream licenses, compensation, IRB). For a journal version, add a short compute/API note, license table for evaluated models, and human-subjects protocol statement even if review was not required.","section":"Checklist / §8"},{"comment":"Notation: ε is used for error rate and also appears near exponential fits; consider e or err for error rate to avoid clash with base of the natural exponential in the Aki formula.","section":"§2, §5"}],"recommendation":"major_revision","confidential_remarks":"The evaluation idea is timely and the honesty about failed pre-registered criteria is a plus. The main risk for the journal is that the abstract’s 85-pair human-consensus headline may not be reproducible from the documented 519-item human sample; if authors cannot supply a clear full-catalog human-consensus pipeline, I would still consider acceptance after revision centered on the judge-baseline discriminator (31 pairs) plus human ranking validation, which is already interesting. Scope is open-weight 3–37B only; that is fine if stated up front."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing: at matched error rate, open-weight models can still differ a lot in upper-tail severity, and this paper actually measures that with a 10k benchmark, 21 models, bootstrap CIs, and human validation instead of just asserting it.\n\nWhat is new is the operational package: ERRORQUAKE-10K, a 9-level 0–4 scale, GR-style b as a matched-accuracy discriminator, the empirical I(b; model | ε) = 1.56 bits / R² = 0.356 decomposition, and the mechanism taxonomy (retrieval → fabrication as severity rises, with size coupling). Binary hallucination work and coarser severity bins already exist; this is the first clean cross-model tail-shape comparison at this scale. They report pre-registered failures honestly (Exp. 3 magnitude prediction fails; S1 coarsening kills ranking; S2 borderline). Human ICC = 0.85, human–judge rank ρ = 0.89, and dense scaling stronger on human data (−0.86) are real credits. Code, data, and Croissant metadata are released.\n\nSoft spots, in proportion. The stress-test is partly right: the 85-pair human-consensus headline rests on ~35 items/model in the 519-item study, so per-model human b CIs are underpowered relative to the judge baseline (31 pairs, jackknife 25–41, aggregation alternatives all ≥3). Appendix B’s 186k “human-consensus” figure is not clearly full item-level re-scoring, so the abstract overweights 85. That does not kill the claim—the judge discriminator, non-parametric tail-ratio check, and rank evidence still show severity is not reducible to ε. Other real limits: 33.5% judge overcall at 2.0, mmin choice flips the scaling sign (upper tail vs bulk), LLM-written queries, and a nearly tautological non-reducibility theorem. Scope is open-weight 3–37B active; no frontier proprietary claim.\n\nThis is for people who select or audit models for factual deployment and for eval researchers who want a second axis next to accuracy. Math is standard Aki MLE + bootstrap; citations are fair. I would send it to peer review. Engage if you care about risk-aware LLM evaluation; treat the 85 figure as the softest number and lean on the robustness suite.","headline":"Useful evaluation axis with a real matched-accuracy discriminator and honest pre-reg failures; the 85-pair human headline is thinner than the abstract sells, but the core claim still holds on the judge baseline and rank evidence.","tokens_in":27892,"tokens_out":602,"would_cite":true,"duration_ms":5866,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"At the same error rate, open-weight LLMs still differ sharply in how severe their worst mistakes are.","keywords":["hallucination severity","error severity distribution","Gutenberg-Richter b-value","open-weight LLMs","matched-accuracy discrimination","LLM evaluation","fabrication vs retrieval"],"falsifier":"Find a substantial set of matched-accuracy model pairs whose human-scored severity tails still produce overlapping b confidence intervals, or show that coarsening or re-anchoring the 9-level severity scale leaves the matched-accuracy discrimination count near zero.","tokens_in":27701,"feed_emoji":"📉","tokens_out":604,"duration_ms":5598,"temperature":0.7,"pith_summary":"Standard hallucination benchmarks count every wrong answer the same way, so a slightly off date and a fabricated court ruling both add one to the error rate. This paper argues that the shape of the severity distribution is a separate and useful fact about a model. Using a 10,000-query benchmark scored on a continuous 0–4 severity scale, the authors fit a Gutenberg–Richter upper-tail slope b for each of 21 open-weight models and show that many model pairs with nearly identical error rates have non-overlapping confidence intervals on b. Human raters confirm the ranking and the reliability of the scale. A short theorem shows that b is not informationally redundant with the error rate, and a taxonomy shows that high-severity errors are disproportionately fabrications rather than simple retrieval slips. The practical upshot is that reporting severity distribution alongside accuracy gives deployment-relevant information that the scalar error rate alone cannot supply.","feed_headline":"Same error rate, different catastrophic tails","feed_subtitle":"Open-weight LLMs with matched accuracy still diverge sharply in how severe their worst mistakes are","key_machinery":"The severity distribution index b — the Gutenberg–Richter upper-tail slope of the magnitude-frequency relation for continuous 0–4 error scores — together with the Non-Reducibility Theorem that b is informationally independent of the error rate ε.","core_discovery":"Across 210 pairs among 21 open-weight models, 85 have disjoint 95% confidence intervals on the upper-tail severity slope b at matched accuracy on human-consensus scoring. Severity profile therefore carries model-discriminative information that the scalar error rate cannot express; the paper proves this non-redundancy formally and ties the heavier tails to a shift from retrieval errors toward fabrications.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Matched accuracy, mismatched catastrophic error tails","85 model pairs split on severity tails at same accuracy","Severity slope b separates models accuracy cannot","Same error rate hides divergent heavy-tailed mistakes","Error severity profiles carry info error rates lack"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That continuous 0–4 severity scores from the dual-judge pipeline and the human consensus are reliable enough, and that the chosen upper-tail cutoff isolates the deployment-relevant slope rather than bulk decay.","fun_headline_variants_meta":{"raw":{"variants":["Matched accuracy, mismatched catastrophic error tails","85 model pairs split on severity tails at same accuracy","Severity slope b separates models accuracy cannot","Same error rate hides divergent heavy-tailed mistakes","Error severity profiles carry info error rates lack"]},"model":"grok-4.5","effort":"low","cost_usd":0.005114,"raw_usage":{"total_tokens":1539,"prompt_tokens":934,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":51140000,"prompt_tokens_details":{"text_tokens":934,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":555,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":934,"tokens_out":50,"duration_ms":4414,"temperature":1.0,"reasoning_tokens":555,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T20:39:16.356242+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Find a substantial set of matched-accuracy model pairs whose human-scored severity tails still produce overlapping b confidence intervals, or show that coarsening or re-anchoring the 9-level severity scale leaves the matched-accuracy discrimination count near zero.","supporting_citations":[],"review_version":1}