Pith. sign in

REVIEW 3 major objections 4 minor

Does a Developed Comorbidity Index Really Add Value? A Selection-Aware Bootstrap for Post-Selection Concordance

T0 review · 3 major / 4 minor · reviewed 2026-07-15 · grok-4.5

Pith's one-line read Standard optimism correction miscalibrates claims that a best-of-several comorbidity index adds discrimination; a selection-aware bootstrap restores coverage.

desk verdict Abstract-only: coherent diagnosis of structural miscalibration in incremental concordance under multi-candidate selection, plus a practical bootstrap fix—worth a serious look if the full paper delivers the sims and code. read the letter →

arxiv 2607.12445 v1 pith:7BTDZW4K submitted 2026-07-14 stat.ME stat.AP

classification stat.MEstat.AP MSC 62F4062P1062N02
keywords comorbidityindexconcordanceoptimismcorrectionwinner'scurseselection-awarebootstrapincrementaldiscriminationpost-selectioninferenceUno's
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Disease-specific comorbidity indices are often built by screening several candidate constructions, publishing the highest-scoring one, and claiming it improves discrimination over a fixed off-the-shelf score such as Charlson or Elixhauser. The paper shows that the optimism correction routinely applied to that claim is not enough: it treats the selected index as if it were the only model ever considered, so it misses the winner's-curse term that arises from choosing the best of many. The resulting confidence interval for the incremental concordance is therefore structurally miscalibrated; it does not shrink with sample size and, under a true null, produces false claims of added value at rates well above the nominal level as the number of candidates grows. The authors supply a drop-in selection-aware bootstrap that re-runs the best-of-several selection inside every resample while holding the comparator fixed, thereby restoring near-nominal coverage. Fully-known-truth simulations and a semi-synthetic survey experiment confirm that the standard interval's coverage collapses from roughly 0.94 with one candidate to 0.70 with a hundred, while the new interval stays calibrated and remains competitive with a carefully calibrated cross-validation procedure.

What carries the argument

The selection-aware bootstrap: inside each resample the entire candidate-screening procedure is re-run, the best-scoring construction is chosen afresh, and its incremental concordance relative to the fixed comparator is recorded, thereby capturing the selection optimism that the ordinary optimism correction leaves out.

What would settle it

In a fully-known-truth simulation with a true null of zero incremental concordance and a large number of similar-quality candidates, check whether the selection-aware bootstrap's 95 percent interval still covers the true zero at approximately the nominal rate while the standard optimism-corrected interval does not.

Watch

Extended reading notes

Core claim

When several candidate comorbidity constructions are screened and only the best is reported, the usual optimism-corrected confidence interval for incremental concordance over a fixed comparator is structurally biased because it omits the selection-induced winner's-curse term; a bootstrap that re-executes the same best-of-several selection inside each resample with the comparator held fixed removes that bias and recovers near-nominal coverage.

Load-bearing premise

That re-running best-of-several selection inside each bootstrap resample, with the comparator held fixed, fully captures the selection-induced optimism that arises under the actual development processes used for comorbidity indices.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript argues that disease-specific comorbidity indices are routinely developed by screening several candidate constructions and reporting the best-scoring one against a fixed comparator (e.g., Charlson or Elixhauser), and that the standard optimism correction does not validate claims of added discriminative value under that practice. Because the correction treats the selected model as if it were the only one fit, it omits the winner’s-curse term from multi-candidate selection; the resulting confidence interval for incremental concordance is therefore structurally miscalibrated (not merely small-sample optimistic) and does not shrink with sample size. At a true null it inflates false claims of added value as more candidates are screened. The authors propose a drop-in selection-aware bootstrap that re-runs best-of-several selection inside each resample with the comparator held fixed. In a fully-known-truth simulation they report that 95% coverage under the standard correction falls from 0.94 (one candidate) to 0.70 (one hundred), while the selection-aware interval stays near nominal, matches a calibrated CV interval, and is at least as powerful at matched error rate. Results are said to hold under Uno’s concordance; a semi-synthetic survey experiment is used to illustrate when the correction matters. Scope is discrimination only; software and results are claimed to reproduce every finding.

Significance. If the structural-miscalibration claim and the coverage recovery hold under realistic selection pipelines, the paper would change reporting practice for a large class of clinical-prediction and comorbidity-index papers: multi-candidate screening is common, and claims of incremental concordance over Charlson/Elixhauser would need selection-aware intervals. Strengths claimed in the abstract include a fully-known-truth simulation with explicit coverage targets, a semi-synthetic check, matching to calibrated CV, reproducibility of every finding via software, and a drop-in procedure rather than a new index. Those are the right kinds of deliverables for a methods contribution in this area. The practical guidance (most needed with many similar-quality candidates and few events per candidate) is actionable if supported by the design.

major comments (3)
  1. Abstract-only review: the load-bearing claim that re-running best-of-several selection inside each bootstrap resample (comparator fixed) fully captures the selection-induced optimism of real comorbidity-index pipelines cannot be verified from the abstract alone. The central coverage result (standard 0.94→0.70 as candidates go 1→100; selection-aware near nominal) is the right kind of evidence, but without the simulation protocol—candidate-generation process, similarity structure among candidates, event-per-candidate regimes, and exact selection rule—it is impossible to judge whether the design represents the pipelines the paper targets. This must be checkable in Methods and code before the structural-bias claim can be accepted.
  2. The abstract asserts that the standard interval is structurally miscalibrated and does not shrink as the sample grows (omitted winner’s-curse term). That is a strong asymptotic claim. The full manuscript needs an explicit decomposition or asymptotic argument separating ordinary optimism from the selection term, and a demonstration that the selection-aware bootstrap’s coverage remains calibrated as n grows under a true null with K>1 candidates. Without that, the ‘structural’ (vs finite-sample) diagnosis remains an assertion.
  3. The semi-synthetic survey experiment is cited as confirming ‘when the correction matters,’ but the abstract does not state the ground-truth construction, the candidate set, or the event counts. For the practical recommendation (report selection-aware intervals when several constructions were tried, especially with few events per candidate) to be load-bearing, the experiment must show coverage or false-claim rates under regimes that match published comorbidity-index developments, not only under a convenient synthetic null.
minor comments (4)
  1. Abstract wording ‘does not shrink as the sample grows’ should be paired in the full text with a precise statement of what is held fixed (K, candidate correlation, event rate) so readers do not over-read it as a universal n-independence claim.
  2. Clarify early whether ‘best-of-several’ is restricted to a fixed finite menu of constructions or also covers stepwise/penalized search over large comorbidity feature sets; the bootstrap recipe and the practical advice differ in those two cases.
  3. The claim that results hold under Uno’s concordance is important for time-to-event comorbidity work; the full paper should state whether Harrell’s C or other concordance estimators were also checked and whether the selection-aware correction is estimator-agnostic.
  4. Software and full reproducibility are promised; the manuscript should point to a permanent archive (DOI/version) and a single script that regenerates the coverage table and the semi-synthetic figure.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity detectable from abstract-only material; methodological claims are evaluated against external known-truth benchmarks, not quantities defined by the same fit.

full rationale

Only the abstract is available, so no internal equations, definitions, or self-citation chains can be inspected for reduction-by-construction. From the abstract alone the paper is a methodological proposal: it argues that standard optimism correction omits a winner's-curse term under multi-candidate selection, and proposes a selection-aware bootstrap that re-runs best-of-several selection inside each resample. The claimed performance (coverage collapse of the standard interval from ~0.94 to ~0.70 as candidates increase; near-nominal recovery for the proposed interval) is evaluated against a fully-known-truth simulation and a semi-synthetic experiment, i.e., external benchmarks with known truth or fixed survey data rather than a quantity defined by the same fitting step. There is no indication that a 'prediction' is a fitted parameter renamed, that a uniqueness theorem is imported from the authors, or that an ansatz is smuggled via self-citation. The reader's residual concern (whether the simulation design fully captures real comorbidity-index pipelines) is a correctness/external-validity issue, not circularity. Honest non-finding: score 0; steps empty.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

Abstract-only: free parameters of the method itself are not enumerated. The central claim rests on standard bootstrap validity under resampling of the selection process, on the fixed-comparator design, and on simulation/semi-synthetic designs that stand in for real comorbidity-index development. No new physical entities; the 'selection-aware bootstrap' is a procedure, not an invented scientific object.

assumptions (3)
  • domain assumption Re-running best-of-several selection inside each bootstrap resample with the comparator held fixed correctly accounts for selection-induced optimism in incremental concordance.
    Load-bearing modeling choice for the proposed interval; stated as the mechanism that removes structural bias.
  • domain assumption Standard optimism correction treats the selected model as if it were the only one ever fit and therefore omits the winner's-curse term.
    Core diagnosis of why existing practice is miscalibrated; if false, the structural-bias claim collapses.
  • standard math Bootstrap resampling and concordance estimators (including Uno's) behave as assumed under the simulation and survey designs used.
    Background statistical regularity needed for coverage claims; not proved in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Does a Developed Comorbidity Index Really Add Value? A Selection-Aware Bootstrap for Post-Selection Concordance." pith.science (2026). https://pith.science/paper/7BTDZW4K

@misc{pith2026260712445,
  author       = {Pith},
  title        = {Pith review of: Does a Developed Comorbidity Index Really Add Value? A Selection-Aware Bootstrap for Post-Selection Concordance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7BTDZW4K}},
  note         = {Machine review of arXiv:2607.12445}
}
read the original abstract

Disease-specific comorbidity indices are routinely developed by building several candidate constructions and reporting the best-scoring one, then claiming it adds discriminative value over a fixed off-the-shelf comparator such as the Charlson or Elixhauser score. We show that the optimism correction in standard use does not make that claim valid. Because it corrects the selected model as if it were the only one ever fit, it omits the winner's-curse term from choosing the best of several candidates; so its confidence interval for the incremental concordance is not merely optimistic in small samples but structurally miscalibrated, and does not shrink as the sample grows. At a true null it inflates false claims of added value above the nominal level, increasingly so as more candidates are screened. We introduce a drop-in selection-aware bootstrap that re-runs the best-of-several selection inside each resample with the comparator held fixed, removing the structural bias. In a fully-known-truth simulation, 95% coverage under the standard correction falls from 0.94 with one candidate to 0.70 with a hundred, while the selection-aware interval holds near nominal; its coverage matches a calibrated cross-validation interval, and at a matched error rate it is at least as powerful. The results hold under Uno's concordance, and a semi-synthetic experiment on real survey data confirms when the correction matters. In practice, if several constructions were tried, report a selection-aware interval, most needed with many similar-quality candidates and few events per candidate. The scope is discrimination only; software and results reproduce every finding.

Discussion (0). Sign in to comment.

Pith tools

Reviewed July 15, 2026 · model on record in the stance chip above.