REVIEW 3 major objections 4 minor
Does a Developed Comorbidity Index Really Add Value? A Selection-Aware Bootstrap for Post-Selection Concordance
T0 review · 3 major / 4 minor · reviewed 2026-07-15 · grok-4.5
Pith's one-line read Standard optimism correction miscalibrates claims that a best-of-several comorbidity index adds discrimination; a selection-aware bootstrap restores coverage.
desk verdict Abstract-only: coherent diagnosis of structural miscalibration in incremental concordance under multi-candidate selection, plus a practical bootstrap fix—worth a serious look if the full paper delivers the sims and code. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The selection-aware bootstrap: inside each resample the entire candidate-screening procedure is re-run, the best-scoring construction is chosen afresh, and its incremental concordance relative to the fixed comparator is recorded, thereby capturing the selection optimism that the ordinary optimism correction leaves out.
What would settle it
In a fully-known-truth simulation with a true null of zero incremental concordance and a large number of similar-quality candidates, check whether the selection-aware bootstrap's 95 percent interval still covers the true zero at approximately the nominal rate while the standard optimism-corrected interval does not.
Extended reading notes
Core claim
When several candidate comorbidity constructions are screened and only the best is reported, the usual optimism-corrected confidence interval for incremental concordance over a fixed comparator is structurally biased because it omits the selection-induced winner's-curse term; a bootstrap that re-executes the same best-of-several selection inside each resample with the comparator held fixed removes that bias and recovers near-nominal coverage.
Load-bearing premise
That re-running best-of-several selection inside each bootstrap resample, with the comparator held fixed, fully captures the selection-induced optimism that arises under the actual development processes used for comorbidity indices.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript argues that disease-specific comorbidity indices are routinely developed by screening several candidate constructions and reporting the best-scoring one against a fixed comparator (e.g., Charlson or Elixhauser), and that the standard optimism correction does not validate claims of added discriminative value under that practice. Because the correction treats the selected model as if it were the only one fit, it omits the winner’s-curse term from multi-candidate selection; the resulting confidence interval for incremental concordance is therefore structurally miscalibrated (not merely small-sample optimistic) and does not shrink with sample size. At a true null it inflates false claims of added value as more candidates are screened. The authors propose a drop-in selection-aware bootstrap that re-runs best-of-several selection inside each resample with the comparator held fixed. In a fully-known-truth simulation they report that 95% coverage under the standard correction falls from 0.94 (one candidate) to 0.70 (one hundred), while the selection-aware interval stays near nominal, matches a calibrated CV interval, and is at least as powerful at matched error rate. Results are said to hold under Uno’s concordance; a semi-synthetic survey experiment is used to illustrate when the correction matters. Scope is discrimination only; software and results are claimed to reproduce every finding.
Significance. If the structural-miscalibration claim and the coverage recovery hold under realistic selection pipelines, the paper would change reporting practice for a large class of clinical-prediction and comorbidity-index papers: multi-candidate screening is common, and claims of incremental concordance over Charlson/Elixhauser would need selection-aware intervals. Strengths claimed in the abstract include a fully-known-truth simulation with explicit coverage targets, a semi-synthetic check, matching to calibrated CV, reproducibility of every finding via software, and a drop-in procedure rather than a new index. Those are the right kinds of deliverables for a methods contribution in this area. The practical guidance (most needed with many similar-quality candidates and few events per candidate) is actionable if supported by the design.
major comments (3)
- Abstract-only review: the load-bearing claim that re-running best-of-several selection inside each bootstrap resample (comparator fixed) fully captures the selection-induced optimism of real comorbidity-index pipelines cannot be verified from the abstract alone. The central coverage result (standard 0.94→0.70 as candidates go 1→100; selection-aware near nominal) is the right kind of evidence, but without the simulation protocol—candidate-generation process, similarity structure among candidates, event-per-candidate regimes, and exact selection rule—it is impossible to judge whether the design represents the pipelines the paper targets. This must be checkable in Methods and code before the structural-bias claim can be accepted.
- The abstract asserts that the standard interval is structurally miscalibrated and does not shrink as the sample grows (omitted winner’s-curse term). That is a strong asymptotic claim. The full manuscript needs an explicit decomposition or asymptotic argument separating ordinary optimism from the selection term, and a demonstration that the selection-aware bootstrap’s coverage remains calibrated as n grows under a true null with K>1 candidates. Without that, the ‘structural’ (vs finite-sample) diagnosis remains an assertion.
- The semi-synthetic survey experiment is cited as confirming ‘when the correction matters,’ but the abstract does not state the ground-truth construction, the candidate set, or the event counts. For the practical recommendation (report selection-aware intervals when several constructions were tried, especially with few events per candidate) to be load-bearing, the experiment must show coverage or false-claim rates under regimes that match published comorbidity-index developments, not only under a convenient synthetic null.
minor comments (4)
- Abstract wording ‘does not shrink as the sample grows’ should be paired in the full text with a precise statement of what is held fixed (K, candidate correlation, event rate) so readers do not over-read it as a universal n-independence claim.
- Clarify early whether ‘best-of-several’ is restricted to a fixed finite menu of constructions or also covers stepwise/penalized search over large comorbidity feature sets; the bootstrap recipe and the practical advice differ in those two cases.
- The claim that results hold under Uno’s concordance is important for time-to-event comorbidity work; the full paper should state whether Harrell’s C or other concordance estimators were also checked and whether the selection-aware correction is estimator-agnostic.
- Software and full reproducibility are promised; the manuscript should point to a permanent archive (DOI/version) and a single script that regenerates the coverage table and the semi-synthetic figure.
Circularity Check
No circularity detectable from abstract-only material; methodological claims are evaluated against external known-truth benchmarks, not quantities defined by the same fit.
full rationale
Only the abstract is available, so no internal equations, definitions, or self-citation chains can be inspected for reduction-by-construction. From the abstract alone the paper is a methodological proposal: it argues that standard optimism correction omits a winner's-curse term under multi-candidate selection, and proposes a selection-aware bootstrap that re-runs best-of-several selection inside each resample. The claimed performance (coverage collapse of the standard interval from ~0.94 to ~0.70 as candidates increase; near-nominal recovery for the proposed interval) is evaluated against a fully-known-truth simulation and a semi-synthetic experiment, i.e., external benchmarks with known truth or fixed survey data rather than a quantity defined by the same fitting step. There is no indication that a 'prediction' is a fitted parameter renamed, that a uniqueness theorem is imported from the authors, or that an ansatz is smuggled via self-citation. The reader's residual concern (whether the simulation design fully captures real comorbidity-index pipelines) is a correctness/external-validity issue, not circularity. Honest non-finding: score 0; steps empty.
Assumptions & free parameters
assumptions (3)
- domain assumption Re-running best-of-several selection inside each bootstrap resample with the comparator held fixed correctly accounts for selection-induced optimism in incremental concordance.
- domain assumption Standard optimism correction treats the selected model as if it were the only one ever fit and therefore omits the winner's-curse term.
- standard math Bootstrap resampling and concordance estimators (including Uno's) behave as assumed under the simulation and survey designs used.
Cite this review
Pith. "Pith review of Does a Developed Comorbidity Index Really Add Value? A Selection-Aware Bootstrap for Post-Selection Concordance." pith.science (2026). https://pith.science/paper/7BTDZW4K
@misc{pith2026260712445,
author = {Pith},
title = {Pith review of: Does a Developed Comorbidity Index Really Add Value? A Selection-Aware Bootstrap for Post-Selection Concordance},
year = {2026},
howpublished = {\url{https://pith.science/paper/7BTDZW4K}},
note = {Machine review of arXiv:2607.12445}
}
read the original abstract
Disease-specific comorbidity indices are routinely developed by building several candidate constructions and reporting the best-scoring one, then claiming it adds discriminative value over a fixed off-the-shelf comparator such as the Charlson or Elixhauser score. We show that the optimism correction in standard use does not make that claim valid. Because it corrects the selected model as if it were the only one ever fit, it omits the winner's-curse term from choosing the best of several candidates; so its confidence interval for the incremental concordance is not merely optimistic in small samples but structurally miscalibrated, and does not shrink as the sample grows. At a true null it inflates false claims of added value above the nominal level, increasingly so as more candidates are screened. We introduce a drop-in selection-aware bootstrap that re-runs the best-of-several selection inside each resample with the comparator held fixed, removing the structural bias. In a fully-known-truth simulation, 95% coverage under the standard correction falls from 0.94 with one candidate to 0.70 with a hundred, while the selection-aware interval holds near nominal; its coverage matches a calibrated cross-validation interval, and at a matched error rate it is at least as powerful. The results hold under Uno's concordance, and a semi-synthetic experiment on real survey data confirms when the correction matters. In practice, if several constructions were tried, report a selection-aware interval, most needed with many similar-quality candidates and few events per candidate. The scope is discrimination only; software and results reproduce every finding.
Reviewed July 15, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.