REVIEW 3 major objections 1 cited by
The C-index Multiverse
T0 review · 3 major / 0 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read As declared in its abstract, the paper claims that the C-index is not a unique number: seemingly equal R and Python implementations can disagree for the same model and data.
desk verdict The submitted full text is a different paper, so the C-index multiverse claim cannot be assessed from this submission. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the 'C-index multiverse': the family of distinct numerical values that different implementations produce for what is nominally the same concordance index. The mechanisms that generate the multiplicity are tie handling, the choice and adjustment for censoring, and the non-standardised way risk is summarised from survival distributions. These choices are what make the metric implementation-dependent rather than a property of model and data alone.
What would settle it
Run one fixed survival dataset and one fitted Cox model through several R and Python C-index functions, varying only tie-handling and censoring-adjustment options; if every implementation returns exactly the same concordance value, the multiverse claim is falsified for that configuration. A systematic sweep across packages and options would settle the claim. For this record, comparing the abstract's survival-analysis claims with the body's aggregation-equation theorems already shows the body text does not support them.
Extended reading notes
Core claim
On its own terms, the paper claims that the concordance index is multiversal: fixing the data, the model, and the training/test split, the C-index value still depends on the software package chosen, because packages differ in tie-breaking conventions, in censoring adjustments (for example Harrell's, Uno's and Antolini's estimators), and in how they convert survival distributions into a scalar risk score. The paper reports numerical demonstrations on publicly available breast cancer data and semi-synthetic examples, across models from Cox proportional hazards to recent deep learning survival methods, showing that these implementation choices change the reported score. A caveat specific to thi
Load-bearing premise
The claim assumes the tested R/Python packages were representative of 'available software' and were invoked correctly; in this record it also assumes the abstract describes the actual paper, since the supplied full text is a different manuscript.
Editorial extensions
If this is right
- Published C-index values should be reported together with the software, version, tie-breaking rule, and censoring adjustment used, otherwise the number is ambiguous.
- Model comparisons that use different packages for different models can be biased by implementation choice rather than actual predictive performance.
- Benchmarking studies and leader boards for survival models need a common, documented protocol for computing the C-index before rankings are meaningful.
- Analysts need unified documentation and a checklist of pitfalls, which the paper positions as a guideline for navigating the multiverse.
Reading between the lines
- The same implementation-dependence likely affects other rank-based survival metrics, such as time-dependent AUC or Brier-score variants, since they share the same tie and censoring conventions; the paper does not claim this.
- A natural testable extension would be a cross-package test suite that pins a reference dataset and reports every package's output under each documented tie and censoring option, making the boundaries of the multiverse explicit.
- In this record the empirical demonstration cannot be verified, because the supplied body text belongs to a different paper; verifying the C-index multiverse requires the actual methods and experiments or a corrected manuscript.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This submission presents, in its abstract, an audit study of concordance-index (C-index) implementations across R and Python packages. The authors claim to demonstrate a 'C-index multiverse': for the same survival model and data, seemingly equivalent implementations yield different C-index values, and they attribute the variation to tie handling, censoring adjustment, and risk-score summarization choices. They further claim to illustrate the consequences on publicly available breast cancer data and semi-synthetic examples, and to offer guidelines and publicly available code. However, the full text supplied with the submission is not the audit paper. It is arXiv:2508.14827, 'Vacuum bubble and fissure formation in collective motion with competing attractive and repulsive forces', a PDE study of pattern formation in interacting particle systems. No section, equation, table, figure, or appendix of the claimed C-index study appears in the manuscript. The central claim and all supporting evidence are therefore absent from the submitted document.
Significance. The scientific question behind the abstract is genuinely significant. If an audit were carefully executed and documented, showing that C-index values vary across nominally equivalent software settings, it would have direct implications for reproducibility, model comparison, and reporting standards in time-to-event analysis. The abstract names concrete mechanisms (tie handling, censoring adjustment, risk-summary choices) and promises a practical guideline and public code, which are potentially valuable contributions. However, none of this can be evaluated from the submitted text. The manuscript contains no machine-checked proofs, no reproducible code listing, no package/version documentation, and no numerical experiments. The strengths that would justify publication are asserted in the abstract but are entirely absent from the manuscript body.
major comments (3)
- [Full Text] The submitted full text is the complete text of a different arXiv paper (2508.14827) on vacuum bubble and fissure formation in collective motion. It contains no part of the C-index audit described in the abstract: no methods, no package list, no datasets, no numerical results, and no discussion of survival analysis. This is not a missing appendix or a presentation issue; the claimed subject of the paper is absent, so the central claim is unsupported in the submitted document.
- [Abstract] The abstract's central claim to 'demonstrate the existence of a C-index multiverse' is not backed by any evidence in the manuscript. Specifically, there is no specification of the R/Python packages and versions tested, no formal statement of the tie-handling or censoring-adjustment rules compared, no description of the breast cancer or semi-synthetic datasets, and no numerical tables or figures showing differences in C-index outputs. The inductive generalization to 'available R and python software' therefore rests on no auditable evidence in this submission.
- [Abstract / Code Availability] The abstract states that 'All code is publicly available at www.github.com/BBolosSierra/CindexMultiverse,' but the full text contains no such URL or repository information, nor any code listing, package manifest, or version record. Even if the link is valid externally, the submitted manuscript provides no way to verify which implementations, options, and data were used, so the reproducibility claim cannot be checked.
Circularity Check
No circularity identified: the supplied full text is a different paper, so the C-index audit's derivation chain is absent; absence of evidence is not circularity.
full rationale
The abstract of arXiv:2508.14821 claims an empirical C-index multiverse among R/Python implementations, with variation from tie handling, censoring adjustment, and risk summary. This is an audit claim: it would stand or fall on the documented package versions, datasets, and code, not on a fitted parameter or a quantity defined in terms of the target result. The full text supplied with the submission, however, is arXiv:2508.14827, a nonlinear-PDE paper on vacuum bubbles, not the C-index audit. It contains no methods, package list, tables, or numerical results from which the multiverse claim could be derived. I therefore cannot walk a derivational chain for the C-index claim, and I cannot exhibit any equation or fitted parameter that reduces to its own input. Under the standing rule that circularity must be shown by quotation and specific reduction, the correct verdict is no circularity found (score 0). The mismatch is a serious completeness/reproducibility problem for this submission—the central empirical claim is unsupported in the provided text—but unsupported is not the same as circular. No step in the visible abstract is self-definitional, no fitted input is relabeled as a prediction, and no load-bearing self-citation is present in the submitted evidence.
Assumptions & free parameters
assumptions (2)
- domain assumption The selected R and Python packages and options are representative of the C-index multiverse and are executed correctly.
- domain assumption There is a single agreed target value for the C-index on a given model and data, so that discrepancies across packages represent a genuine multiverse rather than bugs or misconfiguration.
Cite this review
Pith. "Pith review of The C-index Multiverse." pith.science (2026). https://pith.science/paper/POCEKUBM
@misc{pith2026250814821,
author = {Pith},
title = {Pith review of: The C-index Multiverse},
year = {2026},
howpublished = {\url{https://pith.science/paper/POCEKUBM}},
note = {Machine review of arXiv:2508.14821}
}
read the original abstract
Quantifying out-of-sample discrimination performance for time-to-event outcomes is a fundamental step for model evaluation and selection in the context of predictive modelling. The concordance index, or C-index, is a widely used metric for this purpose, particularly with the growing development of machine learning methods. Beyond differences between proposed C-index estimators (e.g. Harrell's, Uno's and Antolini's), we demonstrate the existence of a C-index multiverse among available R and python software, where seemingly equal implementations can yield different results. This can undermine reproducibility and complicate fair comparisons across models and studies. Key variation sources include tie handling and adjustment to censoring. Additionally, the absence of a standardised approach to summarise risk from survival distributions, result in another source of variation dependent on input types. We demonstrate the consequences of the C-index multiverse when quantifying predictive performance for several survival models (from Cox proportional hazards to recent deep learning approaches) on publicly available breast cancer data, and semi-synthetic examples. Our work emphasises the need for better reporting to improve transparency and reproducibility. This article aims to be a useful guideline, helping analysts when navigating the multiverse, providing unified documentation and highlighting potential pitfalls of existing software. All code is publicly available at: www.github.com/BBolosSierra/CindexMultiverse.
Forward citations
Cited by 1 Pith paper
-
A reproducible and extensible framework for benchmarking competing risks survival models
A reproducible benchmarking framework and a new SHAP extension for competing-risks survival models, with results showing simpler regression models often match deep learning.
Reference graph
Works this paper leans on
-
[1]
K. Balasubramanian, S. Banerjee, and P. Rigollet. On the structure of stationary solutions to mckean-vlasov equations with applications to noisy transformers, 2025
work page 2025
-
[2]
F. Berthelin, D. Chiron, and M. Ribot. Stationary solutions with vacuum for a one-dimensional chemotaxis model with nonlinear pressure.Commun. Math. Sci., 14(1):147–186, 2016
work page 2016
-
[3]
M. Burger and A. Esposito. Porous medium equation and cross-diffusion systems as limit of nonlocal interaction.Nonlinear Anal., 235:Paper No. 113347, 30, 2023
work page 2023
-
[4]
E. Buzano and M. Golubitsky. Bifurcation on the hexagonal lattice and the planar B´ enard problem.Philos. Trans. Roy. Soc. London Ser. A, 308(1505):617–667, 1983
work page 1983
-
[5]
J. A. Carrillo, X. Chen, Q. Wang, Z. Wang, and L. Zhang. Phase transitions and bump solutions of the Keller-Segel model with volume exclusion.SIAM J. Appl. Math., 80(1):232– 261, 2020
work page 2020
-
[6]
J. A. Carrillo and R. S. Gvalani. Phase transitions for nonlinear nonlocal aggregation-diffusion equations.Comm. Math. Phys., 382(1):485–545, 2021
work page 2021
- [7]
-
[8]
P. Chossat and G. Iooss.The Couette-Taylor problem, volume 102 ofApplied Mathematical Sciences. Springer-Verlag, New York, 1994
work page 1994
Show all 16 references
-
[9]
Chossat and R
P. Chossat and R. Lauterbach.Methods in equivariant bifurcations and dynamical systems, volume 15 ofAdvanced Series in Nonlinear Dynamics. World Scientific Publishing Co., Inc., River Edge, NJ, 2000
2000
-
[10]
M. C. Cross and P. C. Hohenberg. Pattern formation outside of equilibrium.Rev. Mod. Phys., 65:851–1112, Jul 1993
1993
-
[11]
P. C. Fife. Pattern formation in gradient systems. InHandbook of dynamical systems, Vol. 2, pages 677–722. North-Holland, Amsterdam, 2002
2002
-
[12]
Golubitsky, I
M. Golubitsky, I. Stewart, and D. G. Schaeffer.Singularities and groups in bifurcation theory. Vol. II, volume 69 ofApplied Mathematical Sciences. Springer-Verlag, New York, 1988
1988
-
[13]
P.-E. Jabin. A review of the mean field limits for Vlasov equations.Kinet. Relat. Models, 7(4):661–711, 2014. 32
2014
-
[14]
Roose and R
D. Roose and R. Szalai. Continuation and bifurcation analysis of delay differential equations. InNumerical continuation methods for dynamical systems, Underst. Complex Syst., pages 359–399. Springer, Dordrecht, 2007
2007
-
[15]
Scheel and A
A. Scheel and A. Stevens. Reversible switching due to attraction and repulsion: clusters, gaps, sorting, and mixing.arXiv preprint, 2025
2025
-
[16]
Shalova and A
A. Shalova and A. Schlichting. Solutions of stationary mckean-vlasov equation on a high- dimensional sphere and other riemannian manifolds, 2025. 33
2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.