{"id":"d980c463-f9b2-4871-891f-0d910d039395","arxiv_id":"2508.14821","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Different R and Python packages compute the same C-index differently, and the paper demonstrates this on breast cancer and semi-synthetic data.","lead":"This paper claims that different R and Python implementations of the C-index produce different values for the same survival model, a 'multiverse' of results. It matters because C-index values are used to compare and select predictive models, and implementation-dependent variation can undermine reproducibility.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Supplied full text is a different paper (arXiv:2508.14827); the C-index audit's methods and results are absent, so the central multiverse claim cannot be assessed from this submission.","rationale":"The reader's verdict of UNVERDICTED is the correct disposition. The reader's stated weakest assumption concerns representativeness and correct execution of the audited software packages; that is a plausible concern once the audit is actually in hand. However, the more fundamental and immediately decisive concern is that the full text accompanying this submission is an unrelated paper. The absence of the C-index audit means there is no body of evidence to support the strongest claim, no way to probe the reader's weakest assumption, and no basis for moving to ACCEPT, CONDITIONAL, or REJECT. I therefore agree with the reader's verdict and recommend no change. I mark agreement as partial because my identified load-bearing concern is the manuscript-mismatch itself rather than the specific representativeness/execution assumption the reader highlighted, though both point to the same practical outcome: the central claim is not currently assessable.","tokens_in":25515,"tokens_out":3046,"duration_ms":37189,"concrete_test":"Fetch the current source of arXiv:2508.14821 from the arXiv API (export.arxiv.org) and diff its body against the supplied full text. If the body is not the C-index audit, the central claim is unverifiable from this submission; if the correct body is retrieved, then run the public GitHub code for the CindexMultiverse repository on the reported breast-cancer and semi-synthetic examples to check whether the claimed package discrepancies reproduce.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims a C-index multiverse among R and Python software, where seemingly equal implementations yield different results due to tie handling, censoring adjustments, and risk-summary choices. The load-bearing condition for that claim is an actual audit: a documented set of packages/versions, concrete tie and censoring rules, datasets, and reproducibility checks. The supplied full text is not that audit; it is the complete text of arXiv:2508.14827, 'Vacuum bubble and fissure formation in collective motion with competing attractive and repulsive forces'. No methods, tables, package list, or numerical results for the C-index study appear in the manuscript under review. The reader's weakest assumption (representativeness and correct execution of the selected packages) is a real downstream concern, but it cannot even be evaluated without the audit body. Per the review rule treating unusual inserted passages as in-scope evidence, the mismatch is itself decisive: an audit paper whose audit content is absent cannot support the inductive generalization that C-index outputs are non-unique across available software. This is not an accusation about the authors' underlying research; the correct paper may be perfectly sound, but this submission does not provide the evidence needed to assess it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This submission presents, in its abstract, an audit study of concordance-index (C-index) implementations across R and Python packages. The authors claim to demonstrate a 'C-index multiverse': for the same survival model and data, seemingly equivalent implementations yield different C-index values, and they attribute the variation to tie handling, censoring adjustment, and risk-score summarization choices. They further claim to illustrate the consequences on publicly available breast cancer data and semi-synthetic examples, and to offer guidelines and publicly available code. However, the full text supplied with the submission is not the audit paper. It is arXiv:2508.14827, 'Vacuum bubble and fissure formation in collective motion with competing attractive and repulsive forces', a PDE study of pattern formation in interacting particle systems. No section, equation, table, figure, or appendix of the claimed C-index study appears in the manuscript. The central claim and all supporting evidence are therefore absent from the submitted document.","tokens_in":25797,"tokens_out":2553,"duration_ms":28702,"significance":"The scientific question behind the abstract is genuinely significant. If an audit were carefully executed and documented, showing that C-index values vary across nominally equivalent software settings, it would have direct implications for reproducibility, model comparison, and reporting standards in time-to-event analysis. The abstract names concrete mechanisms (tie handling, censoring adjustment, risk-summary choices) and promises a practical guideline and public code, which are potentially valuable contributions. However, none of this can be evaluated from the submitted text. The manuscript contains no machine-checked proofs, no reproducible code listing, no package/version documentation, and no numerical experiments. The strengths that would justify publication are asserted in the abstract but are entirely absent from the manuscript body.","major_comments":[{"comment":"The submitted full text is the complete text of a different arXiv paper (2508.14827) on vacuum bubble and fissure formation in collective motion. It contains no part of the C-index audit described in the abstract: no methods, no package list, no datasets, no numerical results, and no discussion of survival analysis. This is not a missing appendix or a presentation issue; the claimed subject of the paper is absent, so the central claim is unsupported in the submitted document.","section":"Full Text"},{"comment":"The abstract's central claim to 'demonstrate the existence of a C-index multiverse' is not backed by any evidence in the manuscript. Specifically, there is no specification of the R/Python packages and versions tested, no formal statement of the tie-handling or censoring-adjustment rules compared, no description of the breast cancer or semi-synthetic datasets, and no numerical tables or figures showing differences in C-index outputs. The inductive generalization to 'available R and python software' therefore rests on no auditable evidence in this submission.","section":"Abstract"},{"comment":"The abstract states that 'All code is publicly available at www.github.com/BBolosSierra/CindexMultiverse,' but the full text contains no such URL or repository information, nor any code listing, package manifest, or version record. Even if the link is valid externally, the submitted manuscript provides no way to verify which implementations, options, and data were used, so the reproducibility claim cannot be checked.","section":"Abstract / Code Availability"}],"minor_comments":[],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract and the full text is so complete that this appears to be an administrative mix-up in the submission rather than a scientific disagreement. If the correct C-index audit manuscript exists, the authors should resubmit it; as it stands, there is nothing to review. The submitted full text is unrelated to the claimed topic, and no local revision can repair the missing content."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — the submission under this arXiv ID is not the C-index paper. The full text supplied is arXiv:2508.14827, a dynamical-systems paper on vacuum bubbles. No methods, package list, version numbers, tie rules, or numerical results for the C-index multiverse appear in the manuscript I was given. That is the load-bearing problem, and it is decisive for any evaluation of this submission.\n\nWhat I can judge is the abstract, which is coherent and describes a genuinely useful project: systematically testing R and Python survival-analysis packages to show that \"equal\" C-index implementations disagree due to tie handling, censoring adjustments, and risk summarization. That kind of cross-software audit is valuable; the field needs unified documentation of these pitfalls. The authors point to a public GitHub repo, which is good practice and makes the work independently checkable once the actual paper is attached. If the audit is executed as described, it would be a real contribution to reproducibility in survival modeling.\n\nThe absence of the body is not a minor issue. The abstract's inductive claim—\"a C-index multiverse among available R and Python software\"—rests entirely on the audit. With no audit in front of me, I cannot tell whether the package choices are representative, whether options are configured correctly, or whether results are reproducible. The reader flagged representativeness and correct execution as the weakest assumptions; those are legitimate, but they are downstream—I cannot even get to them without the methods. The mismatch is also concerning, though I would not read intent into it; it may be an upload error. Nevertheless, the paper as submitted is incomplete.\n\nWho is this for? Methodologically minded survival analysts and tool developers. A reader would get value from the finished version, but not from this submission. My recommendation: this should not go to peer review in its current form. The editor should return it to the authors to supply the correct full text. If the actual C-index audit is attached and matches the abstract's promises, it deserves a serious referee. As it stands, I cannot recommend engagement.","headline":"The submitted full text is a different paper, so the C-index multiverse claim cannot be assessed from this submission.","tokens_in":26225,"tokens_out":1934,"would_cite":false,"duration_ms":19721,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"As declared in its abstract, the paper claims that the C-index is not a unique number: seemingly equal R and Python implementations can disagree for the same model and data.","keywords":["concordance index","C-index","survival analysis","reproducibility","software comparison","tie handling","censoring","model evaluation"],"falsifier":"Run one fixed survival dataset and one fitted Cox model through several R and Python C-index functions, varying only tie-handling and censoring-adjustment options; if every implementation returns exactly the same concordance value, the multiverse claim is falsified for that configuration. A systematic sweep across packages and options would settle the claim. For this record, comparing the abstract's survival-analysis claims with the body's aggregation-equation theorems already shows the body text does not support them.","tokens_in":25468,"feed_emoji":"⚠️","tokens_out":6560,"duration_ms":65452,"temperature":0.7,"pith_summary":"The paper — as declared by its title and abstract — sets out to show that the C-index, a standard measure of how well a survival model ranks event times, is not a single well-defined number in practice. It argues that seemingly equivalent implementations in R and Python can return different concordance values for the same model and the same data, and that the differences trace to three concrete sources: how ties are handled, how the estimator adjusts for censoring, and how risk is summarized from a predicted survival distribution. If correct, this 'C-index multiverse' would mean reported discrimination scores are software-dependent, undermining reproducibility and fair comparisons across studies. The supplied full text, however, is not this paper: it is an unrelated mathematical study of vacuum bubble and fissure formation in aggregation models. The C-index claims therefore cannot be checked against the body text in this record.","feed_headline":"C-index scores differ across equal-looking software","feed_subtitle":"Tie rules, censoring adjustments, and risk summaries make the concordance index implementation-dependent.","key_machinery":"The central object is the 'C-index multiverse': the family of distinct numerical values that different implementations produce for what is nominally the same concordance index. The mechanisms that generate the multiplicity are tie handling, the choice and adjustment for censoring, and the non-standardised way risk is summarised from survival distributions. These choices are what make the metric implementation-dependent rather than a property of model and data alone.","core_discovery":"On its own terms, the paper claims that the concordance index is multiversal: fixing the data, the model, and the training/test split, the C-index value still depends on the software package chosen, because packages differ in tie-breaking conventions, in censoring adjustments (for example Harrell's, Uno's and Antolini's estimators), and in how they convert survival distributions into a scalar risk score. The paper reports numerical demonstrations on publicly available breast cancer data and semi-synthetic examples, across models from Cox proportional hazards to recent deep learning survival methods, showing that these implementation choices change the reported score. A caveat specific to thi","pith_inferences":["The same implementation-dependence likely affects other rank-based survival metrics, such as time-dependent AUC or Brier-score variants, since they share the same tie and censoring conventions; the paper does not claim this.","A natural testable extension would be a cross-package test suite that pins a reference dataset and reports every package's output under each documented tie and censoring option, making the boundaries of the multiverse explicit.","In this record the empirical demonstration cannot be verified, because the supplied body text belongs to a different paper; verifying the C-index multiverse requires the actual methods and experiments or a corrected manuscript."],"forward_implications":["Published C-index values should be reported together with the software, version, tie-breaking rule, and censoring adjustment used, otherwise the number is ambiguous.","Model comparisons that use different packages for different models can be biased by implementation choice rather than actual predictive performance.","Benchmarking studies and leader boards for survival models need a common, documented protocol for computing the C-index before rankings are meaningful.","Analysts need unified documentation and a checklist of pitfalls, which the paper positions as a guideline for navigating the multiverse."],"supporting_citations":[],"fun_headline_variants":["Same model, same data, different C-index","C-index depends on your software","Multiverse of C-index: software choice matters","Tie rules and censoring shift C-index scores","C-index not as universal as it seems"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim assumes the tested R/Python packages were representative of 'available software' and were invoked correctly; in this record it also assumes the abstract describes the actual paper, since the supplied full text is a different manuscript.","fun_headline_variants_meta":{"raw":{"variants":["Same model, same data, different C-index","C-index depends on your software","Multiverse of C-index: software choice matters","Tie rules and censoring shift C-index scores","C-index not as universal as it seems"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000708,"raw_usage":{"total_tokens":3020,"prompt_tokens":734,"completion_tokens":2286,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":2218}},"tokens_in":478,"tokens_out":2286,"duration_ms":16801,"temperature":1.0,"reasoning_tokens":2218,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:13:48.737864+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run one fixed survival dataset and one fitted Cox model through several R and Python C-index functions, varying only tie-handling and censoring-adjustment options; if every implementation returns exactly the same concordance value, the multiverse claim is falsified for that configuration. A systematic sweep across packages and options would settle the claim. For this record, comparing the abstract's survival-analysis claims with the body's aggregation-equation theorems already shows the body text does not support them.","supporting_citations":[],"review_version":1}