REVIEW 3 major objections 3 minor 2 cited by
Competing Risks: Impact on Risk Estimation and Algorithmic Fairness
T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims that treating competing risks as censoring systematically overestimates survival-model risk and widens group disparities.
desk verdict Solid core, overreaching abstract on disparity amplification, and I couldn't read the methods; still worth peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is the misclassification of a competing event as independent censoring. In a competing-risks survival model, each individual has a time to the target event and a time to a competing event; only the earlier is observed. Treating the competing event as independent censoring alters the risk set at later times, so the estimated survival curve no longer targets the cause-specific cumulative incidence but a curve that treats the competing event as if it only removed observation, not the possibility of the target event. The paper's error framework quantifies that difference in terms of the competing event's cumulative incidence, and uses it to derive group-level bias when competing-ev
What would settle it
Simulate survival data with a known target-event hazard and a known competing-event hazard, with groups that differ in competing-event incidence. Fit a standard survival model that codes competing events as censoring and compare its predicted cumulative incidence to the true value. If the model does not consistently overestimate risk, or if the overestimate is the same size for groups with different competing-event rates, the paper's central claim is false. A real-data version: in a cardiovascular cohort with long follow-up, compare predictions from the censoring-treated model against observed
Extended reading notes
Core claim
The central claim is that when a patient's follow-up ends because of a competing event—for example, death from a different cause before the outcome of interest—coding that patient as simply censored changes the quantity being estimated. The estimated probability of the target event is then not the true cumulative incidence but a larger object, because patients who can no longer experience the target event are still counted as still at risk. The paper derives a formal expression for this gap and shows it grows with the cumulative incidence of the competing event. Because competing-event rates differ across demographic groups, the overestimation is group-specific: the group with more competing
Load-bearing premise
The bias formula requires a correctly specified and identifiable model of the competing-risk process, including how the target event and competing event depend on each other; if that dependence is misspecified or cannot be identified from censored data, the claimed error estimates do not necessarily equal the true bias.
Editorial extensions
If this is right
- If competing events are treated as censoring, predicted cumulative incidence of the target event is systematically too high, so clinical thresholds or triage rules based on such models will flag too many people.
- The inflation is not constant across groups: groups with higher competing-event incidence receive larger upward bias, so measured disparities in risk can overstate true differences.
- Accounting for competing risks changes model choice: cause-specific or subdistribution hazard models should be used rather than standard censoring-based estimators.
- Model evaluation metrics like discrimination will also be affected, since ranking and calibration change when risks are inflated differently by group.
- In cardiovascular care specifically, ignoring non-cardiovascular death as a competing event can alter management decisions and disproportionately affect patients at highest risk of other-cause death.
Reading between the lines
- A testable extension: in other domains with competing events—job turnover where resignation and termination compete, recidivism where death or deportation competes—the same group-specific bias should appear whenever competing-event rates differ by group.
- An implication the authors leave implicit: fairness audits of survival models should measure competing-event incidence by group before interpreting risk-score disparities, since apparent disparity may be an artifact of censoring misspecification.
- Recalibrating a censoring-based model on observed rates may hide the bias if the calibration data itself treats competing events as censoring; validation should use actual event status after competing events.
- If competing and target events are dependent, the direction and size of the bias may be more complex, so practitioners should report sensitivity to that dependence.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies the common survival-analysis practice of treating competing events as ordinary censoring. From the abstract, the paper claims to (i) formalize the misspecification error caused by this practice, (ii) prove that it leads to systematic overestimation of the event risk, (iii) show that the error is group-specific and can amplify disparities relevant to algorithmic fairness, and (iv) support the findings with an empirical cardiovascular-management analysis. The mathematical core, as far as the abstract reveals, is the contrast between the censoring-model event probability P(T ≤ t) and the true cumulative incidence F(t) = P(T ≤ t, T < C), with the former dominating the latter. I was not able to read the equations and tables in the supplied full text because the text is severely encoding-corrupted; the assessment below therefore focuses on the abstract's logic and on the claims that are explicitly stated.
Significance. If the proposed framework is correct, the paper would provide a useful formalization of a well-known but often underappreciated bias in survival models, and it would connect that bias to fairness in risk prediction. The overestimation direction is standard and provable under the usual independent-censoring assumption; the potential contribution is the quantitative error decomposition and the explicit group-specific analysis. A careful statement of the conditions under which disparities are amplified would make the fairness claim more than a heuristic. The manuscript is potentially valuable for practitioners building survival models for clinical decision support, but its current abstract-level claims need to be tightened.
major comments (3)
- [Abstract, ¶1] The headline claim that misclassifying competing risks as censoring 'amplif[ies] disparities' is not a logical consequence of the overestimation inequality. For group g, let F_g(t)=P(T_g≤t,T_g<C_g) be the true cumulative incidence and R_g(t)=P(T_g≤t) be the misspecified estimator's target. The bias is Δ_g(t)=R_g(t)-F_g(t), and the disparity between groups A and B changes by Δ_A-Δ_B. Since Δ_g(t)=P(T_g≤t,C_g≤T_g), the sign of Δ_A-Δ_B depends on the joint distributions of (T_g,C_g). If the higher-target-risk group has a much lower competing-event hazard than the lower-target-risk group, Δ_A may be smaller than Δ_B, so the predicted-risk gap can shrink or even reverse. The abstract's later 'potentially exacerbating' is the correct level of claim; the opening 'amplifying disparities' needs either an explicit condition under which Δ_g is monotone in F_g or a softened formulation.
- [Abstract, ¶1 and Section 2 (definitions of T and C)] The unconditional claim that treating competing risks as censoring 'systematically overestimates' risk requires that the censoring-model estimator is consistent for S_T(t)=P(T>t). This holds under independent or at least noninformative censoring. If the competing event time C is dependent on the target event time T, treating C as censoring yields an inconsistent estimator whose limiting value can lie on either side of P(T≤t). The paper should explicitly state the regularity conditions under which the error formula is derived (e.g., C independent of T conditional on covariates), and should discuss the consequences when those conditions fail. As written in the abstract, the theorem is overgeneralized.
- [Abstract, ¶2 ('develop a framework to estimate this error')] The abstract promises a quantitative error estimator but does not state its identification conditions. The bias Δ_g(t) depends on the joint distribution of (T_g,C_g), which is not identified from right-censored data without assumptions (for example, cause-specific hazards with independent censoring, or a subdistribution hazard model). Unless the manuscript provides a precise model and consistency proof for the proposed error estimator, the empirical estimates in the cardiovascular application cannot be interpreted as estimates of the true bias. The authors should add the identification assumptions and, ideally, a small simulation verifying that the estimator recovers the true Δ under those assumptions.
minor comments (3)
- [Abstract, ¶1 and concluding paragraph] The wording alternates between 'amplifying disparities' and 'potentially exacerbating/ accentuating identity'; the abstract should be internally consistent and use the hedged form unless a theorem guarantees amplification.
- [General] The supplied full text is not readable: all equations and tables appear as encoding-corrupted characters. If this reflects the submitted manuscript rather than my rendering pipeline, the authors need to resubmit a machine-readable version. As a result, I could not independently verify the proof structure or the empirical tables.
- [Abstract, ¶2] The term 'competing risks' is used as a synonym for 'competing events'; define the target event and the competing event explicitly in the introduction to avoid ambiguity.
Circularity Check
No significant circularity found; the abstract-level derivation is self-contained and no fitted quantity is renamed as a prediction.
full rationale
The abstract describes a theoretical derivation: treating competing risks as censoring is compared against the true competing-risk structure, and the resulting error is quantified from that structure rather than from the target bias itself. The claimed conclusions—systematic overestimation of risk and group-specific error—are presented as consequences of the misspecification, and the fairness language is explicitly hedged with 'potentially exacerbating'/'potentially accentuating.' There is no equation in the abstract that defines the error in terms of the quantity it is meant to predict, and no parameter is fitted to data and then called a prediction. No self-citation is invoked as load-bearing evidence. Although the supplied full text is heavily garbled and prevents equation-level quotation, nothing in the readable abstract or surrounding context exhibits a reduction of a derived result to its own input. The skeptic's concern about the disparity-amplification claim is a question of scope and conditions, not circularity: even if the sign of disparity amplification depends on competing-hazard orderings, that is an overgeneralization risk, not a case of the derivation being equivalent to its assumptions. Therefore the honest finding is no significant circularity (score 0).
Assumptions & free parameters
assumptions (3)
- domain assumption The competing-risk process is represented by identifiable cause-specific or subdistribution hazards, with a specified dependence between the event of interest and the competing event.
- domain assumption Censoring other than the misclassified competing events is non-informative, i.e., independent of the event time given covariates.
- domain assumption Group-specific error comparisons rely on accurately estimated risk profiles for each demographic group in the cardiovascular data.
Cite this review
Pith. "Pith review of Competing Risks: Impact on Risk Estimation and Algorithmic Fairness." pith.science (2026). https://pith.science/paper/Z2K3IFEO
@misc{pith2026250805435,
author = {Pith},
title = {Pith review of: Competing Risks: Impact on Risk Estimation and Algorithmic Fairness},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z2K3IFEO}},
note = {Machine review of arXiv:2508.05435}
}
read the original abstract
Accurate time-to-event prediction is integral to decision-making, informing medical guidelines, hiring decisions, and resource allocation. Survival analysis, the quantitative framework used to model time-to-event data, accounts for patients who do not experience the event of interest during the study period, known as censored patients. However, many patients experience events that prevent the observation of the outcome of interest. These competing risks are often treated as censoring, a practice frequently overlooked due to a limited understanding of its consequences. Our work theoretically demonstrates why treating competing risks as censoring introduces substantial bias in survival estimates, leading to systematic overestimation of risk and, critically, amplifying disparities. First, we formalize the problem of misclassifying competing risks as censoring and quantify the resulting error in survival estimates. Specifically, we develop a framework to estimate this error and demonstrate the associated implications for predictive performance and algorithmic fairness. Furthermore, we examine how differing risk profiles across demographic groups lead to group-specific errors, potentially exacerbating existing disparities. Our findings, supported by an empirical analysis of cardiovascular management, demonstrate that ignoring competing risks disproportionately impacts the individuals most at risk of these events, potentially accentuating inequity. By quantifying the error and highlighting the fairness implications of the common practice of considering competing risks as censoring, our work provides a critical insight into the development of survival models: practitioners must account for competing risks to improve accuracy, reduce disparities in risk assessment, and better inform downstream decisions.
Forward citations
Cited by 2 Pith papers
-
The C-index illusion: discrimination without calibration in published survival models
Published survival models with high C-index scores can assign systematically wrong probabilities, and discrimination-only evaluation hides this failure.
-
A reproducible and extensible framework for benchmarking competing risks survival models
A reproducible benchmarking framework and a new SHAP extension for competing-risks survival models, with results showing simpler regression models often match deep learning.
Reference graph
Works this paper leans on
-
[1]
��������� ������ ������ �� ���� ���������� ��� ����������� �������� ������� ��������� �������� ����������� ��� ���� ���������� �� ���������� ��������� ����������������������������������� ����� ���� ������� ������� ���������� �� ���������� ��������� �������� ������������� ���������� �� �������� �� ���������������� ��������� ������� ����������� ������ �����...
work page Pith review arXiv 2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.