REVIEW 3 major objections 6 minor 35 references
Fairness Is More Than Algorithms: Racial Disparities in Time-to-Recidivism
T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Racial differences in time-to-recidivism appear within identical COMPAS risk groups after roughly seven months.
desk verdict A clean conditional-independence core and a useful temporal angle, but the headline causal claim is undermined by a censoring assumption that fails under the paper's own alternative. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the race-and-score-conditioned survival curve $S_d(t\mid m)=P[\tau>t \mid D=d, M=m]$, i.e., the probability that a defendant of race $d$ in COMPAS score group $m$ has not been rearrested by time $t$. Theorem 1 shows that under the null of no direct context effect, the counterfactual survival curve after an intervention on race equals this observational curve and hence is identical across races; Lemma 1 turns the contrapositive into a decision rule. The test itself is the non-parametric log-rank test applied to the observed time $T=\min\{\tau,\tau'\}$, with return to custody for non-criminal violations treated as independent right-censoring. The machinery gives a temporal answer: it identifies not only whether disparities exist but when they begin.
What would settle it
Re-analyze the COMPAS cohort with direct measures of post-release context, such as verified housing, employment, and social-support enrollment. If the post-seven-month survival curves for African-American and Caucasian low-risk defendants become indistinguishable after adjusting for those measures, the paper's contextual explanation passes; if they remain different, the claim that context drives the divergence would need revising. A second check is to re-estimate the model treating return to custody as a competing risk rather than independent censoring: if the racial divergence shrinks or disappears, the result hinges on that censoring assumption.
Extended reading notes
Core claim
The central claim is that within identical COMPAS risk-score groups, time-to-recidivism is racially different once follow-up extends past about seven months, and this difference signals a direct effect of non-algorithmic context on recidivism. The authors define counterfactual racial parity as equality of the survival curves $P^{\mathrm{do}(D=d)}[\tau > t \mid M=m]$ across races, and prove (Theorem 1) that under their causal DAG this counterfactual quantity collapses to the observable quantity $P[\tau>t \mid D=d, M=m]$ whenever context $U$ has no direct edge to $\tau$ or $\tau'$. Therefore a statistically significant difference between races in the same score group, assessed with a log-rank test on censored data, rejects the null that context plays no direct role. The COMPAS analysis finds exactly such a rejection for low-risk defendants after roughly seven months: African-American defendants show a faster drop in no-recidivism probability than Caucasian defendants with the same score. Medium- and high-risk groups show no significant divergence.
Load-bearing premise
The argument holds together only if the COMPAS score is a fully informative proxy for all individual characteristics relevant to recidivism, and race itself has no direct effect on recidivism or return-to-custody timing once score and context are fixed; if the score omits race-correlated traits, the observed divergence could appear even with no unobserved contextual factor at work.
Editorial extensions
If this is right
- Binary two-year recidivism audits can miss disparities that appear only after longer follow-up, so evaluation windows should be reported and varied.
- Fixing the risk algorithm will not remove all racial gaps in recidivism, since the gaps arise within groups receiving the same score.
- Low-risk minority defendants, who might otherwise seem least in need of intervention, are the group where long-term divergence is detected.
- Housing, employment, and social support become plausible intervention targets implied by the contextual account, rather than algorithmic tuning.
- The same survival test can be carried to other settings where a score and a censored time-to-event outcome coexist, such as loan repayment versus default.
Reading between the lines
- The seven-month threshold may track the typical duration of post-release supervision or the point at which short-term support programs end; a direct test would compare divergence times across jurisdictions with different supervision lengths.
- The analysis treats return to custody as independent censoring, but if return to custody is more frequent for one race, the log-rank comparison can be distorted; a competing-risk analysis is a natural robustness check the paper does not report.
- If datasets with direct measures of housing, employment, and social support were linked to COMPAS records, the socioeconomic mechanism could be tested head-on rather than inferred from the timing of divergence.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a multi-stage causal framework, represented by a DAG, for studying racial disparities in time-to-recidivism. It introduces a counterfactual notion of racial parity and shows that, under a null hypothesis in which unobserved context U does not directly affect time-to-recidivism or time-to-return-to-custody, the no-recidivism survival curves within a COMPAS risk-score group should be equal across races. The paper then uses a log-rank test on the observed time to rearrest or return to custody as an empirical test of this null. Applying the test to the COMPAS data, the authors find no significant racial differences for medium- and high-risk groups, but report that for low-risk defendants a significant divergence appears after roughly seven months, with African-American defendants showing faster declines in no-recidivism probability. The paper interprets this as evidence that non-algorithmic, potentially socioeconomic, factors affect recidivism trajectories.
Significance. If valid, the framework would be a useful contribution: it moves the algorithmic-fairness debate from binary recidivism outcomes to time-to-event data, explicitly incorporates censoring, and formalizes a contrapositive test for the presence of contextual effects. The paper is clearly written and its empirical strategy is transparent enough to be scrutinized. However, the central inferential claim is compromised by a load-bearing problem: the log-rank test is applied to the observed composite time T = min(tau, tau'), and the independent-censoring assumption under which T identifies the latent recidivism survival curves fails precisely under the alternative the test is designed to detect. In addition, the empirical headline rests on repeated testing without multiple-comparison correction and on a post hoc seven-month threshold. These issues prevent the paper from supporting its strongest conclusion that contextual factors affect recidivism timing, as opposed to affecting the joint recidivism/return-to-custody process.
major comments (3)
- [§3.4 and §4.2] The independent-censoring assumption stated in §3.4, tau independent of tau' given M=m, fails exactly under the alternative H1 that the test is designed to detect. Under H1, U directly affects both tau and tau', so tau and tau' share U as a common cause and are dependent conditional on M. Consequently, the observed time T=min(tau,tau') does not identify S_d(t|m)=P(tau>t|D=d,M=m), and a significant log-rank statistic on T can arise from a race-dependent difference in the return-to-custody process even when the true time-to-recidivism curves are identical. The abstract and §4.2 therefore overstate the finding: the empirical test only supports a difference in the joint recidivism/return-to-custody process, not 'recidivism patterns' or the claim that factors beyond the algorithmic risk assessment significantly influence recidivism patterns. Lemma 1 is stated in terms of P(tau>t|D,M), but Empirical Test 1 replaces this with a test on T; the replacement is valid only under an assumption that is false on the alternative of interest.
- [§4.2 and Appendix A] The significance claim is based on repeated log-rank tests computed over follow-up time without any multiple-comparison correction, and the 'approximately seven months' threshold is selected post hoc. Appendix A further narrows the significant effect to individual COMPAS scores 3 and 4, which multiplies the number of strata tested. Because no adjusted p-values are reported, the reader cannot assess whether the finding survives correction for the number of time points and risk-score strata. The analysis should either pre-specify the follow-up horizon and testing procedure or report family-wise error-rate-adjusted p-values, and it should report stratum-specific sample sizes and event counts. The appendix's caveat that the lack of effect for other scores 'might also be due to limited data' should be stated in the main text alongside the headline result.
- [§3.1 and §3.3] The causal interpretation relies on the structural assumption that race has no direct edge to tau or tau', together with the associated DAG specification. If perceived race directly shapes supervision intensity, charging decisions, or return-to-custody practices beyond the algorithmic score, then a within-M log-rank difference could reflect racial discrimination in the justice process rather than the effect of socioeconomic context U. The paper presents these assumptions without sensitivity analysis, and the data cannot distinguish between the two explanations. Given that the policy conclusion in §4.3 and §5 is specifically about structural socioeconomic factors, the paper should at least discuss how robust the claimed interpretation is to violations of the no-direct-race-edge assumption.
minor comments (6)
- [Figures 2-5] The captions describe subplots (a) survival curves for Caucasian defendants, (b) survival curves for African-American defendants, and (c) p-values, but each panel appears to contain one survival curve plot and one p-value plot per risk group; the captions and figures should be aligned.
- [Additional Key Words and Section 1] The 'Additional Key Words and Phrases' field contains a placeholder string ('Do, Not, Us, This, Code, Put, the, Correct, Terms, for, Your, Paper') and the ACM reference block still says 2018; these should be corrected before submission.
- [§3.4] The 'proportional hazards' bullet is stated as equality of derivatives of log survival curves, which is actually the condition of equal hazard rates; moreover, the log-rank test is a valid test of equality of hazard functions without requiring proportional hazards. The text should be corrected.
- [Figures 4-5] The x-axis is labeled 'Time (months)' but reaches values around 1000, which appear to be days; the units should be unified and clearly stated.
- [§3.4] The notation for observed and expected events is inconsistent: the text defines O_d,m,t and E_d,m,t but later refers to O_d,m and E_m,t without clearly connecting the aggregated quantities.
- [§3.3 and Hypothesis 1] Lemma 1 describes H0 as 'context U does not directly affect time-to-recidivism tau', while Hypothesis 1 defines H0 as no direct effect on both tau and tau'; these statements should be aligned to avoid ambiguity about what is being rejected.
Circularity Check
No significant circularity: the causal test is a contrapositive of a conditional-independence implication, and no fitted parameter is relabeled as a prediction.
full rationale
The derivation chain is self-contained. Theorem 1 derives the observable equality P[tau>t|D=d,M=m] = P[tau>t|M=m] from the DAG's no-direct-edge assumption plus H0 (removal of U->tau and U->tau'), using d-separation and do-calculus; Lemma 1 is the valid contrapositive of that implication, and Empirical Test 1 implements Lemma 1 with a standard log-rank test under explicitly stated assumptions. No parameters are fitted and then 'predicted': the seven-month divergence is a direct log-rank rejection of equality of survival curves, not the output of a model fitted to that same quantity. The only self-citation ([16], cited for 911-call disparities) is background and not load-bearing. A genuine caveat, but not a circularity, is that the independent-censoring assumption tau perpendicular tau' | M=m holds under H0 but can fail under H1; a rejection could in principle be driven by U affecting return-to-custody timing rather than recidivism timing. The paper's Section 5 limitation ('we do not have the privilege of directly verifying if the socioeconomic factors are indeed the additional sources of bias') appropriately concedes this. Because the test statistic is not defined in terms of the causal conclusion and the causal claim rests on stated structural assumptions rather than on a self-referential fit, the circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Follow-up divergence threshold =
~7 months
assumptions (5)
- domain assumption Causal DAG: context U affects race D and decision M and, under H0, has no edge to tau or tau'; D has no direct edge to tau or tau'.
- domain assumption M, the COMPAS risk category, is a fully informative proxy for all observable characteristics relevant to recidivism.
- domain assumption Independent censoring: tau is independent of tau' given M=m.
- domain assumption The COMPAS dataset's recorded outcomes align with the framework definitions: rearrest is the recidivism event, return to custody for non-criminal violations is censoring, and follow-up extends to roughly 36 months.
- standard math Standard log-rank conditions: independent recidivism events across individuals and correct event-time recording.
Cite this review
Pith. "Pith review of Fairness Is More Than Algorithms: Racial Disparities in Time-to-Recidivism." pith.science (2026). https://pith.science/paper/LSFAOFDI
@misc{pith2026250418629,
author = {Pith},
title = {Pith review of: Fairness Is More Than Algorithms: Racial Disparities in Time-to-Recidivism},
year = {2026},
howpublished = {\url{https://pith.science/paper/LSFAOFDI}},
note = {Machine review of arXiv:2504.18629}
}
read the original abstract
Racial disparities in recidivism remain a persistent challenge within the criminal justice system, increasingly exacerbated by the adoption of algorithmic risk assessment tools. Past works have primarily focused on bias induced by these tools, treating recidivism as a binary outcome. Limited attention has been given to non-algorithmic factors (including socioeconomic ones) in driving racial disparities from a systemic perspective. To that end, this work presents a multi-stage causal framework to investigate the advent and extent of disparities by considering time-to-recidivism rather than a simple binary outcome. The framework captures interactions among races, the algorithm, and contextual factors. This work introduces the notion of counterfactual racial disparity and offers a formal test using survival analysis that can be conducted with observational data to assess if differences in recidivism arise from algorithmic bias, contextual factors, or their interplay. In particular, it is formally established that if sufficient statistical evidence for differences across racial groups is observed, it would support rejecting the null hypothesis that non-algorithmic factors (including socioeconomic ones) do not affect recidivism. An empirical study applying this framework to the COMPAS dataset reveals that short-term recidivism patterns do not exhibit racial disparities when controlling for risk scores. However, statistically significant disparities emerge with longer follow-up periods, particularly for low-risk groups. This suggests that factors beyond algorithmic scores, possibly structural disparities in housing, employment, and social support, may accumulate and exacerbate recidivism risks over time. This underscores the need for policy interventions extending beyond algorithmic improvements to address broader influences on recidivism trajectories.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. 2016. Machine Bias. https://www.propublica.org/article/machine-bias-risk-assessments- in-criminal-sentencing
work page 2016
-
[2]
Michelle Bao, Angela Zhou, Samantha Zottola, Brian Brubach, Sarah Desmarais, Aaron Horowitz, Kristian Lum, and Suresh Venkatasubramanian
-
[3]
Chelsea Barabas, Madars Virza, Karthik Dinakar, Joichi Ito, and Jonathan Zittrain. 2018. Interventions over Predictions: Reframing the Ethical Debate for Actuarial Risk Assessment. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (Proceedings of Machine Learning Research, Vol. 81). PMLR, 62–76. https://proceedings.mlr.pres...
work page 2018
-
[4]
Richard Berk, Hoda Heidari, Shahin Jabbari, Michael Kearns, and Aaron Roth. 2021. Fairness in Criminal Justice Risk Assessments: The State of the Art. Sociological Methods & Research 50, 1 (2021), 3–44. https://doi.org/10.1177/0049124118782533
-
[5]
Marco Castillo, Sera Linardi, and Ragan Petrie. 2024. Recidivism and Barriers to Reintegration: A Field Experiment Encouraging Use of Reentry Support. https://ssrn.com/abstract=4891486 Manuscript submitted to ACM 14 Jessy Xinyi Han, Kristjan Greenewald, and Devavrat Shah
work page 2024
-
[6]
Alexandra Chouldechova. 2016. Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. arXiv:1610.07524 [stat.AP] https://arxiv.org/abs/1610.07524
arXiv 2016
-
[7]
Sam Corbett-Davies, Emma Pierson, Avi Feller, Sharad Goel, and Aziz Huq. 2017. Algorithmic Decision Making and the Cost of Fairness. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Halifax, NS, Canada) (KDD ’17). Association for Computing Machinery, New York, NY, USA, 797–806. https://doi.org/10.1145/...
arXiv 2017
-
[8]
Kennedy, and Alexandra Chouldechova
Amanda Coston, Alan Mishler, Edward H. Kennedy, and Alexandra Chouldechova. 2020. Counterfactual Risk Assessments, Evaluation, and Fairness. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (Barcelona Spain). ACM, New York, NY, USA, 582–593. https://doi.org/10.1145/3351095.3372851
arXiv 2020
Show all 35 references
-
[9]
William Dieterich, Christina Mendoza, and Tim Brennan. 2016. COMPAS Risk Scales: Demonstrating Accuracy Equity and Predictive Parity. https://www.documentcloud.org/documents/2998391-ProPublica-Commentary-Final-070616.html#document/p32/a310125
2016
-
[10]
Hyungrok Do, Yuxin Chang, Yoon Sang Cho, Padhraic Smyth, and Judy Zhong. 2023. Fair Survival Time Prediction via Mutual Information Minimization. In Proceedings of the 8th Machine Learning for Healthcare Conference (Proceedings of Machine Learning Research, Vol. 219) , Kaivaly...
2023
-
[11]
Flores, Kristin Bechtel, and Christopher T
Anthony W. Flores, Kristin Bechtel, and Christopher T. Lowenkamp. 2016. False Positives, False Negatives, and False Analyses: A Rejoinder to Machine Bias: There’s Software Used across the Country to Predict Future Criminals. And It’s Biased against Blacks. 80 (2016), 38. https...
2016
-
[12]
Northpointe Institute for Public Management. 1996. COMPAS [Computer software]
1996
-
[13]
Roland G. Fryer. 2019. An Empirical Analysis of Racial Differences in Police Use of Force. Journal of Political Economy 127, 3 (2019), 1210–1261. https://doi.org/10.1086/701423
2019 doi
-
[14]
Gary Goodley, Dominic Pearson, and Paul Morris. 2022. Predictors of Recidivism Following Release from Custody: A Meta-Analysis. Psychology, Crime & Law 28, 7 (2022), 703–729. https://doi.org/10.1080/1068316X.2021.1962866
2022
-
[15]
James Greiner and Donald B
D. James Greiner and Donald B. Rubin. 2011. CAUSAL EFFECTS OF PERCEIVED IMMUTABLE CHARACTERISTICS. The Review of Economics and Statistics 93, 3 (2011), 775–785. http://www.jstor.org/stable/23016076
2011
-
[16]
Craig Watkins, Christopher Winship, Fotini Christia, and Devavrat Shah
Jessy Xinyi Han, Andrew Cesare Miller, S. Craig Watkins, Christopher Winship, Fotini Christia, and Devavrat Shah. 2024. A Causal Framework to Evaluate Racial Bias in Law Enforcement Systems. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 7, 1 (2024), 562–572...
2024
-
[17]
Moritz Hardt, Eric Price, and Nathan Srebro. 2016. Equality of opportunity in supervised learning. In Proceedings of the 30th International Conference on Neural Information Processing Systems (Barcelona, Spain) (NIPS’16). Curran Associates Inc., Red Hook, NY, USA, 3323–3331
2016
-
[18]
Shu Hu and George H. Chen. 2024. Fairness in survival analysis with distributionally robust optimization. J. Mach. Learn. Res. 25, 1, Article 246 (Jan. 2024), 85 pages
2024
-
[19]
Huebner and Timothy S
Beth M. Huebner and Timothy S. Bynum. 2008. The Role of Race and Ethnicity in Parole Decisions. Criminology 46, 4 (2008), 907–938. https: //doi.org/10.1111/j.1745-9125.2008.00130.x
2008
-
[20]
Jacobs and Jennifer L
Leah A. Jacobs and Jennifer L. Skeem. 2021. Neighborhood Risk Factors for Recidivism: For Whom Do They Matter?American Journal of Community Psychology 67, 1-2 (2021), 103–115. https://doi.org/10.1002/ajcp.12463
2021 doi
-
[21]
Hyunzee Jung, Solveig Spjeldnes, and Hide Yamatani. 2010. Recidivism and Survival Time: Racial Disparity among Jail Ex-Inmates. Social Work Research 34, 3 (2010), 181–189. https://doi.org/10.1093/swr/34.3.181
2010 doi
-
[22]
Niki Kilbertus, Mateo Rojas-Carulla, Giambattista Parascandolo, Moritz Hardt, Dominik Janzing, and Bernhard Schölkopf. 2017. Avoiding Dis- crimination through Causal Reasoning. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long B...
2017
-
[23]
Anat Kimchi. 2019. Investigating the Assignment of Probation Conditions: Heterogeneity and the Role of Race and Ethnicity. Journal of Quantitative Criminology 35, 4 (2019), 715–745. https://doi.org/10.1007/s10940-018-9400-2
2019 doi
-
[24]
Matt Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. 2017. Counterfactual Fairness. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17). Curran Associates Inc., Red Hook, NY, USA, 4069–4079
2017
-
[25]
Jeff Larson, Surya Mattu, Lauren Kirchner, and Julia Angwin. 2016. How We Analyzed the COMPAS Recidivism Algorithm. https://www.propublica.org/article/how-we-analyzed-the-compas-recidivism-algorithm
2016
-
[26]
Kennedy, and Alexandra Chouldechova
Alan Mishler, Edward H. Kennedy, and Alexandra Chouldechova. 2021. Fairness in Risk Assessment Instruments: Post-Processing to Achieve Counterfactual Equalized Odds. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (Virtual Event Canada)....
2021
-
[27]
Okonofua, Kimia Saadatian, Joseph Ocampo, Michael Ruiz, and Perfecta Delgado Oxholm
Jason A. Okonofua, Kimia Saadatian, Joseph Ocampo, Michael Ruiz, and Perfecta Delgado Oxholm. 2021. A Scalable Empathic Supervision Intervention to Mitigate Recidivism from Probation and Parole. Proceedings of the National Academy of Sciences 118, 14 (2021), e2018036118. https...
2021 doi
-
[28]
Judea Pearl. 2014. Interpretation and identification of causal mediation.Psychological Methods 19, 4 (2014), 459–481. https://doi.org/10.1037/a0036434
2014 doi
-
[29]
Drago Plečko and Elias Bareinboim. 2024. Causal Fairness Analysis: A Causal Toolkit for Fair Machine Learning. Foundations and Trends ® in Machine Learning 17, 3 (2024), 304–589. https://doi.org/10.1561/2200000106 Manuscript submitted to ACM Fairness Is More Than Algorithms: R...
2024 doi
-
[30]
Marit Rehavi and Sonja B
M. Marit Rehavi and Sonja B. Starr. 2014. Racial Disparity in Federal Criminal Sentences. Journal of Political Economy 122, 6 (2014), 1320–1354. https://doi.org/10.1086/677255
2014 doi
-
[31]
Cynthia Rudin, Caroline Wang, and Beau Coker. 2020. The Age of Secrecy and Unfairness in Recidivism Prediction. Harvard Data Science Review 2, 1 (mar 31 2020). https://hdsr.mitpress.mit.edu/pub/7z10o269
2020
-
[32]
Maya Sen and Omar Wasow. 2016. Race as a Bundle of Sticks: Designs that Estimate Effects of Seemingly Immutable Characteristics. Annual Review of Political Science 19, 1 (2016), 499–522. https://doi.org/10.1146/annurev-polisci-032015-010015 _eprint: https://doi.org/10.1146/ann...
2016 doi
-
[33]
Sahil Verma and Julia Rubin. 2018. Fairness Definitions Explained. In Proceedings of the International Workshop on Software Fairness (Gothenburg, Sweden) (FairWare ’18). Association for Computing Machinery, New York, NY, USA, 1–7. https://doi.org/10.1145/3194770.3194776
2018
-
[34]
Bruce Western and Catherine Sirois. 2019. Racialized Re-entry: Labor Market Inequality After Incarceration. Social Forces 97, 4 (2019), 1517–1542. https://doi.org/10.1093/sf/soy096 Manuscript submitted to ACM 16 Jessy Xinyi Han, Kristjan Greenewald, and Devavrat Shah A Additio...
2019 doi
-
[2021]
CoRR abs/2106.05498 (2021)
It’s COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks. CoRR abs/2106.05498 (2021). arXiv:2106.05498 https://arxiv.org/abs/2106.05498
2021 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.