Pith. sign in

REVIEW 3 major objections 5 minor 35 references

The Dynamic and Endogenous Behavior of Re-Offense Risk: An Agent-Based Simulation Study of Treatment Allocation in Incarceration Diversion Programs

T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read No single risk rule for diversion treatment wins; the best priority flips with system settings.

desk verdict A transparent simulation testbed with a credible two-regime policy claim, but missing error bars and an unvalidated repeat-offense extrapolation keep the central result from being trustworthy yet. read the letter →

arxiv 2601.12441 v2 pith:DO56S4ED submitted 2026-01-18 cs.CY econ.GNq-fin.EC

classification cs.CYecon.GNq-fin.EC
keywords recidivismriskassessmenttreatmentallocationagent-basedsimulationsocialinteractionsprobationincarcerationdiversionpolicyevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that re-offense risk is not a static individual attribute but a dynamic quantity shaped by community feedback, so treatment allocation policies must be evaluated at the system level over long horizons. Using an agent-based simulation calibrated to U.S. probation data, it shows that prioritizing low-risk individuals performs better in the long run under low incarceration rates and long off-probation exposure, while prioritizing high-risk individuals is better in the short term or when incarceration truncates monitoring. The authors conclude that risk-based decision systems should be assessed as sociotechnical systems with long-term accountability. If correct, this undermines any universal risk-threshold rule for diversion decisions.

What carries the argument

The engine is a Cox proportional hazard model with a social-interaction term: an individual's hazard is scaled by γ0·μ, where μ is the community mean offense rate. Because μ aggregates everyone's offending, treatment decisions for one person change the environment for everyone else, creating an endogenous feedback loop. The agent-based simulation embeds this hazard in a full probation/off-probation/incarceration flow and evaluates policies under capacity constraints.

What would settle it

Re-estimate γ0 (the community feedback coefficient) and the baseline survival curve from repeated-offense and post-probation event data. If γ0 is negligible or changes sign, or if the baseline hazard for repeat events differs materially from the first-arrest curve, the two-regime ranking in Section 5.1 may invert or vanish.

Watch

Extended reading notes

Core claim

The central discovery is a two-regime structure in long-run policy performance (Section 5.1). Neither the low-risk nor the high-risk prioritization policy dominates across values of the incarceration probability δinc and the off-probation duration Foff. Low-risk priority wins when δinc is small (offenders stay in the community, so preventing risk accumulation matters more); high-risk priority wins when δinc is large (offenders are removed quickly, so immediate offense prevention is worth more). The age-first-low-risk policy sits between the two. The mechanism is that low-risk treatment prevents drift into high-risk states, whereas high-risk treatment delivers larger immediate absolute risk r

Load-bearing premise

The simulation assumes the baseline survival function and community-feedback coefficient estimated from time-to-first-arrest data (Baltimore, 1986-1989) continue to govern repeated offenses and off-probation dynamics over a 30,000-day horizon.

Editorial extensions

If this is right

  • If correct, risk-threshold tools alone cannot determine whom to treat; the monitoring horizon and incarceration response must be part of the decision.
  • Low-risk prioritization should be favored when long-term community exposure is high (low incarceration, long off-probation periods).
  • High-risk prioritization is preferable when monitoring periods are short or incarceration quickly removes offenders.
  • Policy rankings shrink as capacity grows or arrivals fall, since more resources per person dampen allocation differences.
  • Heterogeneous treatment effects can invert intuition: even when low-risk individuals benefit more, the low-risk policy may raise overall per-capita offenses because group-H offenses dominate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to test dynamic, learning-based allocation that adapts as μ and individual histories evolve; the two-regime structure suggests an adaptive policy could outperform both static thresholds.
  • The framework implies that policies evaluated only on short-term outcomes will systematically overstate high-risk prioritization, since long-term feedback is omitted.
  • A direct empirical test: if community-level offense rates are manipulated (e.g., through targeted enforcement or treatment in a neighborhood), the sign and size of γ0 should predict whether low-risk or high-risk allocation reduces total community offending.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper develops an agent-based simulation of probationers in an incarceration-diversion program, integrating an individual-level Cox hazard model with community-level feedback through the mean offense rate. The model is calibrated to 1986–1989 felony probation data from Baltimore and used to compare four treatment allocation policies (null, low-risk, high-risk, age-first–low-risk) under varying off-probation duration, incarceration probability, treatment effect, arrival rate, and capacity. The main claimed result is a two-regime structure: prioritizing low-risk individuals performs better in the long run under low incarceration probability and long off-probation exposure, while prioritizing high-risk individuals is better in the short run or when incarceration probability is high. The paper argues this shows risk-based diversion policies should be evaluated as sociotechnical systems rather than static prediction tools.

Significance. If the results are robust, the paper makes a useful conceptual contribution: it embeds risk assessment in an endogenous, system-level feedback loop and demonstrates that policy rankings can depend on temporal and systemic parameters, not just on individual risk scores. The simulation framework is clearly described with pseudo-code, and the calibration uses public ICPSR data and a published hazard model, which aids reproducibility. The insight that low-risk prioritization can outperform high-risk prioritization over long horizons under certain monitoring regimes is non-obvious and policy-relevant. However, the central claim rests on two load-bearing assumptions that are not yet adequately supported: the transfer of a first-arrest survival model to repeated offenses over a 30,000-day horizon, and the statistical reliability of the small policy differences shown in the headline figures.

major comments (3)
  1. [§5.1, Fig. 2 and Fig. 3] The central two-regime claim is presented without uncertainty quantification. The plotted differences from null policy are on the order of 0.01 offenses per capita, and the policy curves cross near δinc in the range 0.01–0.06 where differences are close to zero. Table 6 later reports '~0' for several entries, indicating that confidence intervals were computed, but the main figures omit them. Without error bars or a formal test that low-risk and high-risk long-run offense rates are statistically distinguishable at each (Foff, δinc), the regime switching point could be a Monte Carlo artifact. Please add uncertainty bands or significance tests, especially near the claimed crossover.
  2. [§4.2, Eq. (10) and Procedure 3] The baseline survival function S0 and the community-feedback coefficient γ0 are estimated from time-to-first-arrest data (1986–1989 Baltimore), yet they are used to generate every repeated offense and off-probation event over a 30,000-day horizon. This extrapolation from first-event survival to recurrent-event, long-run dynamics is not validated. If repeat-event hazards differ (e.g., due to desistance, aging, or changing offense mix), the two-regime ranking could change. Please provide sensitivity analyses with alternative baseline hazards, or validate against a dataset with multiple offense events per individual.
  3. [§5.2 and Table 6] The conclusion in §5.2 that heterogeneous 'lower-better' treatment effects can make the low-risk policy worse than the high-risk policy is interesting but relies on per-group offense rates reported in Figure 4(b) without confidence intervals. Since this result is used to explain a counterintuitive policy ranking, it should also be accompanied by uncertainty quantification. The current presentation does not establish that the group-level differences are not noise.
minor comments (5)
  1. [Throughout] Typos: 'componets' (p. 6), 'inlcude' (p. 11), 'off-probabtion' (p. 13), 'prohabtion' (p. 8), 'v.s.' should be 'vs.'.
  2. [Table 3] The range '[1, 3,768]' is confusingly formatted; likely '[1, 3,768]' should be '[1, 3768]' or something similar.
  3. [Footnote 2] The citation to 'Sirakaya [30]' for the Bayesian model averaging appears to refer to the dissertation, but the JASA article [31] is the more appropriate citation for the published model. Please verify.
  4. [Table 4] Table 4 mixes the newly fitted parameters (α0, θ1) with coefficients imported from the original Sirakaya model. Please label the columns or add a note clarifying which rows are estimated in this paper and which are taken from prior work.
  5. [§3.3] The 'Age-first–low-risk' policy description says 'lowest risk score h_i' but h_i can be negative; 'lowest' is fine but consider defining the ordering explicitly to avoid ambiguity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central two-regime policy finding is emergent from the calibrated simulation, not equivalent to its fitted inputs.

full rationale

The derivation chain is: fit a Cox proportional-hazard model and baseline survival S0 to first-arrest durations; embed that hazard in an agent-based simulation with a community-feedback term μ and treatment effect β treated as a tunable parameter; then run treatment-allocation policies and read off offense/incarceration metrics. The central claim (Section 5.1) that low-risk and high-risk policies alternate in dominance as δ_inc and F_off vary is not a restatement of any fitted coefficient or policy definition: it emerges from the dynamic interplay of offense-count accumulation, μ feedback, capacity constraints, and incarceration removal. The short-term advantage of the high-risk policy is acknowledged to follow from the homogeneous multiplicative treatment effect in Eq. (10), but the paper presents this as a model implication rather than as an independent empirical prediction, and the long-run regime structure is not forced by that algebraic fact. References [18] and [35] are self-citations but are peripheral literature citations, not load-bearing evidence. The reuse of Sirakaya's h0_i/γ0 is an explicit external calibration choice, not a hidden reduction; no equation in the paper reduces the reported policy rankings to the fitted parameters by construction. Concerns about extrapolating S0 to repeat events or about missing uncertainty bars in Figures 2 and 3 are validity/robustness issues, not circularity.

Assumptions & free parameters 9 free parameters · 6 assumptions · 0 invented entities

The model's core is calibrated to prior data, but the long-run policy conclusion rests on several tuned parameters and domain assumptions, especially the reuse of first-arrest survival dynamics for repeated offenses.

free parameters (9)
  • Treatment effect β = 0.342 baseline; varied {0.105,0.511}; heterogeneous variants
    Not estimated from data; tuned from literature. Policy rankings depend on it (§4.2, Table 5).
  • Incarceration probability δinc = 0.048 baseline; range [0,0.12]
    Tuning parameter; the two-regime switch point is defined in terms of it (§5.1).
  • Off-probation duration Foff = Exp(1/1000) baseline; range Exp(1/365)-Exp(1/2000)
    Tuning parameter; shifts the policy switch point (§5.1).
  • Arrival rate Farrival = Exp(1/5) baseline; range Exp(1/2)-Exp(1/20)
    Tuning parameter; affects per-capita resource scarcity (§5.3).
  • Treatment capacity C = 80 baseline; 100 and 200 in scenarios
    Tuning parameter; capacity constraints define the allocation problem (§5.3).
  • Episode length and horizon TE, Tmax = TE=100 days, Tmax=30,000 days
    Arbitrary choices; long-term metrics depend on equilibrium after 30,000 days (§4.2).
  • α0, θ1 = α0=0.7903, θ1=0.1883
    Fitted to ICPSR data via Cox/Breslow (§4.2, Table 4).
  • γ0 = 0.045
    Coefficient for % rearrested from Sirakaya's prior fitted model, absorbed into h0_i; not re-estimated here (Table 4).
  • Maximum returns Rinc = 30
    Tuning parameter limiting reentry (§3.1).
assumptions (6)
  • standard math Cox proportional hazards structure for inter-offense times (Eq. 1)
    Assumes baseline hazard and multiplicative scaling; standard survival model.
  • domain assumption Community mean offense rate μ proxies criminogenic peer influence and enters log-hazard linearly with positive γ0 (Eq. 5)
    The endogenous feedback mechanism; not derived from first principles in this paper.
  • ad hoc to paper Baseline survival S0 from time-to-first-arrest applies to all subsequent offenses and off-probation events (Procedure 3, Eq. 10)
    No validation that repeat-event hazards follow the same baseline.
  • domain assumption System is stationary and reaches equilibrium within 30,000 days; long-term metrics measured at equilibrium (Section 4.2)
    The 82-year horizon far exceeds the observation window; equilibrium is a simulation artifact assumption.
  • domain assumption 1986-1989 Baltimore City data and parameter ranges inform current U.S. probation policy (Section 4)
    External validity is assumed; no modern validation is provided.
  • domain assumption Treatment decisions are made only at arrival; no dynamic reassignment (Section 3.2)
    Simplifies the model but excludes adaptive policies.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Dynamic and Endogenous Behavior of Re-Offense Risk: An Agent-Based Simulation Study of Treatment Allocation in Incarceration Diversion Programs." pith.science (2026). https://pith.science/paper/DO56S4ED

@misc{pith2026260112441,
  author       = {Pith},
  title        = {Pith review of: The Dynamic and Endogenous Behavior of Re-Offense Risk: An Agent-Based Simulation Study of Treatment Allocation in Incarceration Diversion Programs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DO56S4ED}},
  note         = {Machine review of arXiv:2601.12441}
}
read the original abstract

Incarceration-diversion treatment programs aim to improve societal reintegration and reduce recidivism, but limited capacity forces policymakers to make prioritization decisions that often rely on risk assessment tools. While predictive, these tools typically treat risk as a static, individual attribute, which overlooks how risk evolves over time and how treatment decisions shape outcomes through social interactions. In this paper, we develop a new framework that models reoffending risk as a human-system interaction, linking individual behavior with system-level dynamics and endogenous community feedback. Using an agent-based simulation calibrated to U.S. probation data, we evaluate treatment allocation policies under different capacity constraints and incarceration settings. Our results show that no single prioritization policy dominates. Instead, policy effectiveness depends on temporal windows and system parameters: prioritizing low-risk individuals performs better when long-term trajectories matter, while prioritizing high-risk individuals becomes more effective in the short term or when incarceration leads to shorter monitoring periods. These findings highlight the need to evaluate risk-based decision systems as sociotechnical systems with long-term accountability, rather than as isolated predictive tools.

Figures

Figures reproduced from arXiv: 2601.12441 by the authors.

Figure 1
Figure 1. Flow chart of a probationer’s journey through the system. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Offense and incarceration per capita v.s. [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Offenses per capita across different treatment effects v.s. [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Impact of heterogeneous treatment effects on policy performance. [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

35 extracted references · 9 canonical work pages

  1. [1]

    Ronald L. Akers. 2011. Rational choice, deterrence, and social learning theory in criminology: the path not taken. InCrime Opportunity Theories. Routledge

  2. [2]

    Andrews, James Bonta, and J

    D.A. Andrews, James Bonta, and J. Stephen Wormith. 2011. The risk-need-responsivity (RNR) model: does adding the good lives model contribute to effective crime prevention?Criminal Justice and Behavior38, 7 (July 2011), 735–755. https://doi.org/10.1177/0093854811406356

  3. [3]

    D. A. Andrews, James Bonta, and R. D. Hoge. 1990. Classification for effective rehabilitation: rediscovering psychology.Criminal Justice and Behavior17, 1 (March 1990), 19–52. https://doi.org/10.1177/0093854890017001004

  4. [4]

    Michael Applegarth, Raven A

    D. Michael Applegarth, Raven A. Lewis, and Rachael M. Rief. 2023. Imperfect tools: A research note on developing, applying, and increasing under- standing of criminal justice risk assessments.Criminal Justice Policy Review34, 4 (Aug. 2023), 319–336. https://doi.org/10.1177/08874034231180505

  5. [5]

    Emily Bazelon. 2005. Sentencing by the numbers.The New York Times(Jan. 2005). https://www.nytimes.com/2005/01/02/magazine/sentencing-by- the-numbers.html

  6. [6]

    Richard Berk, Hoda Heidari, Shahin Jabbari, Michael Kearns, and Aaron Roth. 2021. Fairness in criminal justice risk assessments: The state of the art.Sociological Methods & Research50, 1 (Feb. 2021), 3–44. https://doi.org/10.1177/0049124118782533

  7. [7]

    Career Criminals

    Alfred Blumstein. 1986.Criminal Careers and "Career Criminals. " Volume I. National Academy Press, Washington, District of Columbia

  8. [8]

    N. E. Breslow. 1975. Analysis of survival data under the proportional hazards model.International Statistical Review / Revue Internationale de Statistique43, 1 (1975), 45–57. jstor:1402659 https://www.jstor.org/stable/1402659

Show all 35 references
  1. [9]

    Alexandra Chouldechova and Aaron Roth. 2020. A snapshot of the frontiers of fairness in machine learning.Commun. ACM63, 5 (April 2020), 82–89. https://doi.org/10.1145/3376898

  2. [10]

    D. R. Cox. 1972. Regression models and life-tables.Journal of the Royal Statistical Society: Series B (Methodological)34, 2 (1972), 187–202. https://onlinelibrary.wiley.com/doi/abs/10.1111/j.2517-6161.1972.tb00899.x

  3. [11]

    Durlauf and Yannis M

    Steven N. Durlauf and Yannis M. Ioannides. 2010. Social interactions.Annual Review of Economics2, Volume 2, 2010 (Sept. 2010), 451–478. https://www.annualreviews.org/content/journals/10.1146/annurev.economics.050708.143312

  4. [12]

    Federal Bureau of Investigation. 2019. Crime in the U.S., 2019. https://ucr.fbi.gov/crime-in-the-u.s/2019/crime-in-the-u.s.-2019/tables/table- 25/table-25.xls

  5. [13]

    Paul Gendreau, Tracy Little, and Claire Goggin. 1996. A meta-analysis of the predictors of adult offender recidivism: what works!Criminology34, 4 (1996), 575–608. https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1745-9125.1996.tb01220.x

  6. [14]

    James E. Gentle. 2003.Random Number Generation and Monte Carlo Methods(2nd ed ed.). Springer, New York

  7. [15]

    Harding, Jeffrey D

    David J. Harding, Jeffrey D. Morenoff, Anh P. Nguyen, and Shawn D. Bushway. 2017. Short- and long-term effects of imprisonment on future felony convictions and prison admissions.Proceedings of the National Academy of Sciences114, 42 (Oct. 2017), 11103–11108. https://www.pnas.o...

  8. [16]

    Jon Kleinberg, Himabindu Lakkaraju, Jure Leskovec, Jens Ludwig, and Sendhil Mullainathan. 2018. Human decisions and machine predictions.The Quarterly Journal of Economics133, 1 (Feb. 2018), 237–293. https://doi.org/10.1093/qje/qjx032

  9. [17]

    Mario Koran, Justin Mayo, and Taylor Glascock. 2024. 10 guards, 900 inmates and the dire results of warnings ignored.The New York Times(Feb. 2024). https://www.nytimes.com/2024/02/02/us/wi-prison-staffing-shortage.html

  10. [18]

    Bingxuan Li, Pengyi Shi, and Amy Ward. 2024. Latent feature mining for predictive model enhancement with large language models. arXiv:2410.04347 http://arxiv.org/abs/2410.04347

  11. [19]

    Lipsey, Nana A

    Mark W. Lipsey, Nana A. Landenberger, and Sandra J. Wilson. 2007. Effects of cognitive-behavioral programs for criminal offenders.Campbell Systematic Reviews3, 1 (2007), 1–27. https://onlinelibrary.wiley.com/doi/abs/10.4073/csr.2007.6

  12. [20]

    Lance Lochner. 2007. Individual perceptions of the criminal justice system.American Economic Review97, 1 (March 2007), 444–460. https: //www.aeaweb.org/articles?id=10.1257/aer.97.1.444

  13. [21]

    Loughran, Ray Paternoster, Aaron Chalfin, and Theodore Wilson

    Thomas A. Loughran, Ray Paternoster, Aaron Chalfin, and Theodore Wilson. 2016. Can rational choice be considered a general theory of crime? Evidence from individual-level panel data.Criminology54, 1 (2016), 86–112. https://onlinelibrary.wiley.com/doi/abs/10.1111/1745-9125.12097

  14. [22]

    Lowenkamp, Edward J

    Christopher T. Lowenkamp, Edward J. Latessa, and Alexander M. Holsinger. 2006. The risk principle in action: what have we learned from 13,676 offenders and 97 correctional programs?Crime & Delinquency52, 1 (Jan. 2006), 77–93. https://doi.org/10.1177/0011128705281747 16 Zhang, ...

  15. [23]

    Matthew Makarios, Kimberly Gentry Sperber, and Edward J. Latessa. 2014. Treatment dosage and the risk principle: a refinement and extension. Journal of Offender Rehabilitation53, 5 (July 2014), 334–350. https://doi.org/10.1080/10509674.2014.922157

  16. [24]

    Michael D. Maltz. 1984.Recidivism. Academic Press, Orlando

  17. [25]

    Michael D. Maltz. 1996. From poisson to the present: Applying operations research to problems of crime and justice.Journal of Quantitative Criminology12, 1 (March 1996), 3–61. https://doi.org/10.1007/BF02354470

  18. [26]

    Charles F. Manski. 2000. Economic analysis of social interactions.Journal of Economic Perspectives14, 3 (Sept. 2000), 115–136. https://www.aeaweb. org/articles?id=10.1257/jep.14.3.115

  19. [27]

    United States Bureau of the Census. 2009. County and City Data Book [United States], 1988. https://www.icpsr.umich.edu/web/ICPSR/studies/9251

  20. [28]

    Joan Petersilia. 1997. Probation in the united states.Crime and Justice22 (Jan. 1997), 149–200. https://www.journals.uchicago.edu/doi/abs/10.1086/ 449262

  21. [29]

    Brian A Reaves. 2009. Felony defendants in large urban counties, 2009 - statistical tables. (2009)

  22. [30]

    Sirakaya

    S. Sirakaya. 2003.Essays on social interactions and evolution in economics. Ph. D. Dissertation. University of Wisconsin-Madison

  23. [31]

    Sirakaya

    S. Sirakaya. 2006. Recidivism and social interactions.J. Amer. Statist. Assoc.101, 475 (2006), 863–877

  24. [32]

    Sonja B. Starr. 2014. Opinion | Sentencing, by the numbers.The New York Times(Aug. 2014). https://www.nytimes.com/2014/08/11/opinion/ sentencing-by-the-numbers.html

  25. [33]

    United States Department of Justice Office of Justice Programs Bureau of Justice Statistics. 2005. Recidivism of felons on probation, 1986-1989: [united states]. https://www.icpsr.umich.edu/web/NACJD/studies/9574#

  26. [34]

    Glenn Thrush. 2024. Staffing crisis at federal prisons highlighted in oregon.The New York Times(May 2024). https://www.nytimes.com/2024/05/22/ us/politics/oregon-prison-staffing-shortage.html

  27. [35]

    Zhiqiang Zhang, Pengyi Shi, and Amy Ward. 2025. Admission decisions under imperfect classification: An application in criminal justice. Social Science Research Network:5214197 https://papers.ssrn.com/abstract=5214197

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.