{"id":"8afa61a5-36df-44d1-ae2e-8fb09f079690","arxiv_id":"2601.12441","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"Treatment-priority effectiveness in probation is regime-dependent: low-risk prioritization wins long-term under lenient incarceration, high-risk wins short-term or under strict incarceration.","lead":"This paper simulates how different treatment-priority rules for probation programs affect re-offense over time, including social spillovers through community offense rates. It finds no policy always wins: prioritizing low-risk people works better over long horizons, while prioritizing high-risk people works better short-term or when incarceration is stricter.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Main simulation figures lack uncertainty quantification, so the claimed two-regime policy crossover in §5.1 may be sampling noise rather than a real effect.","rationale":"The paper's central claim is an empirical finding about policy ranking, so statistical uncertainty is foundational. The reader's weakest assumption—portability of the first-arrest survival function—affects the quantitative validity of the whole simulation, but the two-regime pattern could survive changes to the baseline if the relative ordering is driven by structural feedback. However, if the differences between policies are within Monte Carlo noise, there is no ordering to explain. The paper's own use of '~0' in Table 6 indicates that some effects are not significant, making the absence of error bars in Figures 2 and 3 especially conspicuous. This concern does not require rejecting the paper; it is a fixable reporting gap. Since the reader already assigned a conditional verdict, my analysis reinforces that conditionality without moving to a stronger verdict.","tokens_in":169,"tokens_out":10901,"duration_ms":134673,"concrete_test":"For the parameter grid in Fig. 2 (Foff ∈ {Exp(1/365), Exp(1/730), Exp(1/1000), Exp(1/2000)}; δinc ∈ {0.02, 0.04, 0.06, 0.08, 0.10, 0.12}; β=0.342; C=80), run at least 100 independent replications per policy and compute 95% confidence intervals for the difference in long-run per-capita offenses (high-risk minus low-risk), using the same equilibrium period as the paper. If the intervals include zero at both sides of the claimed crossover for any Foff, the two-regime claim is not established; if the intervals are disjoint with opposite signs across the boundary, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Figures 2 and 3 present per-policy offense/incarceration differences as single curves with no error bars or replication counts. The central claim in §5.1 is that low-risk and high-risk policies 'alternate in dominance' as δinc and Foff vary. With y-axis increments of ~0.01 (long-term offenses), the policy differences at any one δinc are comparable in magnitude to typical Monte Carlo noise in an agent-based simulation with 9,374 initial agents and 30,000-day horizons. Near the claimed switching points (δinc≈0.01–0.06), the curves cross within a range where the plotted differences are close to zero. Table 6 acknowledges that some entries are statistically indistinguishable ('~0' indicates overlapping confidence intervals), showing the authors did compute uncertainties, but the headline figures omit them. Without a formal test of whether high-risk and low-risk long-run offense rates differ at each (Foff, δinc), the regime boundary could be an artifact of random seed selection. This concern is logically prior to the reader's external-validity objection: if the policy differences are not statistically distinguishable, the two-regime structure is unsupported even if the survival-time extrapolation happens to be valid.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops an agent-based simulation of probationers in an incarceration-diversion program, integrating an individual-level Cox hazard model with community-level feedback through the mean offense rate. The model is calibrated to 1986–1989 felony probation data from Baltimore and used to compare four treatment allocation policies (null, low-risk, high-risk, age-first–low-risk) under varying off-probation duration, incarceration probability, treatment effect, arrival rate, and capacity. The main claimed result is a two-regime structure: prioritizing low-risk individuals performs better in the long run under low incarceration probability and long off-probation exposure, while prioritizing high-risk individuals is better in the short run or when incarceration probability is high. The paper argues this shows risk-based diversion policies should be evaluated as sociotechnical systems rather than static prediction tools.","tokens_in":7,"tokens_out":3319,"duration_ms":101469,"significance":"If the results are robust, the paper makes a useful conceptual contribution: it embeds risk assessment in an endogenous, system-level feedback loop and demonstrates that policy rankings can depend on temporal and systemic parameters, not just on individual risk scores. The simulation framework is clearly described with pseudo-code, and the calibration uses public ICPSR data and a published hazard model, which aids reproducibility. The insight that low-risk prioritization can outperform high-risk prioritization over long horizons under certain monitoring regimes is non-obvious and policy-relevant. However, the central claim rests on two load-bearing assumptions that are not yet adequately supported: the transfer of a first-arrest survival model to repeated offenses over a 30,000-day horizon, and the statistical reliability of the small policy differences shown in the headline figures.","major_comments":[{"comment":"The central two-regime claim is presented without uncertainty quantification. The plotted differences from null policy are on the order of 0.01 offenses per capita, and the policy curves cross near δinc in the range 0.01–0.06 where differences are close to zero. Table 6 later reports '~0' for several entries, indicating that confidence intervals were computed, but the main figures omit them. Without error bars or a formal test that low-risk and high-risk long-run offense rates are statistically distinguishable at each (Foff, δinc), the regime switching point could be a Monte Carlo artifact. Please add uncertainty bands or significance tests, especially near the claimed crossover.","section":"§5.1, Fig. 2 and Fig. 3"},{"comment":"The baseline survival function S0 and the community-feedback coefficient γ0 are estimated from time-to-first-arrest data (1986–1989 Baltimore), yet they are used to generate every repeated offense and off-probation event over a 30,000-day horizon. This extrapolation from first-event survival to recurrent-event, long-run dynamics is not validated. If repeat-event hazards differ (e.g., due to desistance, aging, or changing offense mix), the two-regime ranking could change. Please provide sensitivity analyses with alternative baseline hazards, or validate against a dataset with multiple offense events per individual.","section":"§4.2, Eq. (10) and Procedure 3"},{"comment":"The conclusion in §5.2 that heterogeneous 'lower-better' treatment effects can make the low-risk policy worse than the high-risk policy is interesting but relies on per-group offense rates reported in Figure 4(b) without confidence intervals. Since this result is used to explain a counterintuitive policy ranking, it should also be accompanied by uncertainty quantification. The current presentation does not establish that the group-level differences are not noise.","section":"§5.2 and Table 6"}],"minor_comments":[{"comment":"Typos: 'componets' (p. 6), 'inlcude' (p. 11), 'off-probabtion' (p. 13), 'prohabtion' (p. 8), 'v.s.' should be 'vs.'.","section":"Throughout"},{"comment":"The range '[1, 3,768]' is confusingly formatted; likely '[1, 3,768]' should be '[1, 3768]' or something similar.","section":"Table 3"},{"comment":"The citation to 'Sirakaya [30]' for the Bayesian model averaging appears to refer to the dissertation, but the JASA article [31] is the more appropriate citation for the published model. Please verify.","section":"Footnote 2"},{"comment":"Table 4 mixes the newly fitted parameters (α0, θ1) with coefficients imported from the original Sirakaya model. Please label the columns or add a note clarifying which rows are estimated in this paper and which are taken from prior work.","section":"Table 4"},{"comment":"The 'Age-first–low-risk' policy description says 'lowest risk score h_i' but h_i can be negative; 'lowest' is fine but consider defining the ordering explicitly to avoid ambiguity.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely topic and the simulation is methodical, but the main regime-switching result needs to be made statistically credible. I would recommend requesting uncertainty quantification in the headline figures and sensitivity analysis for the survival-model extrapolation before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new claim here is the two-regime structure: over long horizons and low incarceration stringency, treating low-risk probationers first beats treating high-risk ones, while high-risk prioritization wins when monitoring is short or incarceration removes people quickly. That reframes risk-tool evaluation from static individual prediction to system-level accountability, and it deserves a careful look. The agent-based testbed itself is transparent, the calibration is described well enough to reproduce in principle, and the authors are appropriately cautious in their language. They also know how to quantify uncertainty when they want to—Table 6 reports confidence intervals and flags statistically indistinguishable entries.\n\nThe soft spots are real. First, the headline figures (Figures 2 and 3) showing the two-regime crossover have no error bars or replication counts. Near the switching points the policy differences are on the order of 0.01 on the vertical axis, which is exactly where Monte Carlo noise in a 9,000-agent, 30,000-day simulation could be. Without error bars you cannot tell whether the crossover is a real effect or a particular seed. Table 6 shows they have the machinery to fix this; the central figures just omit it.\n\nSecond, the survival model that drives every re-offense event is estimated from time-to-first-arrest in 1986-1989 Baltimore probation data and then applied to repeated offenses over a 30,000-day horizon. Nothing validates that first-arrest dynamics govern repeat-event dynamics across decades. The community-feedback coefficient gamma0 and the baseline survival function S0 are load-bearing, and the portability assumption is untested. If it fails, the two-regime ranking could shift. The treatment effect and incarceration probability are tuning parameters, which is acceptable for a sensitivity study, but the central claim rests on that extrapolation.\n\nThere is also a mild circularity in evaluating long-term outcomes on the same model that generated them, but that is common in simulation work and not a red flag by itself.\n\nThis is a serious, well-written paper. It deserves peer review, not desk rejection, and the authors should be asked to release code and data, add uncertainty quantification to the main figures, and test or at least discuss the repeat-offense extrapolation. I'd bring it to a reading group, but I wouldn't cite the two-regime result as established until those issues are resolved.","headline":"A transparent simulation testbed with a credible two-regime policy claim, but missing error bars and an unvalidated repeat-offense extrapolation keep the central result from being trustworthy yet.","tokens_in":15844,"tokens_out":2805,"would_cite":false,"duration_ms":28892,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No single risk rule for diversion treatment wins; the best priority flips with system settings.","keywords":["recidivism","risk assessment","treatment allocation","agent-based simulation","social interactions","probation","incarceration diversion","policy evaluation"],"falsifier":"Re-estimate γ0 (the community feedback coefficient) and the baseline survival curve from repeated-offense and post-probation event data. If γ0 is negligible or changes sign, or if the baseline hazard for repeat events differs materially from the first-arrest curve, the two-regime ranking in Section 5.1 may invert or vanish.","tokens_in":15028,"feed_emoji":"⚖️","tokens_out":2630,"duration_ms":27552,"temperature":0.7,"pith_summary":"The paper argues that re-offense risk is not a static individual attribute but a dynamic quantity shaped by community feedback, so treatment allocation policies must be evaluated at the system level over long horizons. Using an agent-based simulation calibrated to U.S. probation data, it shows that prioritizing low-risk individuals performs better in the long run under low incarceration rates and long off-probation exposure, while prioritizing high-risk individuals is better in the short term or when incarceration truncates monitoring. The authors conclude that risk-based decision systems should be assessed as sociotechnical systems with long-term accountability. If correct, this undermines any universal risk-threshold rule for diversion decisions.","feed_headline":"Risk-based treatment priority flips with incarceration policy","feed_subtitle":"Agent-based model: low-risk priority wins long-term under lenient monitoring; high-risk wins short-term.","key_machinery":"The engine is a Cox proportional hazard model with a social-interaction term: an individual's hazard is scaled by γ0·μ, where μ is the community mean offense rate. Because μ aggregates everyone's offending, treatment decisions for one person change the environment for everyone else, creating an endogenous feedback loop. The agent-based simulation embeds this hazard in a full probation/off-probation/incarceration flow and evaluates policies under capacity constraints.","core_discovery":"The central discovery is a two-regime structure in long-run policy performance (Section 5.1). Neither the low-risk nor the high-risk prioritization policy dominates across values of the incarceration probability δinc and the off-probation duration Foff. Low-risk priority wins when δinc is small (offenders stay in the community, so preventing risk accumulation matters more); high-risk priority wins when δinc is large (offenders are removed quickly, so immediate offense prevention is worth more). The age-first-low-risk policy sits between the two. The mechanism is that low-risk treatment prevents drift into high-risk states, whereas high-risk treatment delivers larger immediate absolute risk r","pith_inferences":["A natural extension is to test dynamic, learning-based allocation that adapts as μ and individual histories evolve; the two-regime structure suggests an adaptive policy could outperform both static thresholds.","The framework implies that policies evaluated only on short-term outcomes will systematically overstate high-risk prioritization, since long-term feedback is omitted.","A direct empirical test: if community-level offense rates are manipulated (e.g., through targeted enforcement or treatment in a neighborhood), the sign and size of γ0 should predict whether low-risk or high-risk allocation reduces total community offending."],"forward_implications":["If correct, risk-threshold tools alone cannot determine whom to treat; the monitoring horizon and incarceration response must be part of the decision.","Low-risk prioritization should be favored when long-term community exposure is high (low incarceration, long off-probation periods).","High-risk prioritization is preferable when monitoring periods are short or incarceration quickly removes offenders.","Policy rankings shrink as capacity grows or arrivals fall, since more resources per person dampen allocation differences.","Heterogeneous treatment effects can invert intuition: even when low-risk individuals benefit more, the low-risk policy may raise overall per-capita offenses because group-H offenses dominate."],"fun_headline_variants":["Treatment priority flips with incarceration rate","Low-risk wins long-term, high-risk short-term","Who to treat first? Depends on jail policy","Agent simulation: risk priority shifts with jail use"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The simulation assumes the baseline survival function and community-feedback coefficient estimated from time-to-first-arrest data (Baltimore, 1986-1989) continue to govern repeated offenses and off-probation dynamics over a 30,000-day horizon.","fun_headline_variants_meta":{"raw":{"variants":["Treatment priority flips with incarceration rate","Low-risk wins long-term, high-risk short-term","Who to treat first? Depends on jail policy","Agent simulation: risk priority shifts with jail use"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1037,"prompt_tokens":718,"completion_tokens":319,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":462,"completion_tokens_details":{"reasoning_tokens":272}},"tokens_in":462,"tokens_out":319,"duration_ms":5224,"temperature":1.0,"reasoning_tokens":272,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T09:46:02.605000+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate γ0 (the community feedback coefficient) and the baseline survival curve from repeated-offense and post-probation event data. If γ0 is negligible or changes sign, or if the baseline hazard for repeat events differs materially from the first-arrest curve, the two-regime ranking in Section 5.1 may invert or vanish.","supporting_citations":[],"review_version":1}