{"id":"f9da627c-4362-48ed-aa13-8121c8916eb4","arxiv_id":"1909.00066","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper introduces counterfactual performance and fairness metrics for risk assessments, estimates them with doubly robust methods, and proves that observational fairness parity implies counterfactual parity only under strong balance conditions.","lead":"Risk assessment tools used to guide decisions should be judged on outcomes under a chosen action, not on historical outcomes that depend on past decisions. This paper gives a doubly robust way to estimate such counterfactual performance and fairness metrics, and shows that standard fairness corrections can make the counterfactual version of the same metric worse.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Child-welfare demonstration that DR evaluation reverses the model ranking rests on untestable exchangeability (assumption 2 in §3.2.2); unmeasured confounding could bias all counterfactual estimates, and a sensitivity analysis would show whether the real-data conclusion is robust.","rationale":"I read the paper as making two separable claims: (i) counterfactual metrics defined under a baseline intervention are the right evaluation targets for risk assessments, and DR estimators consistently estimate them under standard causal assumptions; (ii) on synthetic and child welfare data, the proposed evaluation changes the apparent ranking of observational and counterfactual models, and fairness corrections can harm counterfactual parity. The theory (Theorems 1-3) is internally consistent: I checked the algebra in Appendix B, including the cancellation steps in the sufficiency proofs, and found no error. The synthetic experiment is well designed because both potential outcomes are observed, so the 'true counterfactual' column provides a ground-truth check; the DR curves align with it. The fairness experiments illustrate Theorem 1 as claimed. The vulnerable link is the child welfare analysis, which is the only place where the paper claims to 'demonstrate' the approach on real data. That demonstration requires Y0 ⟂ T | X. The paper states this is untestable and gives a qualitative argument, but in a setting where human screeners make decisions with access to live calls, unmeasured confounders are credible. If they exist, the DR estimates are biased, and the rank reversal observed in Figures 2-3 and the downstream task results may not reflect true counterfactual performance. This is exactly the reader's weakest assumption. Because the central methodological contribution remains valid under stated assumptions and is verified in simulation, I would not change the ACCEPT verdict, but I would treat the child-welfare conclusions as conditional on the plausibility of exchangeability and encourage a formal sensitivity analysis.","tokens_in":26349,"tokens_out":10792,"duration_ms":90955,"concrete_test":"Re-run the child welfare evaluation under a range of simulated unmeasured confounders U: generate U from a Bernoulli with prevalence p, log-odds ratios α for T and β for Y0 given X, re-estimate propensity and outcome regressions including U, and recompute the DR calibration/PR/ROC curves and the Table 1 task-adaptation AUROC for the counterfactual and observational models over a grid of (α, β). Report the smallest confounding strength at which the DR ranking reverses or the task-adaptation gap disappears; compare that strength with the magnitude of associations among measured covariates. If modest confounding reverses the conclusions, the real-data claim should be downgraded; if only implausibly strong confounding does, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's foundational theoretical and synthetic contributions are sound, but the real-data conclusion — that the counterfactual model is better calibrated and better discriminating under DR evaluation in child welfare (§3.4.2, Figures 2-3) and transfers to downstream tasks (§3.4.4) — is identified only under assumption (2), Y0 ⟂ T | X. The Allegheny County screen-in decision is made by call workers using information that is plausibly richer than the 1000 structured features; unrecorded judgment, interview tone, or family context could jointly influence screening and re-referral under no investigation. If such a confounder U exists, then E[Y | X, T=0] ≠ E[Y0 | X], and every DR estimator in §3.3.3 is biased for its counterfactual target; the observed ranking of the observational and counterfactual models under DR evaluation could be an artifact of confounding rather than a true difference in counterfactual risk. The paper explicitly acknowledges that exchangeability is untestable and argues it 'may be reasonable' because measured variables capture most screener information, but provides no quantitative support for this. This does not threaten the methodological contribution, which is valid under its stated assumptions and is verified on synthetic data; it limits the strength of the empirical claim. A calibrated sensitivity analysis is the concrete missing piece.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that standard evaluation and fairness metrics for risk assessment instruments (RAIs) are distorted by historical treatment decisions, and it introduces counterfactual analogs targeting outcomes under a baseline intervention. It defines doubly robust (DR) estimators for counterfactual TPR, precision, FPR, and calibration (§3.3.3), proves three theorems characterizing when observational fairness (base rate parity, predictive parity, equalized odds) implies the corresponding counterfactual fairness (§4.1), and demonstrates the framework on a synthetic dataset with known potential outcomes and on Allegheny County child welfare hotline data (§3.4). The empirical sections show that DR evaluation reverses the model ranking seen under observational evaluation, and that fairness-corrective methods such as reweighing and post-processing can increase counterfactual disparity.","tokens_in":26605,"tokens_out":9012,"duration_ms":80912,"significance":"If the results hold, the paper provides a principled correction to the evaluation of risk assessments in decision-support settings, where the observed outcome is affected by the historical intervention. The theoretical contributions are strong: Appendix A gives clean identifications under consistency, exchangeability, and positivity; Appendix B gives algebraic proofs of the balance conditions; Theorems 1–3 are substantive and correctly derived. The synthetic experiments are particularly convincing because they evaluate against the true potential outcomes, and the code is publicly released. The real-data analysis is a useful illustration but is limited by the untestable exchangeability assumption; this does not undermine the methodological core, which stands on its own.","major_comments":[{"comment":"The real-data conclusion that the counterfactual model outperforms the observational model under DR evaluation is identified only under the exchangeability assumption Y0 ⟂ T | X. The authors acknowledge that this assumption is untestable and argue it is plausibly satisfied, but the strength of the empirical claims (e.g., 'DR evaluation shows that the counterfactual model is well-calibrated and the observational model underestimates risk') goes beyond what can be supported without a sensitivity analysis. Because the child welfare demonstration is a stated contribution, I request a sensitivity analysis that quantifies how large unmeasured confounding would need to be to alter the model ranking, or, failing that, a clear statement in the conclusions that these results are conditional on exchangeability.","section":"§3.4.2, assumption (2)"}],"minor_comments":[{"comment":"In the sentence about 'pure predition' settings, 'predition' should be 'prediction'.","section":"§1"},{"comment":"The heading 'Theorem 3 (Eqalized Odds)' contains a typo; it should be 'Equalized Odds'.","section":"§4.1.3"},{"comment":"The caption uses 'counterfatual' instead of 'counterfactual'.","section":"Appendix D, Figure 9"},{"comment":"The calibration confidence interval is computed with the plug-in bin proportion in the denominator and treats the residual variance as known; the text should note that uncertainty in the estimated bin proportion is ignored, or cite a delta-method treatment.","section":"§3.3.3"},{"comment":"The AUROC for the observational model on the placement task has a confidence interval of (0.46,0.49), which is below 0.5; a brief comment on whether this is statistically distinguishable from random would be helpful.","section":"§3.4.4, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The methodological core is sound and the synthetic validation is convincing. My one substantive concern is the strength of the real-data claims given the untestable exchangeability assumption; a sensitivity analysis would resolve this. The paper fits the journal's scope and is likely to be of high interest to the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: read this for the theory, not the child welfare results. The paper gives the first doubly robust estimators for counterfactual classification metrics in risk-assessment settings, and it proves necessary-and-sufficient balance conditions showing that observational fairness parity does not imply counterfactual parity. Those two contributions are real and correct. The synthetic experiments check the estimators against the true potential outcomes and the code is public, which gives me confidence in the machinery.\n\nWhere the paper is weakest is the real-data demonstration. The Allegheny County analysis argues the counterfactual model is better calibrated and better discriminated under DR evaluation, but that conclusion rides entirely on exchangeability (assumption 2, Y0 ⟂ T | X). The authors concede it is untestable. If call workers use information beyond the 1000 structured features—judgment, tone, context—the DR estimates are biased and the model ranking could be an artifact. There is no sensitivity analysis to show how strong an unmeasured confounder would have to be to change the conclusion. That makes the child welfare claims illustrative rather than decisive.\n\nThe fairness section is the strongest part. Theorems 1-3 are algebraically correct under the stated assumptions, and the paper is upfront about how heavy the independence conditions are. The synthetic demonstration that reweighing and post-processing can inflate counterfactual disparity is clean and well-executed. It is a proof of concept; the real-data fairness results are thin because the child welfare data is balanced on the metrics they examine.\n\nMinor issues: two figures lack the promised confidence intervals, and the case-study data is not available for independent verification. Neither undermines the contribution.\n\nBottom line: this deserves a serious referee. The methodological core is sound and novel, the synthetic validation is honest, and the theoretical results will be useful to anyone thinking about causal fairness metrics. The child welfare application should be read as a motivating case study, not as the evidence for the method. Send it to review; it will get cited.","headline":"The methodological core—DR estimators for counterfactual metrics and the balance conditions linking observational to counterfactual fairness—is solid and novel; read the child welfare claims as an illustration, not decisive evidence, since they depend on an untestable exchangeability assumption.","tokens_in":27163,"tokens_out":2103,"would_cite":true,"duration_ms":18714,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Risk assessments should be judged by risk under a baseline intervention, not by outcomes already changed by historical decisions; doubly robust estimators make this possible, and observational fairness parity rarely carries over.","keywords":["counterfactual risk assessment","doubly robust estimation","algorithmic fairness","potential outcomes","decision support","child welfare screening","equalized odds","base rate parity"],"falsifier":"On a dataset where screening decisions were randomized, or where an exogenous shock changed screen-in rates, estimate the counterfactual metrics with the paper's doubly robust procedure and compare them to direct randomized estimates; disagreement beyond sampling error would falsify the identifying assumptions. Without such data, a sensitivity analysis that injects an unmeasured confounder correlated with both screen-in and re-referral, and asks how strong the association must be to erase the counterfactual model's advantage, would settle how load-bearing the exchangeability assumption is.","tokens_in":26145,"feed_emoji":"⚖️","tokens_out":7326,"duration_ms":62808,"temperature":0.7,"pith_summary":"This paper argues that risk assessment tools used in decision support, such as child welfare hotline screening, are typically trained and evaluated on observed outcomes, so they measure risk under whatever historical decisions were made rather than under the intervention the tool is meant to inform. The authors define counterfactual analogues of standard performance metrics (true positive rate, precision, false positive rate, calibration) in terms of the outcome under a baseline treatment, and provide doubly robust estimators for them. They prove that observational fairness parity implies counterfactual parity only under special balance conditions, so equalizing a standard metric such as base rate parity or equalized odds can, and in their synthetic experiments does, increase disparity in the counterfactual version. On a large child welfare hotline dataset, doubly robust evaluation reverses the usual verdict: a model of risk under screen-out is better calibrated and better at identifying cases later placed out of home or given services than the standard observational model. The paper's core point is that predictive accuracy and fairness are the wrong targets when decisions change outcomes; the right targets are counterfactual.","feed_headline":"Fairness fixes can worsen the disparity they target","feed_subtitle":"Risk tools read historical outcomes as risk; counterfactual metrics flip which model wins and why fixes backfire.","key_machinery":"The load-bearing object is the counterfactual risk model $E[Y^0 \\mid X]$, the probability of the adverse outcome under the baseline treatment such as no investigation, together with its doubly robust estimator $$\\widehat{DR}_{$Y^{0}$} = \\frac{1}{n}\\sum_{i=1}^n \\left[ \\frac{1 - T_i}{1 - \\hat{\\pi}(X_i)}(Y_i - \\hat{s}_0(X_i)) + \\hat{s}_0(X_i) \\right],$$ which is consistent if either the propensity model or the outcome regression is correctly specified. The second engine is the family of balance conditions (balBP, balPP, balEO), explicit algebraic conditions on the joint distribution of potential outcomes, treatment, and group membership that are necessary and sufficient for observational parity to imply counterfactual parity. The doubly robust estimators for TPR, precision, FPR, and calibration each substitute the augmented outcome $\\frac{1-T}{1-\\hat{\\pi}(X)}(Y - \\hat{s}_0(X)) + \\hat{s}_0(X)$ into the relevant expectation, which corrects for the fact that treated cases' observed outcomes were altered by the decision itself.","core_discovery":"The central claim is that in decision settings the quantity a risk assessment should estimate is the potential outcome under a baseline intervention, $E[Y^0 \\mid X]$, rather than the observed outcome $E[Y \\mid X]$, and that evaluation and fairness metrics should be defined against $Y^0$. The paper identifies each counterfactual metric under consistency, exchangeability, and weak positivity, and proposes doubly robust estimators that combine a plug-in outcome regression with an inverse-probability-weighted residual correction; under sample splitting and $n^{-1/4}$ nuisance convergence these are $\\sqrt{n}$-consistent and asymptotically normal. The theoretical core is a set of three theorems: observational base rate parity implies counterfactual base rate parity if and only if a balance condition (balBP) holds, with analogous necessary-and-sufficient conditions for predictive parity (balPP) and equalized odds (balEO), and with independence conditions given as sufficient cases. The paper further shows empirically that two standard fairness-corrective procedures, reweighing for demographic parity and post-processing for equalized odds, produce counterfactual disparity where none existed before when treatment assignment was already biased. In the child welfare application, the doubly robust evaluation shows the counterfactual model is well-calibrated by race and that the observational model underestimates re-referral risk, conclusions the standard observational evaluation would invert.","pith_inferences":["A testable audit follows directly from the theorems: estimate the propensity score and the potential-outcome distributions by group and check whether a balance condition such as $\\mathrm{balBP}$ is approximately satisfied; if it is not, observational parity claims provide no evidence about counterfactual parity.","Because a baseline intervention must be chosen for every counterfactual metric, operationalizing this approach requires a policy judgment, usually 'no intervention,' which regulators and agencies would need to make explicit; part of the fairness debate thereby shifts from statistics to normative choice.","The same doubly robust machinery can evaluate not only risk models but the decisions themselves: ranking metrics for responsiveness-targeted interventions, continuous treatment doses, and off-policy comparisons of alternative screening thresholds are direct extensions the paper sketches but does not implement."],"forward_implications":["Evaluating a risk model against observed outcomes overstates the performance of observational models, since treated cases' outcomes were partly determined by the intervention; replacing the observed outcome with the counterfactual outcome under the baseline changes which model is selected.","Doubly robust counterfactual evaluation is computable for all standard classification metrics, comes with confidence intervals, and is consistent when at least one of the propensity or outcome-regression models is correctly specified.","Observational fairness parity transfers to counterfactual parity only under balance conditions that require, roughly, no residual treatment bias within risk strata and no group differences in risk under treatment, conditions unlikely to hold in child welfare and criminal justice.","Fairness-corrective methods that equalize observed metrics can induce counterfactual disparity that harms the group historically less likely to receive beneficial treatment, because the correction bakes historical bias into the score.","In the child welfare data, the observational model trained on observed re-referrals performs worse than a random classifier at predicting downstream outcomes such as out-of-home placement and services, while the counterfactual model transfers to those related risk tasks."],"supporting_citations":[{"why":"Supplies the doubly robust estimation framework that the counterfactual metrics are built upon.","marker":"[40]"},{"why":"Provides the semiparametric theory and doubly robust loss estimation used for counterfactual evaluation and model selection.","marker":"[51]"},{"why":"Introduces doubly robust policy evaluation in offline bandit settings, the direct antecedent of the proposed evaluation method.","marker":"[14]"},{"why":"Proposes counterfactual risk models for decision support and evaluates them on observed outcomes, the practice the paper argues is insufficient.","marker":"[44]"},{"why":"Defines equalized odds, the fairness notion whose counterfactual analogue the paper analyzes and whose post-processing the experiments apply.","marker":"[17]"},{"why":"Provides the reweighing procedure that the fairness experiments show can induce counterfactual base rate disparity.","marker":"[20]"},{"why":"Supplies the post-processing method for equalized odds and calibration used in the synthetic fairness experiments.","marker":"[37]"},{"why":"The prior child welfare screening study whose data and risk assessment context motivate and validate the counterfactual evaluation.","marker":"[8]"}],"fun_headline_variants":["Risk tools reflect past policies, not future outcomes","Fairness corrections can amplify real-world bias","Counterfactual metrics reveal fairness paradox","Retraining on history misses counterfactual risk","Doubly robust fairness: standard tests mislead"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The empirical and real-data conclusions assume exchangeability, $Y^0 \\perp T \\mid X$: the measured features capture every variable that jointly affects whether a call is screened in and whether the family is re-referred within six months. The paper states this assumption is untestable; if unmeasured confounding is present, the doubly robust estimates are biased and the claim that the counterfactual model outperforms the observational model on child welfare data is not supported.","fun_headline_variants_meta":{"raw":{"variants":["Risk tools reflect past policies, not future outcomes","Fairness corrections can amplify real-world bias","Counterfactual metrics reveal fairness paradox","Retraining on history misses counterfactual risk","Doubly robust fairness: standard tests mislead"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000385,"raw_usage":{"total_tokens":2101,"prompt_tokens":1078,"completion_tokens":1023,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":694,"completion_tokens_details":{"reasoning_tokens":955}},"tokens_in":694,"tokens_out":1023,"duration_ms":9203,"temperature":1.0,"reasoning_tokens":955,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:04:05.466312+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a dataset where screening decisions were randomized, or where an exogenous shock changed screen-in rates, estimate the counterfactual metrics with the paper's doubly robust procedure and compare them to direct randomized estimates; disagreement beyond sampling error would falsify the identifying assumptions. Without such data, a sensitivity analysis that injects an unmeasured confounder correlated with both screen-in and re-referral, and asks how strong the association must be to erase the counterfactual model's advantage, would settle how load-bearing the exchangeability assumption is.","supporting_citations":[{"cited_title":"In Advances in Neural Information Processing Systems","cited_arxiv_id":null,"evidence_quote":"Supplies the doubly robust estimation framework that the counterfactual metrics are built upon."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the semiparametric theory and doubly robust loss estimation used for counterfactual evaluation and model selection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proposes counterfactual risk models for decision support and evaluates them on observed outcomes, the practice the paper argues is insufficient."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines equalized odds, the fairness notion whose counterfactual analogue the paper analyzes and whose post-processing the experiments apply."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the reweighing procedure that the fairness experiments show can induce counterfactual base rate disparity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The prior child welfare screening study whose data and risk assessment context motivate and validate the counterfactual evaluation."}],"review_version":1}