{"id":"5df6e06b-a6b7-4aa7-bef2-a2c1403d0f38","arxiv_id":"2607.03161","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Variance-subtracted doubly robust proxy errors plus conformal p-values and Benjamini–Hochberg yield asymptotic FDR control for selecting reliable CATE predictions under heteroskedasticity.","lead":"The paper introduces Denoised Conformal Alignment, a way to pick only those people for whom a black-box treatment-effect model is accurate enough, while controlling the share of bad picks. It matters because hospitals and platforms already act on model-chosen subsets, and ordinary uncertainty guarantees do not protect those subsets.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Asymptotic FDR control can fail to transfer when m E[bΔ_cal] does not vanish under realistic nuisance rates and multi-arm overlap stress.","rationale":"The reader correctly isolates Assumption 4.2 / vanishing mislabeling as the weakest link and already flags the K=5 FDR inflation. The stress-test sharpens that point: the only quantitative bridge from oracle finite-sample FDR (Lemma 4.4) to the causal claim is the additive term m E[bΔ_cal] in Proposition 4.5, and the paper supplies neither a non-asymptotic bound under realistic nuisance rates nor a demonstration that this term is small precisely when heteroskedasticity and overlap stress make denoising necessary. Experiments that keep realized FDR near the diagonal under moderate \rho do not close the gap unless they also show bΔ_cal = o(1/m). Because the limitation is already acknowledged and the rest of the pipeline (oracle validity, denoising motivation, power barrier) is coherent, the verdict remains CONDITIONAL rather than REJECT; the concrete calibration-size \times mislabeling audit above would decide whether the condition can be relaxed or must stay in the statement.","tokens_in":46913,"tokens_out":838,"duration_ms":9112,"concrete_test":"Re-run the Setting-2 (hard-overlap + heavy-tail) and multi-treatment K=5 protocols with fixed α grid, reporting realized FDR together with empirical bΔ_cal (oracle vs denoised labels on D_cal) and the product m · bΔ_cal for n_cal ∈ {200,500,1000,2000} and m ∈ {500,1000,2000}. If m · bΔ_cal remains ≥ 0.05–0.1 while realized FDR exceeds α by a comparable amount, the asymptotic transfer fails in the paper’s own stress regimes.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central safety claim (Theorem 4.6) is asymptotic FDR control for BH on proxy p-values. Proposition 4.5 gives the finite-sample bound FDR ≤ α + m E[bΔ_cal | g], so the theorem reduces to showing m E[bΔ_cal] \to 0. Lemma B.6 obtains bΔ_cal \to 0 from mean-square consistency of (μ̂, ê, V̂) plus P(A = c) = 0 (Assumptions 4.2–4.3). That step is load-bearing and fragile: (i) the paper itself notes that typical empirical-process rates give m E[bΔ_cal] = O(m √(log m / n_cal)), which vanishes only under the strong regime n_cal ≫ m^{2} log m; (ii) under hard overlap / heavy tails / multi-arm dilution, inverse-propensity factors make Var(φ | X) large and V̂ inconsistent at the rates needed for boundary stability (Lemma B.5), so threshold labels 1{Ǎ ≥ c} flip and bΔ_cal stays order-1; (iii) the authors’ own K=5 multi-treatment experiment (Appendix C.5, Fig. 29) already exhibits realized FDR inflation under that stress. Thus the strongest claim holds only when nuisance/variance estimation is good enough that proxy/oracle labels agree near c at a rate faster than 1/m; the paper does not establish that this holds under the same regimes where denoising is most needed for power.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies selective deployment of black-box CATE predictors: select a subset of candidates whose CATE prediction errors fall below a user tolerance c while controlling post-selection FDR. Because unit-level CATE errors are unobservable, the authors form doubly robust proxy errors from pseudo-outcomes, then construct a variance-subtracted (denoised) proxy to recover ranking signal under heteroskedasticity. An alignment model maps covariates to these scores; split-conformal left-tail p-values and Benjamini–Hochberg yield the selected set (Denoised Conformal Alignment, DCA). Theory gives oracle finite-sample FDR control, a finite-sample perturbation bound driven by proxy/oracle threshold mislabeling, asymptotic FDR under identification and nuisance consistency, and a signal-to-noise barrier explaining power collapse of naive proxies. Experiments on heteroskedastic, hard-overlap, covariate-shift, semi-synthetic (IHDP/NLSM), and multi-treatment settings report improved power with controlled or conservative FDR, plus an honest multi-arm failure case.","tokens_in":47305,"tokens_out":1588,"duration_ms":24771,"significance":"If the results hold as claimed, the paper supplies a deployment-first primitive that standard marginal conformal CATE intervals do not: post-selection FDR control for which CATE predictions are safe to act on. The isolation of validity to boundary-label stability (rather than pointwise variance perfection) is a clean conceptual contribution, and the bias–variance decomposition of DR proxy error motivates denoising in a way that is both theoretically and empirically useful. Strengths include careful adaptation of the Jin–Candès / Gui et al. conformal-alignment template to counterfactual labels, explicit finite-sample perturbation and margin lemmas, extensive stress tests (including covariate-shift weighting and an admitted K=5 multi-treatment breakdown), and reproducible experimental protocols. The work is a meaningful step for selective causal decision-making, provided the asymptotic rate conditions and practical transfer under hard nuisance estimation are stated and stress-tested more carefully.","major_comments":[{"comment":"Theorem 4.6 vs. its proof (and Prop. 4.5): Proposition 4.5 gives FDR(S) ≤ α + m E[bΔ_cal | g]. The proof of Theorem 4.6 only concludes limsup FDR ≤ α “along any growth regime such that m E[bΔ_cal] → 0,” while the theorem statement asserts the limsup as reference and test sizes grow without a relative-rate condition. The paper’s own empirical-process remark (after Prop. 4.5) requires n_cal ≫ m² log m for the inflation term to vanish—an extremely strong regime that is not reflected in the theorem statement or main experimental sample sizes. Please restate Theorem 4.6 with an explicit growth condition (or prove a weaker rate under which m E[bΔ_cal] → 0 from Assumptions 4.2–4.3), and discuss finite-sample implications when m is large relative to n_cal.","section":"§4.2–4.3, Prop. 4.5, Thm. 4.6"},{"comment":"Transfer of asymptotic FDR where denoising is most needed: Lemma B.6 obtains bΔ_cal → 0 from mean-square nuisance/variance consistency plus P(A=c)=0. Under hard overlap, heavy tails, and multi-arm dilution, inverse-propensity factors inflate Var(φ|X) and make V̂ hard to estimate at rates that preserve boundary labels (Lemma B.5). The authors’ own multi-treatment K=5 experiment (Appendix C.5, Fig. 29) already shows realized FDR inflation under that stress. The central safety claim therefore holds only when proxy/oracle labels agree near c at a rate faster than 1/m—precisely the regimes where naive proxies fail and denoising is advertised as necessary. Please either (i) provide finite-sample or high-probability bounds linking overlap/tail conditions to bΔ_cal, or (ii) clearly demote the safety claim in those regimes and report realized FDR more systematically as a function of overlap and n","section":"Assump. 4.2, Lem. B.5–B.6, App. C.5 Fig. 29"},{"comment":"Covariate-shift extension and BH: Section 3.4 and Appendix C.2.3 replace conformal counts by importance-weighted sums and apply BH, while acknowledging that weighted p-values need not satisfy PRDS and that finite-sample BH control is not established (WCS-style pruning is deferred). Figure 6 and Setting 3–4 experiments nonetheless present the weighted procedure as maintaining FDR stability. Either supply conditions under which weighted BH controls FDR (even asymptotically under proxy stability), or reframe the weighted results as empirical diagnostics only and avoid implying the same FDR guarantee as the unweighted case.","section":"§3.4, Eq. (7), App. C.2.3, Fig. 6"}],"minor_comments":[{"comment":"Clarify the operational choice of ρ for purely real data (no oracle CATE on a validation slice). The fixed validation rule in C.1.4 is fine for semi-synthetic settings; a short practical default (e.g., conservative ρ grid + proxy-stability criterion) would help deployment readers.","section":"§3.2, App. C.1.4"},{"comment":"Notation: the manuscript mixes eA, Ã, Ǎ, and ˇA for raw/denoised proxies across main text and figures; unify symbols and define them once near Eq. (5).","section":"§3.1–3.2, Fig. 1–2"},{"comment":"Assumption 4.1 is stated for covariates/scores conditional on g; briefly note that sample splitting of D into D_tr1/D_tr2/D_cal is what makes this plausible, as done later in the appendix.","section":"Assump. 4.1, Alg. 1"},{"comment":"Figure 5 panel labels and ρ>1 stress tests are useful; state explicitly in the caption that ρ>1 is outside the main theory’s [0,1] range and is only a stress test.","section":"Fig. 5, §5"},{"comment":"Related work: Jin & Candès (2026) on weighted conformal p-values under shift is cited; a one-sentence contrast with mFDR-type thresholds (Sun & Cai) is already present—consider moving a short version into the main related-work paragraph for readers who skip the appendix.","section":"App. A"}],"recommendation":"major_revision","confidential_remarks":"Solid contribution for a top ML/stats venue if the theorem statement is aligned with the proof’s rate condition and the hard-nuisance transfer is handled honestly. The K=5 failure and the n_cal ≫ m² log m remark are already in the manuscript; requiring the authors to surface them in the main theorem statement and discussion is the right bar, not a reason to reject. Scope fits stat.ML / selective inference well."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: if you want to act only on CATE predictions that are accurate enough, marginal conformal intervals do not give you that, and raw DR proxy errors often rank noise instead of accuracy. This paper builds a practical fix—subtract estimated conditional variance from the DR squared proxy, train an alignment score on the denoised labels, then run conformal p-values + BH.\n\nWhat is actually new is the variance-subtracted proxy plus the boundary-stability framing. They do not claim pointwise perfect V̂; they show FDR inflation is controlled by the calibration mislabeling rate bΔ_cal (Prop. 4.5), and that denoising restores ranking separability when Var(φ|X) swamps A² (Prop. 4.7). Oracle finite-sample FDR follows the Jin–Candès leave-one-out template carefully; the asymptotic causal theorem is honest about what is finite-sample versus asymptotic. Experiments are the strong part: heteroskedastic Gaussians, hard overlap + heavy tails, weighted shift, IHDP/NLSM, and ablations on ρ, misspecified V, and split size. Denoising clearly recovers power where plug-in, ensemble std, and CQR width stall, while realized FDR stays near or below target in the main regimes.\n\nSoft spots are real but proportionate. The safety claim is asymptotic and load-bearing on m E[bΔ_cal] → 0 under nuisance/variance consistency; their own rate note (n_cal ≫ m² log m) and the K=5 multi-arm FDR inflation under overlap stress show the transfer can fail exactly where proxies are noisiest. ρ is a free operational knob (they tune on a validation slice, not test outcomes), and weighted BH under shift is a drop-in asymptotic extension rather than a full finite-sample WCS argument. None of that undoes the core contribution; it bounds the claim.\n\nMath and citations look solid: Gui et al., Jin–Candès, DR/DML, and causal conformal work are used correctly, not as decoration. Code and protocols are detailed enough to reproduce.\n\nThis is for people who ship selective treatment policies or wrap black-box CATE models. I would bring it to reading group, cite the denoising + boundary-stability idea when I need post-selection CATE reliability, and send it to peer review. Engage.","headline":"Useful deployment wrapper for CATE selection: variance-denoised DR proxies + conformal alignment give asymptotic FDR control when proxy/oracle labels stay stable near the tolerance, with real power gains where naive proxies collapse.","tokens_in":47925,"tokens_out":586,"would_cite":true,"duration_ms":7645,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Subtracting estimated noise from causal error proxies restores reliable selective deployment under FDR control.","keywords":["conditional average treatment effect","selective deployment","false discovery rate","conformal prediction","doubly robust proxies","heteroskedasticity","denoising"],"falsifier":"In a high-heteroskedasticity synthetic design with known true CATE, run Denoised Conformal Alignment at moderate alpha: if realized FDR on the selected set systematically exceeds alpha while naive DR proxies remain empty, or if moderate variance subtraction fails to raise power above the naive baseline, the central claim fails.","tokens_in":47769,"feed_emoji":"🎯","tokens_out":554,"duration_ms":5813,"temperature":0.7,"pith_summary":"When a model ranks people by predicted treatment effect and you act only on the top of that ranking, average accuracy guarantees no longer protect the people you actually treat. This paper targets that post-selection risk: pick a subset whose CATE prediction error is below a user tolerance while keeping the expected fraction of bad picks under a false-discovery limit. True CATE errors are counterfactual, so the authors build doubly robust proxy errors from pseudo-outcomes; under heteroskedasticity those proxies are dominated by variance, so ranking collapses to noise. Denoised Conformal Alignment subtracts an estimated conditional variance, trains an alignment score on the cleaned proxies, and feeds conformal p-values into Benjamini–Hochberg. Validity depends on stable reliable/unreliable labels near the tolerance boundary, not on perfect variance recovery, and experiments recover substantial selection yield where naive proxies fail.","feed_headline":"Denoising causal error proxies unlocks safe selective deployment","feed_subtitle":"Variance subtraction restores ranking power while conformal FDR control holds on selected units","key_machinery":"Denoised proxy error: square-root of the positive part of raw DR squared error minus rho times an estimated conditional variance. It turns a noise-dominated ranking into a score that can be aligned and conformalized so that FDR is controlled by vanishing mislabeling near the tolerance.","core_discovery":"Under standard identification and sample-split nuisance consistency, variance-subtracted doubly robust proxy errors make proxy-versus-oracle threshold labels agree often enough that conformal alignment plus Benjamini–Hochberg yields asymptotic FDR control for selecting units whose unobserved CATE prediction error lies below a tolerance; the same denoising restores the ranking signal that heteroskedastic noise otherwise erases.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Denoised conformal proxies control FDR for CATE selection","Variance subtraction restores ranking in selective CATE deployment","Proxy denoising enables reliable black-box CATE unit selection","Conformal alignment after variance subtract yields FDR control","Denoised doubly robust errors select low-error CATE predictions"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The nuisance and variance models trained on a held-out split must be accurate enough that almost no calibration units flip between “reliable” and “unreliable” at the chosen error tolerance; when overlap is poor or tails are heavy that agreement can fail.","fun_headline_variants_meta":{"raw":{"variants":["Denoised conformal proxies control FDR for CATE selection","Variance subtraction restores ranking in selective CATE deployment","Proxy denoising enables reliable black-box CATE unit selection","Conformal alignment after variance subtract yields FDR control","Denoised doubly robust errors select low-error CATE predictions"]},"model":"grok-4.5","effort":"low","cost_usd":0.005568,"raw_usage":{"total_tokens":1452,"prompt_tokens":690,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":55680000,"prompt_tokens_details":{"text_tokens":690,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":682,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":690,"tokens_out":80,"duration_ms":6138,"temperature":1.0,"reasoning_tokens":682,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T04:25:26.498153+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In a high-heteroskedasticity synthetic design with known true CATE, run Denoised Conformal Alignment at moderate alpha: if realized FDR on the selected set systematically exceeds alpha while naive DR proxies remain empty, or if moderate variance subtraction fails to raise power above the naive baseline, the central claim fails.","supporting_citations":[],"review_version":1}