Pith. sign in

REVIEW 3 major objections 5 minor 40 references

Denoised Conformal Alignment for Reliable Selection of Conditional Average Treatment Effect Predictions

T0 review · 3 major / 5 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read Subtracting estimated noise from causal error proxies restores reliable selective deployment under FDR control.

desk verdict Useful deployment wrapper for CATE selection: variance-denoised DR proxies + conformal alignment give asymptotic FDR control when proxy/oracle labels stay stable near the tolerance, with real power gains where naive proxies collapse. read the letter →

arxiv 2607.03161 v1 pith:4LYSJ3MP submitted 2026-07-03 stat.ML cs.LG

classification stat.MLcs.LG
keywords conditionalaveragetreatmenteffectselectivedeploymentfalsediscoveryrateconformalpredictiondoublyrobustproxiesheteroskedasticitydenoising
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

When a model ranks people by predicted treatment effect and you act only on the top of that ranking, average accuracy guarantees no longer protect the people you actually treat. This paper targets that post-selection risk: pick a subset whose CATE prediction error is below a user tolerance while keeping the expected fraction of bad picks under a false-discovery limit. True CATE errors are counterfactual, so the authors build doubly robust proxy errors from pseudo-outcomes; under heteroskedasticity those proxies are dominated by variance, so ranking collapses to noise. Denoised Conformal Alignment subtracts an estimated conditional variance, trains an alignment score on the cleaned proxies, and feeds conformal p-values into Benjamini–Hochberg. Validity depends on stable reliable/unreliable labels near the tolerance boundary, not on perfect variance recovery, and experiments recover substantial selection yield where naive proxies fail.

What carries the argument

Denoised proxy error: square-root of the positive part of raw DR squared error minus rho times an estimated conditional variance. It turns a noise-dominated ranking into a score that can be aligned and conformalized so that FDR is controlled by vanishing mislabeling near the tolerance.

What would settle it

In a high-heteroskedasticity synthetic design with known true CATE, run Denoised Conformal Alignment at moderate alpha: if realized FDR on the selected set systematically exceeds alpha while naive DR proxies remain empty, or if moderate variance subtraction fails to raise power above the naive baseline, the central claim fails.

Watch

Extended reading notes

Core claim

Under standard identification and sample-split nuisance consistency, variance-subtracted doubly robust proxy errors make proxy-versus-oracle threshold labels agree often enough that conformal alignment plus Benjamini–Hochberg yields asymptotic FDR control for selecting units whose unobserved CATE prediction error lies below a tolerance; the same denoising restores the ranking signal that heteroskedastic noise otherwise erases.

Load-bearing premise

The nuisance and variance models trained on a held-out split must be accurate enough that almost no calibration units flip between “reliable” and “unreliable” at the chosen error tolerance; when overlap is poor or tails are heavy that agreement can fail.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies selective deployment of black-box CATE predictors: select a subset of candidates whose CATE prediction errors fall below a user tolerance c while controlling post-selection FDR. Because unit-level CATE errors are unobservable, the authors form doubly robust proxy errors from pseudo-outcomes, then construct a variance-subtracted (denoised) proxy to recover ranking signal under heteroskedasticity. An alignment model maps covariates to these scores; split-conformal left-tail p-values and Benjamini–Hochberg yield the selected set (Denoised Conformal Alignment, DCA). Theory gives oracle finite-sample FDR control, a finite-sample perturbation bound driven by proxy/oracle threshold mislabeling, asymptotic FDR under identification and nuisance consistency, and a signal-to-noise barrier explaining power collapse of naive proxies. Experiments on heteroskedastic, hard-overlap, covariate-shift, semi-synthetic (IHDP/NLSM), and multi-treatment settings report improved power with controlled or conservative FDR, plus an honest multi-arm failure case.

Significance. If the results hold as claimed, the paper supplies a deployment-first primitive that standard marginal conformal CATE intervals do not: post-selection FDR control for which CATE predictions are safe to act on. The isolation of validity to boundary-label stability (rather than pointwise variance perfection) is a clean conceptual contribution, and the bias–variance decomposition of DR proxy error motivates denoising in a way that is both theoretically and empirically useful. Strengths include careful adaptation of the Jin–Candès / Gui et al. conformal-alignment template to counterfactual labels, explicit finite-sample perturbation and margin lemmas, extensive stress tests (including covariate-shift weighting and an admitted K=5 multi-treatment breakdown), and reproducible experimental protocols. The work is a meaningful step for selective causal decision-making, provided the asymptotic rate conditions and practical transfer under hard nuisance estimation are stated and stress-tested more carefully.

major comments (3)
  1. [§4.2–4.3, Prop. 4.5, Thm. 4.6] Theorem 4.6 vs. its proof (and Prop. 4.5): Proposition 4.5 gives FDR(S) ≤ α + m E[bΔ_cal | g]. The proof of Theorem 4.6 only concludes limsup FDR ≤ α “along any growth regime such that m E[bΔ_cal] → 0,” while the theorem statement asserts the limsup as reference and test sizes grow without a relative-rate condition. The paper’s own empirical-process remark (after Prop. 4.5) requires n_cal ≫ m² log m for the inflation term to vanish—an extremely strong regime that is not reflected in the theorem statement or main experimental sample sizes. Please restate Theorem 4.6 with an explicit growth condition (or prove a weaker rate under which m E[bΔ_cal] → 0 from Assumptions 4.2–4.3), and discuss finite-sample implications when m is large relative to n_cal.
  2. [Assump. 4.2, Lem. B.5–B.6, App. C.5 Fig. 29] Transfer of asymptotic FDR where denoising is most needed: Lemma B.6 obtains bΔ_cal → 0 from mean-square nuisance/variance consistency plus P(A=c)=0. Under hard overlap, heavy tails, and multi-arm dilution, inverse-propensity factors inflate Var(φ|X) and make V̂ hard to estimate at rates that preserve boundary labels (Lemma B.5). The authors’ own multi-treatment K=5 experiment (Appendix C.5, Fig. 29) already shows realized FDR inflation under that stress. The central safety claim therefore holds only when proxy/oracle labels agree near c at a rate faster than 1/m—precisely the regimes where naive proxies fail and denoising is advertised as necessary. Please either (i) provide finite-sample or high-probability bounds linking overlap/tail conditions to bΔ_cal, or (ii) clearly demote the safety claim in those regimes and report realized FDR more systematically as a function of overlap and n
  3. [§3.4, Eq. (7), App. C.2.3, Fig. 6] Covariate-shift extension and BH: Section 3.4 and Appendix C.2.3 replace conformal counts by importance-weighted sums and apply BH, while acknowledging that weighted p-values need not satisfy PRDS and that finite-sample BH control is not established (WCS-style pruning is deferred). Figure 6 and Setting 3–4 experiments nonetheless present the weighted procedure as maintaining FDR stability. Either supply conditions under which weighted BH controls FDR (even asymptotically under proxy stability), or reframe the weighted results as empirical diagnostics only and avoid implying the same FDR guarantee as the unweighted case.
minor comments (5)
  1. [§3.2, App. C.1.4] Clarify the operational choice of ρ for purely real data (no oracle CATE on a validation slice). The fixed validation rule in C.1.4 is fine for semi-synthetic settings; a short practical default (e.g., conservative ρ grid + proxy-stability criterion) would help deployment readers.
  2. [§3.1–3.2, Fig. 1–2] Notation: the manuscript mixes eA, Ã, Ǎ, and ˇA for raw/denoised proxies across main text and figures; unify symbols and define them once near Eq. (5).
  3. [Assump. 4.1, Alg. 1] Assumption 4.1 is stated for covariates/scores conditional on g; briefly note that sample splitting of D into D_tr1/D_tr2/D_cal is what makes this plausible, as done later in the appendix.
  4. [Fig. 5, §5] Figure 5 panel labels and ρ>1 stress tests are useful; state explicitly in the caption that ρ>1 is outside the main theory’s [0,1] range and is only a stress test.
  5. [App. A] Related work: Jin & Candès (2026) on weighted conformal p-values under shift is cited; a one-sentence contrast with mFDR-type thresholds (Sun & Cai) is already present—consider moving a short version into the main related-work paragraph for readers who skip the appendix.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: asymptotic FDR control is a standard proxy-to-oracle perturbation argument, not a result forced by definition or self-citation.

full rationale

The paper's load-bearing safety claim (Theorem 4.6) is asymptotic FDR control for BH on conformal p-values built from denoised DR proxy labels. The derivation chain is: (i) oracle finite-sample FDR for true error labels (Lemma 4.4, recovering Gui et al. 2024 / Jin–Candès 2023 under exchangeability); (ii) finite-sample perturbation FDR ≤ α + m E[bΔ_cal] when oracle labels are replaced by proxies (Proposition 4.5); (iii) bΔ_cal → 0 under nuisance/variance consistency and no mass at c (Lemma B.6), so the additive term vanishes. This is a standard asymptotic transfer argument: validity is conditional on proxy/oracle threshold-label agreement near c, not on redefining the target as the estimator. Denoising (Eq. 5) is a constructed score, not a self-definition of reliability; ρ and c are free operational knobs chosen on a validation slice, and the theorems do not claim that any particular ρ is forced. Power optimality (Proposition 4.9) is only among score-threshold rules for a fixed learned alignment score—a restricted-class result, not global optimality by construction. Key citations (Gui et al. 2024, Jin–Candès 2023, Chernozhukov et al. 2018) are external; there is no load-bearing self-citation uniqueness theorem. The paper's own multi-treatment K=5 FDR inflation is an empirical limitation under failed assumptions, not circularity. The derivation is self-contained against external conformal-selection theory and does not reduce predictions to fitted inputs by construction.

Assumptions & free parameters 4 free parameters · 6 assumptions · 2 invented entities

The central asymptotic FDR claim rests on standard causal identification and split-conformal exchangeability, plus consistency of nuisance/variance estimators and continuity at the tolerance boundary. Operational free parameters are the tolerance c and denoising strength ρ; the method also invents the denoised proxy score and the alignment predictor trained on it. No new physical entities are postulated. Weighted covariate-shift validity and multi-treatment extensions add further regularity that the main theorem does not fully cover.

free parameters (4)
  • denoising strength ρ = grid-selected; representative values 0.65 / 0.15 / 0.60 depending on setting
    Controls variance subtraction in Eq. (5); chosen from a prespecified grid on an outcome-observed validation slice. Empirically regime-dependent (e.g., ~0.65 in Setting 1, ~0.15 under hard overlap/heavy tails).
  • reliability tolerance c = application-chosen; often calibration median/quantile in experiments
    User/application threshold defining H0: A ≥ c. In experiments often set from calibration-side quantiles or median squared oracle error; changes yield and null set.
  • target FDR level α = user-specified in (0,1)
    BH level; free deployment choice, not estimated from data for the theory claim.
  • sample-split sizes and learner hyperparameters = protocol defaults in Appendix C
    Dtr1/Dtr2/Dcal/Dtest proportions and RF/logistic defaults affect finite-sample nuisance quality and thus bΔ_cal; not part of the asymptotic statement but load-bearing in practice.
assumptions (6)
  • domain assumption Unconfoundedness and overlap identify τ(x) (Assumption 2.1).
    Standard potential-outcomes identification; without it CATE is not the right target and DR centering fails.
  • domain assumption Conditional exchangeability of calibration and test candidates given trained g and τ̂ (Assumption 4.1 / B.1).
    Required for conformal rank p-values; sample splitting is used to make it plausible.
  • domain assumption Mean-square consistency of μ̂t, ê, and V̂ as |Dtr1|→∞ (Assumption 4.2).
    Drives proxy→oracle label stability and asymptotic FDR; not free of estimation quality.
  • standard math P(Ai = c) = 0 (Assumption 4.3).
    Avoids boundary atoms so continuous mapping applies to threshold indicators.
  • standard math BH controls FDR for valid (or approximately valid) p-values under the paper’s exchangeability/PRDS-style arguments.
    Uses classical BH theory plus Jin–Candès conformal selection template; weighted case is weaker.
  • ad hoc to paper For covariate shift, density ratio w is known/estimable and conditional mechanism Y|(X,T) is invariant.
    Drop-in weighted p-values (Eq. 7); main FDR theorem is for the unweighted exchangeable case.
invented entities (2)
  • Denoised proxy error Ǎi(ρ)
    purpose: Recover a ranking signal for unobservable CATE error by subtracting estimated conditional variance from squared DR proxy error.
    Defined in Eq. (5); central algorithmic object. Independent evidence is empirical power recovery and the SNR barrier proposition, not an external measurement of Ǎ itself.
  • Alignment predictor g mapping X to predicted denoised error
    purpose: Provide observable selection scores for test candidates who lack outcomes.
    Trained on Dtr2; standard supervised wrapper around the invented proxy labels.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Denoised Conformal Alignment for Reliable Selection of Conditional Average Treatment Effect Predictions." pith.science (2026). https://pith.science/paper/4LYSJ3MP

@misc{pith2026260703161,
  author       = {Pith},
  title        = {Pith review of: Denoised Conformal Alignment for Reliable Selection of Conditional Average Treatment Effect Predictions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4LYSJ3MP}},
  note         = {Machine review of arXiv:2607.03161}
}
read the original abstract

In selective deployment, practitioners act only on a model-chosen subset of individuals based on predicted conditional average treatment effects, but marginal conformal guarantees need not control reliability on that selected subset. We study reliable selection for black-box CATE predictors: selecting candidates whose CATE errors are below a tolerance while controlling the false discovery rate (FDR). Since CATE errors are unobservable, we construct doubly robust proxy errors from pseudo-outcomes; however, naive proxies can lose power under heteroskedasticity because variance overwhelms the reliability signal. We propose Denoised Conformal Alignment, which subtracts an estimated conditional variance component and combines conformal calibration with Benjamini--Hochberg selection. Our analysis shows that validity is governed by stability of proxy/oracle threshold labels, rather than pointwise perfection of the variance estimator. Experiments show substantially improved power while maintaining FDR control across challenging settings.

Figures

Figures reproduced from arXiv: 2607.03161 by the authors.

Figure 1
Figure 1. Conceptual illustration. (Left) Reliable selection. Given covariates X, our goal is to identify individuals or groups whose CATE predictions are reliable, rather than relying on average effects or marginal guarantees. (Right) Denoised error proxies. Since the true CATE prediction error Ai is counterfactual and unobservable, we construct a noisy observable proxy Aei from doubly robust pseudo-outcomes, and denoise it … view at source ↗
Figure 2
Figure 2. End-to-end pipeline for denoised proxy construction and selective inference. (Left) Initialization and proxy construction. Using Dtr1, we estimate the base CATE predictor τˆ(X) and nuisance components; DR pseudo-outcomes ϕ form the raw proxy Ae. (Right) Denoising, alignment, and selection. Variance subtraction denoises Ae into Aˇ, which trains an alignment model on Dtr2 and is calibrated on Dcal to rank candidates a… view at source ↗
Figure 3
Figure 3. Gaussian outcomes with heteroskedastic proxy noise. Realized FDR (top) and power (bottom) versus target α for σ ∈ {0.5, 1.0, 1.5} with ρ = 0.65. Across all σ, DCA stays close to or below the nominal FDR line while achieving much higher power, indicating that variance-aware denoising mitigates the heteroskedastic proxy-noise barrier. More results are in Appendix C. that gains are not driven by a favorable split and o… view at source ↗
Figures from the paper (26 more)
Figure 4
Figure 4. Figure 4: Hard overlap with heavy-tailed outcomes. Realized FDR (top) and power (bottom) versus target α for σ ∈ {0.5, 1.0, 1.5} with a conservative denoising strength ρ = 0.15. Even when inverse-propensity factors and heavy-tailed residuals make raw DR proxies unstable, DCA rec…
Figure 5
Figure 5. Figure 5: Denoising ablations and robustness to imperfect variance estimates. Panel (a) compares no denoising, moderate subtraction, full subtraction, and over-subtraction. Panel (b) perturbs the variance model through well-specified, misspecified, under-estimated, and over-esti…
Figure 6
Figure 6. Figure 6: Covariate shift with weighted conformal calibration. Under Psrc(X) ̸= Ptgt(X) with invariant conditional causal mechanism, importance-weighted conformal p-values provide a drop-in calibration modification. DCA maintains higher selection yield than proxy and uncertainty…
Figure 7
Figure 7. Figure 7: Setting 1 sensitivity (Gaussian + heteroskedastic). Realized FDR and power across ρ ∈ {0.25, 0.5, 0.75, 1.0, 1.5} and σbase ∈ {0.5, 1.0, 1.5}. A wide range of moderate denoising strengths improves selection yield while preserving conservative or near-nominal FDR, consi…
Figure 8
Figure 8. Figure 8: Setting 2 sensitivity (hard overlap + heavy tails). Compared to Setting 1, the optimal denoising regime shifts toward smaller ρ: conservative subtraction improves yield while maintaining controlled/conservative realized FDR, whereas overly aggressive subtraction degrad…
Figure 9
Figure 9. Figure 9: Setting 3 sensitivity (covariate shift). Across shift types and noise scales, importance weighting stabilizes calibration under Psrc(X) ̸= Ptgt(X), while moderate denoising strengths continue to improve power without inducing fragile tuning behavior. T ~ Bernoulli(e_tr…
Figure 10
Figure 10. Figure 10: Setting 4 representative (covariance/rotation shift). Under geometry-only covari￾ate shift with invariant Y | (X, T), importance-weighted conformal calibration remains a drop￾in fix: it stabilizes FDR behavior while allowing DCA to preserve higher selection yield than…
Figure 11
Figure 11. Figure 11: Setting 4 sensitivity (covariance/rotation shift). Across geometry-shift constructions and noise scales, weighted p-values maintain calibration under Psrc(X) ̸= Ptgt(X), while moderate denoising continues to improve power without relying on fragile tuning. C.3.1 Sampl…
Figure 12
Figure 12. Figure 12: Calibration-size scaling. DCA remains useful across calibration sizes, indicating that the baseline sample splitting cost does not erase the benefit of denoising. Target 0.2 0.0 0.2 0.4 0.6 0.8 1.0 Realized FDR Nref = 3000 A: 3-fold Split (default) B: CrossFit + 2-fol…
Figure 13
Figure 13. Figure 13: Strict splitting versus cross-fitting. Cross-fitted variants across sample sizes show that moderate data reuse can stabilize selections without removing the denoising gain, while strict splitting remains the clean validity-first baseline. C.3.2 Denoising strength, ove…
Figure 14
Figure 14. Figure 14: ρ–overlap interaction. FDR and power over denoising strength and overlap quality in Setting 1. Harder overlap shifts the best denoising region toward smaller ρ, consistent with conservative subtraction under noisier variance estimates. Target 0.2 0.0 0.2 0.4 0.6 0.8 1…
Figure 15
Figure 15. Figure 15: Tolerance sensitivity using calibration-error quantiles. Varying c over calibration-side quantiles preserves the same safety–power pattern, indicating that DCA is not a knife-edge tolerance choice. counterpart of the boundary-stability argument in Proposition 4.5: den…
Figure 16
Figure 16. Figure 16: Tolerance sensitivity using absolute thresholds. Across noise levels and absolute choices of c, denoising improves selection yield while maintaining stable realized FDR. 0.0 0.2 0.4 0.6 0.8 1.0 Realized FDR B = 2 bins Standard Jitter Smooth y = 0.0 0.2 0.4 0.6 0.8 1.0…
Figure 17
Figure 17. Figure 17: Tie handling under truncation. Standard, jittered, and smooth tie-breaking variants show similar FDR/power behavior, suggesting that truncation-induced ties do not drive the empirical results. C.3.3 Tie handling under truncation Because Aˇ2 = (Ae2 − ρVb)+ can produce …
Figure 18
Figure 18. Figure 18: CATE estimator agnosticism. DCA versus naive proxy selection across base CATE estimators. Denoising consistently improves selection yield at comparable FDR behavior. C.3.5 Selection paradigms and downstream value [PITH_FULL_IMAGE:figures/full_fig_p044_18.png]
Figure 19
Figure 19. Figure 19: Influence-function proxy comparison. Naive and denoised DR/IF proxies across noise levels. Denoising improves the reliability-ranking signal for both proxy families. Target 0.2 0.0 0.2 0.4 0.6 0.8 1.0 Realized FDR = 0.5 Naive DR Proxy Denoised DR (DCA) Naive R-learner…
Figure 20
Figure 20. Figure 20: Orthogonal and R-learner proxy comparison. Variance-aware denoising remains beneficial for AIPW-orthogonal and R-learner-style proxy constructions, confirming proxy-level compatibility. use the observed outcome Y = yfactual and treatment T = treatment, and take covari…
Figure 21
Figure 21. Figure 21: Selection paradigm comparison. DCA, naive DR, ensemble uncertainty, Top-K, and oracle references. DCA combines FDR control with substantially higher power than proxy or uncertainty baselines. 0.2 0.4 0.6 0.8 1.0 Target 0.0 0.1 0.2 0.3 0.4 Per-Capita Policy Value Total…
Figure 22
Figure 22. Figure 22: Downstream policy value. Per-unit policy value and number of treated units under different selection policies. DCA converts improved reliable selection into higher downstream utility. What this benchmark tests (complement to synthetic settings). IHDP differs from the …
Figure 23
Figure 23. Figure 23: Small-sample behavior. DCA remains competitive in small-N regimes, with power improving as calibration and nuisance estimation become more stable. Target 0.0 0.2 0.4 0.6 0.8 1.0 Realized FDR d=10 Naive (Plug-in) Ours (Denoised) Target 0.0 0.2 0.4 0.6 0.8 1.0 Realized …
Figure 24
Figure 24. Figure 24: High-dimensional nuisance difficulty. Increasing dimension stresses nuisance and variance estimation; DCA retains the denoising advantage when the learned proxy labels remain sufficiently stable. downscale fold sizes to keep each nonempty # code safeguard fit nuisance…
Figure 25
Figure 25. Figure 25: Setting 5 (IHDP) sensitivity. Realized FDR and power across ρ ∈ {0.1, 0.15, 0.2, 0.25, 0.5} on semi-synthetic IHDP with simulator-provided ground-truth CATE. Mod￾erate denoising improves selection yield while maintaining conservative or near-nominal FDR; the represent…
Figure 26
Figure 26. Figure 26: NLSM semi-synthetic benchmark. Three NLSM data-generating mechanisms at multiple noise levels. DCA’s safety–power pattern persists beyond the custom synthetic settings. 0.0 0.2 0.4 0.6 0.8 1.0 Target FDR level 0 20 40 60 80 100 Number of Selected Individuals (a) Selec…
Figure 27
Figure 27. Figure 27: NSW job-training real-data analysis. Pure real-data deployment analysis on the NSW benchmark. Since true CATE errors are unobserved, the figure provides qualitative policy and selected-subpopulation diagnostics rather than oracle FDR/power. C.4.3 NSW job-training real…
Figure 28
Figure 28. Figure 28: Multi-treatment K = 3. Realized FDR and power versus target α. DCA remains well-calibrated and more powerful than the naive proxy baseline; ρ = 0.25 provides the best power– calibration trade-off in this setting. Overlap-aware stabilization of DR pseudo-outcomes. Repl…
Figure 29
Figure 29. Figure 29: Multi-treatment K = 5 failure under overlap stress. Realized FDR exceeds the nominal line across α even for small denoising strengths, indicating breakdown of calibration/selection under finite-sample overlap degradation in multi-arm settings. group BH with groups ind…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

40 extracted references · 3 linked inside Pith

  1. [1]

    Proceedings of the National Academy of Sciences , volume=

    Metalearners for estimating heterogeneous treatment effects using machine learning , author=. Proceedings of the National Academy of Sciences , volume=

  2. [2]

    ICML , pages=

    Estimating individual treatment effect: generalization bounds and algorithms , author=. ICML , pages=

  3. [3]

    Journal of the American Statistical Association , volume=

    Estimation and inference of heterogeneous treatment effects using random forests , author=. Journal of the American Statistical Association , volume=

  4. [4]

    Biometrika , volume=

    Quasi-oracle estimation of heterogeneous treatment effects , author=. Biometrika , volume=

  5. [5]

    Journal of Economic perspectives , volume=

    The state of applied econometrics: Causality and machine learning , author=. Journal of Economic perspectives , volume=

  6. [6]

    2005 , publisher=

    Algorithmic learning in a random world , author=. 2005 , publisher=

  7. [7]

    Foundations and Trends

    A gentle introduction to conformal prediction and distribution-free uncertainty quantification , author=. Foundations and Trends

  8. [8]

    Journal of the Royal Statistical Society Series B , volume=

    Conformal inference of counterfactuals and individual treatment effects , author=. Journal of the Royal Statistical Society Series B , volume=

Show all 40 references
  1. [9]

    Journal of the American Statistical Association , volume=

    An exact and robust conformal inference methods for predictive inference with misspecified conformal models , author=. Journal of the American Statistical Association , volume=

  2. [10]

    Advances in Neural Information Processing Systems , year=

    Selective classification for deep neural networks , author=. Advances in Neural Information Processing Systems , year=

  3. [11]

    Journal of Machine Learning Research , volume=

    Selection by prediction with conformal p-values , author=. Journal of Machine Learning Research , volume=

  4. [12]

    Advances in Neural Information Processing Systems , volume=

    Conformal Alignment: Knowing When to Trust Foundation Models with Guarantees , author=. Advances in Neural Information Processing Systems , volume=

  5. [13]

    Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=

    Conformalized survival analysis , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2023 , publisher=

  6. [14]

    Journal of the Royal Statistical Society: Series B , volume=

    Controlling the false discovery rate: a practical and powerful approach to multiple testing , author=. Journal of the Royal Statistical Society: Series B , volume=

  7. [15]

    Journal of the American Statistical Association , volume=

    Oracle and adaptive compound decision rules for false discovery rate control , author=. Journal of the American Statistical Association , volume=

  8. [16]

    Journal of Educational Psychology , volume=

    Estimating causal effects of treatments in randomized and nonrandomized studies , author=. Journal of Educational Psychology , volume=

  9. [17]

    2009 , publisher=

    Causality , author=. 2009 , publisher=

  10. [18]

    Biometrika , volume=

    The central role of the propensity score in observational studies for causal effects , author=. Biometrika , volume=

  11. [19]

    Biometrics , volume=

    Doubly robust estimation in missing data and causal inference models , author=. Biometrics , volume=. 2005 , publisher=

  12. [20]

    Journal of the American statistical Association , volume=

    Estimation of regression coefficients when some regressors are not always observed , author=. Journal of the American statistical Association , volume=

  13. [21]

    Electronic Journal of Statistics , volume=

    Towards optimal doubly robust estimation of heterogeneous causal effects , author=. Electronic Journal of Statistics , volume=

  14. [22]

    The Econometrics Journal , volume=

    Double/debiased machine learning for treatment and structural parameters , author=. The Econometrics Journal , volume=

  15. [23]

    Advances in Neural Information Processing Systems , volume=

    Conformalized Quantile Regression , author=. Advances in Neural Information Processing Systems , volume=

  16. [24]

    Journal of Machine Learning Research , volume=

    On the foundations of noise-free selective classification , author=. Journal of Machine Learning Research , volume=

  17. [25]

    Econometrica , volume=

    Causal inference under network interference: A spectral approach , author=. Econometrica , volume=

  18. [26]

    Journal of the American Statistical Association , volume=

    Causal inference under network interference with noise , author=. Journal of the American Statistical Association , volume=. 2021 , note=

  19. [27]

    Advances in Neural Information Processing Systems , volume=

    Conformal prediction under covariate shift , author=. Advances in Neural Information Processing Systems , volume=

  20. [28]

    Biometrika , volume=

    The role of the propensity score in estimating dose-response functions , author=. Biometrika , volume=. 2000 , note=

  21. [29]

    Statistical Science , pages=

    Estimation of causal effects with multiple treatments: a review and new ideas , author=. Statistical Science , pages=. 2017 , note=

  22. [30]

    Journal of Econometrics , volume=

    Robust inference on average treatment effects with possibly more covariates than observations , author=. Journal of Econometrics , volume=. 2015 , note=

  23. [31]

    Journal of Business & Economic Statistics , volume=

    Multiway cluster robust double/debiased machine learning , author=. Journal of Business & Economic Statistics , volume=. 2022 , note=

  24. [32]

    Biometrika , volume=

    Model-free selective inference under covariate shift via weighted conformal p-values , author=. Biometrika , volume=. 2026 , note=

  25. [33]

    arXiv preprint arXiv:2411.17983 , year=

    Optimized conformal selection: Powerful selective inference after conformity score optimization , author=. arXiv preprint arXiv:2411.17983 , year=

  26. [34]

    arXiv preprint arXiv:2508.12085 , year=

    Unified conformalized multiple testing with full data efficiency , author=. arXiv preprint arXiv:2508.12085 , year=

  27. [35]

    Gui, Yu and Jin, Ying and Nair, Yash and Ren, Zhimei , journal=

  28. [36]

    Derandomized novelty detection with

    Bashari, Meysam and Epstein, Aviram and Romano, Yaniv and Sesia, Matteo , booktitle=. Derandomized novelty detection with

  29. [37]

    International Conference on Artificial Intelligence and Statistics , pages=

    Adaptive, distribution-free prediction intervals for deep neural networks , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2020 , note=

  30. [38]

    arXiv preprint arXiv:2501.18991 , year=

    Optimal Transport-based Conformal Prediction , author=. arXiv preprint arXiv:2501.18991 , year=

  31. [39]

    Journal of Computational and Graphical Statistics , volume=

    Bayesian nonparametric modeling for causal inference , author=. Journal of Computational and Graphical Statistics , volume=. 2011 , note=

  32. [40]

    Journal of the American Statistical Association , volume=

    Decomposing treatment effect variation , author=. Journal of the American Statistical Association , volume=. 2019 , note=

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.