Pith. sign in

REVIEW 4 cited by

Orthogonal Causal Calibration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01933 v2 pith:JVHY37TC submitted 2024-06-04 stat.ML cs.LGmath.STstat.MEstat.TH

Orthogonal Causal Calibration

classification stat.ML cs.LGmath.STstat.MEstat.TH
keywords calibrationcausalalgorithmcalibratingconditionaleffectsestimatorsloss
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Estimates of heterogeneous treatment effects such as conditional average treatment effects (CATEs) and conditional quantile treatment effects (CQTEs) play an important role in real-world decision making. Given this importance, one should ensure these estimates are calibrated. While there is a rich literature on calibrating estimators of non-causal parameters, very few methods have been derived for calibrating estimators of causal parameters, or more generally estimators of quantities involving nuisance parameters. In this work, we develop general algorithms for reducing the task of causal calibration to that of calibrating a standard (non-causal) predictive model. Throughout, we study a notion of calibration defined with respect to an arbitrary, nuisance-dependent loss $\ell$, under which we say an estimator $\theta$ is calibrated if its predictions cannot be changed on any level set to decrease loss. For losses $\ell$ satisfying a condition called universal orthogonality, we present a simple algorithm that transforms partially-observed data into generalized pseudo-outcomes and applies any off-the-shelf calibration procedure. For losses $\ell$ satisfying a weaker assumption called conditional orthogonality, we provide a similar sample splitting algorithm the performs empirical risk minimization over an appropriately defined class of functions. Convergence of both algorithms follows from a generic, two term upper bound of the calibration error of any model. We demonstrate the practical applicability of our results in experiments on both observational and synthetic data. Our results are exceedingly general, showing that essentially any existing calibration algorithm can be used in causal settings, with additional loss only arising from errors in nuisance estimation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Doubly cross-fit debiased machine learning of heterogeneous treatment effects under principal stratification

    stat.ME 2026-06 unverdicted novelty 7.0

    Proposes a doubly cross-fit doubly robust machine learner for conditional principal causal effects under principal ignorability with odds ratio sensitivity, with limit theory and application to an acute lung injury trial.

  2. Calibeating Prediction-Powered Inference

    stat.ML 2026-04 unverdicted novelty 7.0

    Post-hoc calibration of miscalibrated black-box predictions on a labeled sample improves efficiency of prediction-powered inference for semisupervised mean estimation.

  3. Bellman Calibration for $V$-Learning in Offline Reinforcement Learning

    stat.ML 2025-12 unverdicted novelty 7.0

    Bellman calibration supplies a new reliability criterion and post-hoc recalibration method for value functions in offline RL, with finite-sample guarantees at one-dimensional nonparametric rates that avoid Bellman com...

  4. Doubly cross-fit debiased machine learning of heterogeneous treatment effects under principal stratification

    stat.ME 2026-06 conditional novelty 6.5

    A doubly cross-fit DR sieve learner identifies and estimates within-stratum heterogeneous treatment effects under principal ignorability and odds-ratio sensitivity, with oracle rates and uniform bands.