REVIEW 2 major objections 3 minor 2 cited by
On the Role of Surrogates in Conformal Inference of Individual Causal Effects
T0 review · 2 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Surrogates shorten conformal prediction intervals for individual treatment effects without losing coverage, and the gain equals the surrogates' added predictiveness.
desk verdict A serious surrogate-assisted conformal extension with a load-bearing identification gap for post-treatment surrogates; the efficiency-gain claims need revision or a new assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the efficient influence function (the gradient whose variance is the semiparametric efficiency lower bound) $\psi^{(S2)}_1$ for the $(1-\alpha)$-quantile $r_{\alpha,1}$ of the non-conformity score $R(W,u\{Y(1)\})$ under covariate shift. It has three mean-zero terms: a covariate term $D(1-A)\{m_1(r,X)-(1-\alpha)\}$; a surrogate-augmentation term $A\pi_A(X)e_D(X,0)\{\tilde m_1(r,X,S)-m_1(r,X)\}$ that vanishes when the surrogate adds no predictive information; and an outcome term $AD\pi_A(X)e_D(X,0)e_D(X,1)^{-1}\{\mathbf 1(R_1<r)-\tilde m_1(r,X,S)\}$. A parallel influence function (Theorem 4) handles the nesting quantile $r_{\gamma,C}$ of the pseudo-outcome intervals $C_i$, and Theorem 3 shows the coverage slack decomposes into an $O(|\mathcal I_2|^{-1/2})$ sampling term plus products of nuisance estimation errors. The non-conformity score itself is a conformalized quantile residual, and the nuisance functions—propensity scores $e_A$, $e_D$ and conditional score CDFs $m_a$, $\tilde m_a$—are estimated once at a single initial quantile via localized debiased machine learning, so one binary classification fit per nuisance suffices.
What would settle it
Simulate full potential outcomes with the surrogate under one treatment driving the outcome under the other treatment (say $Y(0)=\beta_0+\beta_1 X+\beta_2 S(1)+\epsilon$) while keeping Assumptions 3–5 satisfied, run SCIENCE, and measure the source-data coverage of $u\{Y(0)\}$ among treated units; coverage falling below $1-\alpha$ as $\beta_2$ grows would falsify the identification claim.
Extended reading notes
Core claim
The paper's central claim is that the quantiles of counterfactual non-conformity scores—the ingredients of conformal prediction intervals for individual treatment effects—can be estimated semi-parametrically efficiently by the influence functions of Theorems 2 and 4, yielding intervals that achieve the nominal coverage level $1-\alpha$ both marginally and group-conditionally, on source data where one potential outcome is observed and, through a nested conformal step on pseudo-outcome intervals $C_i$, on target data where primary outcomes are missing. The efficiency bound for the surrogate-assisted setting drops below the no-surrogate bound by exactly $\mathbb{E}\big[\frac{1-e_D(X,1)}{e_D(X,1)}\,\frac{e_D(X,0)^2\{1-e_A(X)\}^2}{e_A(X)}\,\mathrm{var}\{\tilde m_1(r_{\alpha,1},X,S)\,|\,X\}\big]$, so the value of a surrogate is precisely its residual predictiveness for the primary outcome given covariates, amplified by missingness. The coverage slack in Theorem 3 is a product of nuisance estimation errors, which the paper calls rate double robustness, so consistent nonparametric estimators of the propensity scores and conditional score distributions suffice for nominal coverage.
Load-bearing premise
The load-bearing premise is that the surrogate observed under the actual treatment is non-informative about the counterfactual outcome being predicted, conditional on baseline covariates; the paper's stated assumptions do not identify the distribution of the counterfactual non-conformity score evaluated at that observed surrogate, so this unstated cross-world independence must hold for the intervals to be calibrated.
Editorial extensions
If this is right
- SCIENCE's prediction intervals achieve the nominal $(1-\alpha)$ coverage marginally and within prespecified groups, on the source data (one potential outcome observed) and on the target data (no primary outcomes).
- Surrogates reduce interval width exactly when they are observed alongside missing primary outcomes: the Setting-2 efficiency bound beats Settings 1 and 3 by the surrogate-predictiveness variance $\mathrm{var}\{\tilde m_a(r,X,S)\,|\,X\}$ scaled by the missingness odds, and surrogates observed only in the source data give no gain.
- The coverage slack is a product of nuisance estimation errors, so any consistent estimators of the propensity scores and conditional score distributions—including flexible nonparametric learners—suffice for nominal coverage asymptotically.
- In the Moderna COVE trial, Day 29 neutralizing and binding antibody surrogates shorten the intervals for the Day 57 individual vaccine effect on antibody titer, with both markers together giving the shortest intervals.
- Because only a marginal contrast (e.g., $u\{Y(1)\}-u\{Y(0)\}$) is required, the same machinery covers ITEs, log-risk-ratio contrasts, and utility-weighted ordinal contrasts.
Reading between the lines
- Editorial inference: the identification argument silently requires a cross-world condition—the surrogate measured under one treatment must carry no information about the outcome under the other treatment, given baseline covariates—because the counterfactual score is evaluated at the observed post-treatment surrogate; this condition is untested, and the paper's stated assumptions do not imply it.
- Editorial inference: since no statistical-surrogacy assumption is imposed, arbitrary machine-learning predictions can serve as surrogates, making SCIENCE a conformal analogue of prediction-powered inference for individual effects, with the efficiency gain sized by the predictor's accuracy.
- Testable extension: one could verify the cross-world condition empirically in placebo-controlled trials with a surrogate measured in both arms, by testing whether $S(1)$ predicts $Y(0)$ within strata of $X$; a nonzero association flags intervals whose counterfactual half may be miscalibrated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops SCIENCE, a conformal prediction framework for individual causal effects such as Y(1)-Y(0) under covariate shift between source and target data, with surrogate outcomes S observed in addition to baseline covariates X. For each treatment arm a, the method estimates the (1-alpha)-quantile of a non-conformity score R(W, u_a(Y(a))) using efficient influence functions, constructs prediction intervals for counterfactual outcomes, and then uses a nested conformal step to build intervals for target data with missing primary outcomes. Theorems 2-4 and Corollaries 1-2 provide semiparametric efficiency bounds and product-bias rate double robustness, and simulations plus a COVE trial analysis are used to show shorter intervals when surrogates are incorporated.
Significance. If the identification of the counterfactual non-conformity score were valid, the paper would be a meaningful contribution: the EIF derivations in the supplement are careful, Theorem 3 gives an explicit PAC-type bound with product-bias terms, the efficiency-gain formulas in Corollaries 1-2 are nontrivial, and the open-source R implementation plus the COVE application make the work practically oriented. The baseline results with W=X provide a sound doubly robust conformal procedure under covariate shift. However, the central surrogate-assisted claim depends on an identification step that is not justified for post-treatment surrogates; since the paper explicitly motivates post-treatment S and its main simulations and data analysis use such surrogates, the distinctive contribution is not established.
major comments (2)
- [Theorem 1, Setting 2; Section 2.1] The display in Theorem 1 for Setting 2 equates E_W{P_Y(R_a < r_{\alpha,a} | A=a,D=1,W) | A=1-a,D=1} with E_X[E_S{P_Y(R_a < r_{\alpha,a} | A=a,D=1,S,X) | A=a,X} | A=1-a,D=1]. This step replaces the target surrogate distribution S(1-a)|X,A=1-a with the training surrogate distribution S(a)|X,A=a. Assumption 3 gives A independent of {Y(a),S(a)} given X for each a separately, which does not imply S(0) and S(1) have the same distribution given X and does not identify cross-world quantities such as (Y(0),S(1)). Since Section 2.1 explicitly allows S to be post-treatment, and the COVE analysis in Section 7 uses Day 29 antibody markers measured after the first dose, the quantile r_{\alpha,a} of R(X,S(1-a),Y(a)) is not identified under the stated assumptions. If W is instead interpreted as (X,S(a)), the identification is formally correct, but then the score is not computable for A=1-a individuals; Section 4.3 defines R_{a,i} only for A_i=a using observed S_i, so the estimator does not target the quantity in the theorem. Corollaries 1-2 and the efficiency claims are therefore unsupported for post-treatment surrogates.
- [Section 4.3; Section 6] The implementation estimates a different quantile from the one needed for the coverage guarantee. The estimators \hat r^{(S2)}_{\alpha,a} are formed from source scores R_{a,i} that use W_i=(X_i,S_i) with S_i=S_i(A_i), and the weights \hat \pi_A(X) shift only the X distribution from arm a to arm 1-a; the surrogate distribution remains S(a)|X,A=a. The target event for an A=1-a individual, however, is R(X,S(1-a),u_a(Y(a))), whose surrogate component is not available in the source data. The simulation DGP in Section 6 draws S(0) and S(1) with different means, so the empirical coverage reported in Figures 4-5 evaluates the wrong score and does not validate the method. The same misspecification propagates to the nested target-data intervals in Section 5.
minor comments (3)
- [Section 2.1] The paper warns that post-treatment S cannot be used as a covariate without additional assumptions, but this warning is never reconciled with the use of S in the non-conformity score; the paper should either state the additional assumption or remove S from the score for counterfactual outcomes.
- [Section 3.2] The notation W is ambiguous in the target distribution P(W,Y(a)|A=1-a,D=1): it is not specified whether W contains S(a) or S(1-a), and this ambiguity is central to the identification gap; a precise definition of which potential surrogate appears in R_a should be given.
- [Table 5] The columns of Table 5 are difficult to read because the coverage and width results are interleaved; separating them into two sub-tables or using a clearer longtable format would improve readability.
Circularity Check
No circular reasoning found: the EIF, coverage, and efficiency-gain claims are derived from the stated semiparametric model without fitted-to-target reductions. The paper's load-bearing gap is an identification issue for post-treatment surrogates in Theorem 1 (Setting 2), which is a correctness risk, not a circular step.
full rationale
The derivation chain is self-contained and externally anchored rather than circular. The efficient influence functions in Theorems 2 and 4 are obtained by standard pathwise differentiation (Appendix A.1) of the stated semiparametric model; the double-robust coverage claims in Theorem 3 follow from explicit second-order remainder decompositions (Appendix A.2) whose empirical-process bound is adapted from Yang et al. (2024, Lemma 8), an external source; and the efficiency-gain formulas in Corollaries 1-2 are direct variance algebra (Appendices A.3 and A.5). No parameter is fitted to a target and then relabeled as a prediction: the estimator br_{alpha,a} is a Z-estimator solving sum_i psi(r) >= 0 with nuisance functions learned on an independent fold, and coverage is then proven rather than imposed. The self-citations (Gilbert et al. 2022, 2024; Han et al. 2022) appear only in contextual roles (immune-correlates literature and the COVE risk-score construction) and carry none of the load-bearing argument. The genuine problem in this paper is an identification gap, not circularity. Section 2.1 explicitly warns that surrogates 'can be post-treatment variables' and that S 'is observed only under the actual treatment but not under the counterfactual treatment,' yet the Setting 2 display in Theorem 1 replaces the target-arm outer expectation over W with the training-arm conditional expectation E_S{... | A=a, X} under A=1-a, which is valid only if S(0) and S(1) are equal in distribution given X. Assumption 3 (A independent of {Y(a), S(a)} given X, per arm) does not imply this cross-arm equality, so the target quantile r_{alpha,a} of the score R(X, S(1-a), Y(a)) under A=1-a is not identified for post-treatment surrogates, and the claimed surrogate efficiency gains (Corollaries 1-2 and the COVE application using Day 29 markers) inherit this gap. This is a substantive mathematical flaw for referees to weigh, but it is not an instance of the paper's derivation reducing to its own inputs, so the circularity score is 1.
Assumptions & free parameters
assumptions (7)
- domain assumption Assumption 1: theta_i = u(Y_i(1)) - u(Y_i(0)) (marginal contrast estimand)
- domain assumption Assumption 2: overlap c0 <= P(A=1|x), P(D=1|x,a) <= 1-c0
- domain assumption Assumption 3: A independent of {Y(a), S(a)} given X
- domain assumption Assumption 4: D independent of S(a) given X, A
- domain assumption Assumption 5: D independent of Y(a) given S(a), X, A (MAR)
- standard math Regularity conditions A1-A2: bounded nuisance estimators and monotone estimated CDFs
- ad hoc to paper Implicit availability of the counterfactual surrogate in the score for counterfactual intervals
Cite this review
Pith. "Pith review of On the Role of Surrogates in Conformal Inference of Individual Causal Effects." pith.science (2026). https://pith.science/paper/LZ7MIMDL
@misc{pith2026241212365,
author = {Pith},
title = {Pith review of: On the Role of Surrogates in Conformal Inference of Individual Causal Effects},
year = {2026},
howpublished = {\url{https://pith.science/paper/LZ7MIMDL}},
note = {Machine review of arXiv:2412.12365}
}
read the original abstract
Learning the Individual Treatment Effect (ITE) is essential for personalized decision-making, yet causal inference has traditionally focused on aggregated treatment effects. While integrating conformal prediction with causal inference can provide valid uncertainty quantification for ITEs, the resulting prediction intervals are often excessively wide, limiting their practical utility. To address this limitation, we introduce \underline{S}urrogate-assisted \underline{C}onformal \underline{I}nference for \underline{E}fficient I\underline{N}dividual \underline{C}ausal \underline{E}ffects (SCIENCE), a framework designed to construct more efficient prediction intervals for ITEs. SCIENCE accommodates the covariate shifts between source data and target data and applies to various data configurations, including semi-supervised and surrogate-assisted semi-supervised learning. Leveraging semi-parametric efficiency theory, SCIENCE produces rate double-robust prediction intervals under mild rate convergence conditions, permitting the use of flexible non-parametric models to estimate nuisance functions. We quantify efficiency gains by comparing semi-parametric efficiency bounds with and without the surrogates. Simulation studies demonstrate that our surrogate-assisted intervals offer substantial efficiency improvements over existing methods while maintaining valid group-conditional coverage. Applied to the phase 3 Moderna COVE COVID-19 vaccine trial, SCIENCE illustrates how multiple surrogate markers can be leveraged to generate more efficient prediction intervals.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 2 Pith papers
-
Towards the Efficient Inference by Incorporating Automated Computational Phenotypes under Covariate Shift
Doubly robust, semiparametrically efficient estimators that incorporate automated computational phenotypes (ACPs) into semi-supervised inference under covariate shift, with explicit efficiency gains driven by ACPs in ...
-
Targeted Data Fusion for Region-Specific Survival Effects in the AMP HIV Prevention Trials
A federated survival estimator for multi-site trials that adaptively discards incompatible sites, proving no loss of efficiency relative to using only the target site.
Reference graph
Works this paper leans on
-
[1]
Alonso, A., Van der Elst, W., Molenberghs, G., Buyse, M. and Burzykowski, T. (2016). An information-theoretic approach for the evaluation of surrogate endpoints based on causal inference, Biometrics 72: 669–677. Angelopoulos, A. N., Bates, S., Fannjiang, C., Jordan, M. I. and Zrnic, T. (2023). Prediction-powered inference, Science 382: 669–674. URL: https...
arXiv 2016
-
[2]
A.6 Additional technical details A.6.1 Proof of Lemma 1 The proof of Lemma 2.2 in Kallus and Mao (2020) shows that D ⊥ {Y (a), S(a)} |X, A under Assumptions 4 and
work page 2020
-
[5]
A.6.2 Proof of Lemma 2 On the one hand, we can show that P (R1 ≤ rα,1 | A = 0, D=
Under Assumption 3, we have ( A, D) ⊥ {Y (a), S(a)} |X, which proves Condition (a), that is, D ⊥ {Y (a), S(a)} |X, and Conditions (b) and (c) A ⊥ {Y (a), S(a))} |X, Dby Theorem 17.2 in Wasserman (2013). A.6.2 Proof of Lemma 2 On the one hand, we can show that P (R1 ≤ rα,1 | A = 0, D=
work page 2013
-
[7]
Here, we use the fact that PI2{ bψ(S2) 1 (rα,1, W)} ≥0 by definition
≥ 0 + I1 + I2. Here, we use the fact that PI2{ bψ(S2) 1 (rα,1, W)} ≥0 by definition. The term I1 is negligible if ψ(S2) 1 (rα,1, W) belongs to a Donsker class (Van der Vaart, 2000). Even if the Donsker condition is not met, the sample-splitting procedure in Chernozhukov et al. (2018) can be used to assure that I1 is negligible, where the first data fold I...
work page 2018
-
[8]
Tchetgen, E. J. T. and VanderWeele, T. J. (2012). On causal inference in the presence of interference, Statistical Methods in Medical Research 21: 55–75. Van der Vaart, A. W. (2000). Asymptotic Statistics , Vol. 3, Cambridge University Press. Vovk, V., Gammerman, A. and Shafer, G. (2005). Algorithmic learning in a random world , Vol. 29, Springer. Vovk, V...
work page 2012
-
[10]
The proof of Lemma A.1 is adapted from Theorem 3 in Yang et al
r log(1/δ) + 1 N | I1 ! ≥ 1 − δ. The proof of Lemma A.1 is adapted from Theorem 3 in Yang et al. (2024). Specifically, we expand the numerator of I1 into the following four parts, P{ bψ(S2) 1 (rα,1, W)} −PI2{ bψ(S2) 1 (rα,1, W)} = " 1 |I2| X i∈I2 AiDibπA(Xi)beD(Xi, 0)1(R1 < rα,1) beD(Xi,
work page 2024
-
[11]
) ≥ t) ≤ exp − 2t2 N which leads to P R(S2) 4 ≥ (1 − α) s 1 2N log 1 δ ! ≤ δ. (14) Combining (11), (12), (13), and (14) together with the help of union bound, we can show that with probability larger than 1 − δ, sup r |R(S2) 1 (rα,1) +R(S2) 2 (rα,1) +R(S2) 3 (rα,1) +R(S2) 4 | ≲ π0e0(e−1 0 + m0 + e−1 0 ˜m0) s log(1/δ) + 1 |I2| , which completes the proof f...
work page 2024
-
[12]
e2 D(X, 1){eA(X)}2 1 − eA(X) var{ ˜m0(rα,0, X, S) | X, A= 0} , where V (S1) 0 = var{ψ(S1) 0 (rα,0, W; m, eD, πA)}, V (S2) 0 = var{ψ(S2) 0 (rα,0, W; m, ˜m, eD, πA)}. 39 B Additional Simulation Studies In this section, we assess the group-conditional performance of the prediction intervals produced by the SCIENCE framework for categorical outcomes. The data...
work page 2023
Show all 12 references
-
[32]
and Cand` es, E
Sesia, M. and Cand` es, E. J. (2020). A comparison of some conformal quantile regression methods, Stat 9: e261. Sugiyama, M., Krauledat, M. and M¨ uller, K.-R. (2007). Covariate shift adaptation by importance weighted cross validation., Journal of Machine Learning Research
2020
-
[36]
Elliott, M. R. (2023). Surrogate endpoints in clinical trials, Annual Review of Statistics and its Application 10: 75–96. 24 FDA (1992). Accelerated approval of new drugs for serious or life-threatening illnesses, Federal Register 57: 58942–58960. FDA (2021). Accelerated appro...
2023 arXiv
-
[96]
and Cand` es, E
Lei, L. and Cand` es, E. J. (2021). Conformal inference of counterfactuals and individual treatment effects, Journal of the Royal Statistical Society Series B: Statistical Methodol- ogy 83: 911–938. Li, Y., Taylor, J. M. and Elliott, M. R. (2010). A bayesian approach to surrog...
2021 arXiv
-
[5749]
and Mathew, T
Krishnamoorthy, K. and Mathew, T. (2009). Statistical tolerance regions: theory, applica- tions, and computation , John Wiley & Sons. Kuchibhotla, A. K. and Berk, R. A. (2023). Nested conformal prediction sets for classifica- tion with applications to probation data, The Annal...
2009
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.