{"id":"4c475fa8-7ea8-4df6-8c90-60b77a6ad50a","arxiv_id":"2412.12365","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Surrogate assisted conformal inference yields shorter valid prediction intervals for individual treatment effects, with the new insight that surrogates help only when measured in the target data.","lead":"This paper proposes SCIENCE, a conformal inference method that uses surrogate markers to build narrower prediction intervals for individual treatment effects while keeping the promised coverage. It matters because individual-level causal intervals are often too wide to guide medical decisions, and surrogates such as antibody levels are widely available.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-treatment surrogates in non-conformity scores are not identified: Theorem 1 silently equates the target and training surrogate distributions, which Assumption 3 does not justify.","rationale":"The reader's weakest assumption correctly identifies the load-bearing gap: the non-conformity score for a counterfactual outcome is evaluated at the observed surrogate from the assigned treatment arm, while calibration data come from the opposite arm, and no stated assumption links S(1) to S(0) or to Y(0). My pass sharpens this into a concrete technical step: in Theorem 1, Setting 2, the inner expectation is explicitly taken over the conditional distribution of S given A=a, whereas the outer target expectation conditions on A=1−a; equating the two requires S(0)|A=0,X and S(1)|A=1,X to have the same distribution, which Assumption 3 does not deliver. This is not merely a notational concern: it is the exact mechanism by which surrogates enter the conformal score, and it underpins the efficiency-gain formulas in Corollaries 1 and 2. The paper is otherwise careful about the post-treatment nature of S, and the simulation study uses S(a) generated separately for each arm, so the gap is substantive rather than stylistic. I do not escalate to rejection because a viable repair exists: either restrict S to pre-treatment baseline variables, in which case S(0)=S(1) and the identification goes through, or redefine the non-conformity score to exclude S and analyze surrogate efficiency separately. Either repair changes the contribution's scope, so the conditional verdict is appropriate. I also note the separate sign inconsistency in the main-text Corollary 1 relative to the supplement, but I did not use it as the primary concern because the identification gap is more fundamental.","tokens_in":34370,"tokens_out":8321,"duration_ms":77253,"concrete_test":"Construct a data-generating process satisfying Assumptions 3–5 with post-treatment surrogates: take X~N(0,1), A independent of all potential outcomes given X, S(0)=X+ε_0, S(1)=X+ε_1, Y(a)=S(a)+ν_a, with ε_0, ε_1, ν_0, ν_1 mutually independent and D independent of everything. Compute the true (1−α)-quantile of R(X, S(1), Y(0)) under A=1, D=1 and compare it with the quantity that Theorem 1 estimates using A=0 data built from S(0). If these disagree, run the SCIENCE procedure on a large sample and measure empirical coverage of the resulting interval for Y(0) among treated units; rejection of nominal coverage confirms that the identification step in Theorem 1 is invalid for post-treatment surrogates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on identifying rα,a, the quantile of R(W, u_a{Y(a)}) under A=1−a. For an individual with A=1, the score uses the observed surrogate S(1) together with the counterfactual outcome Y(0): R(X, S(1), Y(0)). Training data from A=0 provide R(X, S(0), Y(0)). The first identification display in Setting 2 of Theorem 1 writes 1−α = E_W{P_Y(R_a < r | A=a, D=1, W) | A=1−a, D=1} and then equals E_X[E_S{P_Y(R_a < r | A=a, D=1, S, X) | A=a, X} | A=1−a, D=1]. The step from the first expression to the second replaces the target surrogate distribution, S(1−a) given X and A=1−a, with the training surrogate distribution, S(a) given X and A=a. This is only valid if S(0)|X,A=0 and S(1)|X,A=1 are equal in distribution. Assumption 3 states A ⊥ {Y(a), S(a)} | X for each a separately; it does not imply S(0) =d S(1) given X, nor does it identify the cross-world pair {Y(0), S(1)}. Since Section 2.1 explicitly allows S to be post-treatment, and the COVE application uses Day 29 markers measured after the first dose, the observed surrogate under the assigned arm and the surrogate under the opposite arm are generally distinct. Therefore the quantile rα,a is not identified under the stated assumptions. Every efficiency result in Corollaries 1 and 2 is stated for scores that include S, so the claimed surrogate gains are unsupported for post-treatment surrogates; if W is instead taken to be only X, the surrogates play no role and the method reduces to the existing doubly robust conformal machinery without surrogate efficiency gains.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops SCIENCE, a conformal prediction framework for individual causal effects such as Y(1)-Y(0) under covariate shift between source and target data, with surrogate outcomes S observed in addition to baseline covariates X. For each treatment arm a, the method estimates the (1-alpha)-quantile of a non-conformity score R(W, u_a(Y(a))) using efficient influence functions, constructs prediction intervals for counterfactual outcomes, and then uses a nested conformal step to build intervals for target data with missing primary outcomes. Theorems 2-4 and Corollaries 1-2 provide semiparametric efficiency bounds and product-bias rate double robustness, and simulations plus a COVE trial analysis are used to show shorter intervals when surrogates are incorporated.","tokens_in":34756,"tokens_out":16683,"duration_ms":147697,"significance":"If the identification of the counterfactual non-conformity score were valid, the paper would be a meaningful contribution: the EIF derivations in the supplement are careful, Theorem 3 gives an explicit PAC-type bound with product-bias terms, the efficiency-gain formulas in Corollaries 1-2 are nontrivial, and the open-source R implementation plus the COVE application make the work practically oriented. The baseline results with W=X provide a sound doubly robust conformal procedure under covariate shift. However, the central surrogate-assisted claim depends on an identification step that is not justified for post-treatment surrogates; since the paper explicitly motivates post-treatment S and its main simulations and data analysis use such surrogates, the distinctive contribution is not established.","major_comments":[{"comment":"The display in Theorem 1 for Setting 2 equates E_W{P_Y(R_a < r_{\\alpha,a} | A=a,D=1,W) | A=1-a,D=1} with E_X[E_S{P_Y(R_a < r_{\\alpha,a} | A=a,D=1,S,X) | A=a,X} | A=1-a,D=1]. This step replaces the target surrogate distribution S(1-a)|X,A=1-a with the training surrogate distribution S(a)|X,A=a. Assumption 3 gives A independent of {Y(a),S(a)} given X for each a separately, which does not imply S(0) and S(1) have the same distribution given X and does not identify cross-world quantities such as (Y(0),S(1)). Since Section 2.1 explicitly allows S to be post-treatment, and the COVE analysis in Section 7 uses Day 29 antibody markers measured after the first dose, the quantile r_{\\alpha,a} of R(X,S(1-a),Y(a)) is not identified under the stated assumptions. If W is instead interpreted as (X,S(a)), the identification is formally correct, but then the score is not computable for A=1-a individuals; Section 4.3 defines R_{a,i} only for A_i=a using observed S_i, so the estimator does not target the quantity in the theorem. Corollaries 1-2 and the efficiency claims are therefore unsupported for post-treatment surrogates.","section":"Theorem 1, Setting 2; Section 2.1"},{"comment":"The implementation estimates a different quantile from the one needed for the coverage guarantee. The estimators \\hat r^{(S2)}_{\\alpha,a} are formed from source scores R_{a,i} that use W_i=(X_i,S_i) with S_i=S_i(A_i), and the weights \\hat \\pi_A(X) shift only the X distribution from arm a to arm 1-a; the surrogate distribution remains S(a)|X,A=a. The target event for an A=1-a individual, however, is R(X,S(1-a),u_a(Y(a))), whose surrogate component is not available in the source data. The simulation DGP in Section 6 draws S(0) and S(1) with different means, so the empirical coverage reported in Figures 4-5 evaluates the wrong score and does not validate the method. The same misspecification propagates to the nested target-data intervals in Section 5.","section":"Section 4.3; Section 6"}],"minor_comments":[{"comment":"The paper warns that post-treatment S cannot be used as a covariate without additional assumptions, but this warning is never reconciled with the use of S in the non-conformity score; the paper should either state the additional assumption or remove S from the score for counterfactual outcomes.","section":"Section 2.1"},{"comment":"The notation W is ambiguous in the target distribution P(W,Y(a)|A=1-a,D=1): it is not specified whether W contains S(a) or S(1-a), and this ambiguity is central to the identification gap; a precise definition of which potential surrogate appears in R_a should be given.","section":"Section 3.2"},{"comment":"The columns of Table 5 are difficult to read because the coverage and width results are interleaved; separating them into two sub-tables or using a clearer longtable format would improve readability.","section":"Table 5"}],"recommendation":"reject","confidential_remarks":"The paper's technical machinery (EIFs, rate double robustness) appears sound for the no-surrogate setting and for a correctly targeted surrogate setting, but the main contribution, efficiency gains from post-treatment surrogates, is not identified under the stated assumptions. If the authors add a credible assumption such as S(0)=S(1), i.e., effectively pre-treatment surrogates, the article would be a valid but substantially narrower contribution; under the current scope and application, rejection is warranted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful part. The paper extends Yang et al.'s doubly robust conformal framework to surrogate-assisted settings, derives EIFs for the relevant quantiles, and characterizes when surrogates help: they pay off only when observed in the target data. That message is clear and matches the numerical results. The supplement's EIF algebra is careful, Theorem 3's product-bias coverage bound is serviceable, and the COVE analysis is a responsible application with an R package.\n\nThe problem is the identification of the counterfactual non-conformity score in Setting 2. The score is R(W, u_a{Y(a)}) with W=(X,S), and S is explicitly allowed to be post-treatment. For an A=1 individual, the target score uses the observed S(1) with the counterfactual Y(0), while the training data from A=0 use S(0) and Y(0). Theorem 1's Setting 2 formula first writes the target expectation as E_W{P_Y(R_a<r | A=a,D=1,W) | A=1-a,D=1}, then replaces the outer distribution of W|A=1-a with an inner expectation over S|A=a,X. That step assumes S(1-a)|X has the same distribution as S(a)|X, which Assumption 3 does not give. Assumption 3 says A is independent of {Y(a),S(a)} given X, for each a separately; it doesn't equate S(0) and S(1). The COVE example measures Day 29 markers after vaccine, so this isn't a mere technicality. A simple counterexample (independent A, S(0)~N(0,1), S(1)~N(1,1), Y(0)=S(0)+ε) shows the identification fails. Since Corollary 1's efficiency gain is derived from this identification, the claimed surrogate gains for source-data intervals are unsupported. The second-stage target inference using S in the pseudo-outcome score might survive, but the paper doesn't separate it from the flawed step.\n\nThere's also a sign error: the main-text Corollary 1 writes V(S2)-V(S1) as a positive quantity, while the supplement and the surrounding text derive V(S1)-V(S2) as that positive quantity, i.e., surrogates reduce variance. The text describes a gain, so the formula has the wrong sign.\n\nThis is a major, load-bearing problem, but the paper deserves serious referee time. The core machinery for Setting 1 is fine, and the Setting 2 idea may be fixable by assuming pre-treatment surrogates or by rebuilding the scores on X only for counterfactuals, with S used only in the target-data stage. A good referee could help the authors find the right repair. I'd send it out.","headline":"A serious surrogate-assisted conformal extension with a load-bearing identification gap for post-treatment surrogates; the efficiency-gain claims need revision or a new assumption.","tokens_in":35266,"tokens_out":8695,"would_cite":false,"duration_ms":76798,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G15","62D05","62G20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Surrogates shorten conformal prediction intervals for individual treatment effects without losing coverage, and the gain equals the surrogates' added predictiveness.","keywords":["surrogate outcomes","conformal prediction","individual treatment effects","semiparametric efficiency","covariate shift","double robustness","semi-supervised learning","prediction intervals"],"falsifier":"Simulate full potential outcomes with the surrogate under one treatment driving the outcome under the other treatment (say $Y(0)=\\beta_0+\\beta_1 X+\\beta_2 S(1)+\\epsilon$) while keeping Assumptions 3–5 satisfied, run SCIENCE, and measure the source-data coverage of $u\\{Y(0)\\}$ among treated units; coverage falling below $1-\\alpha$ as $\\beta_2$ grows would falsify the identification claim.","tokens_in":34163,"feed_emoji":"🎯","tokens_out":20277,"duration_ms":155672,"temperature":0.7,"pith_summary":"The paper claims that conformal prediction intervals for individual treatment effects can be made materially shorter by bringing in surrogate outcomes—biomarkers, intermediate variables, or machine-learning predictions—without sacrificing nominal coverage. The proposed framework, SCIENCE, covers source data with full outcomes and target data with only covariates (or covariates plus surrogates), and guarantees marginal and group-conditional coverage under covariate shift. Using semiparametric efficiency theory, the paper derives efficient influence functions for the quantiles of the non-conformity scores and shows the resulting intervals are rate double-robust: coverage bias is a product of nuisance-estimation errors, so flexible machine-learning estimators suffice. A central result is that surrogates pay off exactly when they help impute missing primary outcomes; surrogates observed only where primary outcomes are already measured give no efficiency gain. In the Moderna COVE COVID-19 vaccine trial, Day 29 antibody surrogates shorten the prediction intervals for Day 57 individual vaccine effects.","feed_headline":"Surrogates shrink individual-treatment-effect intervals","feed_subtitle":"A conformal framework keeps coverage valid and shows surrogates pay off only when they fill in missing outcomes.","key_machinery":"The engine is the efficient influence function (the gradient whose variance is the semiparametric efficiency lower bound) $\\psi^{(S2)}_1$ for the $(1-\\alpha)$-quantile $r_{\\alpha,1}$ of the non-conformity score $R(W,u\\{Y(1)\\})$ under covariate shift. It has three mean-zero terms: a covariate term $D(1-A)\\{m_1(r,X)-(1-\\alpha)\\}$; a surrogate-augmentation term $A\\pi_A(X)e_D(X,0)\\{\\tilde m_1(r,X,S)-m_1(r,X)\\}$ that vanishes when the surrogate adds no predictive information; and an outcome term $AD\\pi_A(X)e_D(X,0)e_D(X,1)^{-1}\\{\\mathbf 1(R_1<r)-\\tilde m_1(r,X,S)\\}$. A parallel influence function (Theorem 4) handles the nesting quantile $r_{\\gamma,C}$ of the pseudo-outcome intervals $C_i$, and Theorem 3 shows the coverage slack decomposes into an $O(|\\mathcal I_2|^{-1/2})$ sampling term plus products of nuisance estimation errors. The non-conformity score itself is a conformalized quantile residual, and the nuisance functions—propensity scores $e_A$, $e_D$ and conditional score CDFs $m_a$, $\\tilde m_a$—are estimated once at a single initial quantile via localized debiased machine learning, so one binary classification fit per nuisance suffices.","core_discovery":"The paper's central claim is that the quantiles of counterfactual non-conformity scores—the ingredients of conformal prediction intervals for individual treatment effects—can be estimated semi-parametrically efficiently by the influence functions of Theorems 2 and 4, yielding intervals that achieve the nominal coverage level $1-\\alpha$ both marginally and group-conditionally, on source data where one potential outcome is observed and, through a nested conformal step on pseudo-outcome intervals $C_i$, on target data where primary outcomes are missing. The efficiency bound for the surrogate-assisted setting drops below the no-surrogate bound by exactly $\\mathbb{E}\\big[\\frac{1-e_D(X,1)}{e_D(X,1)}\\,\\frac{e_D(X,0)^2\\{1-e_A(X)\\}^2}{e_A(X)}\\,\\mathrm{var}\\{\\tilde m_1(r_{\\alpha,1},X,S)\\,|\\,X\\}\\big]$, so the value of a surrogate is precisely its residual predictiveness for the primary outcome given covariates, amplified by missingness. The coverage slack in Theorem 3 is a product of nuisance estimation errors, which the paper calls rate double robustness, so consistent nonparametric estimators of the propensity scores and conditional score distributions suffice for nominal coverage.","pith_inferences":["Editorial inference: the identification argument silently requires a cross-world condition—the surrogate measured under one treatment must carry no information about the outcome under the other treatment, given baseline covariates—because the counterfactual score is evaluated at the observed post-treatment surrogate; this condition is untested, and the paper's stated assumptions do not imply it.","Editorial inference: since no statistical-surrogacy assumption is imposed, arbitrary machine-learning predictions can serve as surrogates, making SCIENCE a conformal analogue of prediction-powered inference for individual effects, with the efficiency gain sized by the predictor's accuracy.","Testable extension: one could verify the cross-world condition empirically in placebo-controlled trials with a surrogate measured in both arms, by testing whether $S(1)$ predicts $Y(0)$ within strata of $X$; a nonzero association flags intervals whose counterfactual half may be miscalibrated."],"forward_implications":["SCIENCE's prediction intervals achieve the nominal $(1-\\alpha)$ coverage marginally and within prespecified groups, on the source data (one potential outcome observed) and on the target data (no primary outcomes).","Surrogates reduce interval width exactly when they are observed alongside missing primary outcomes: the Setting-2 efficiency bound beats Settings 1 and 3 by the surrogate-predictiveness variance $\\mathrm{var}\\{\\tilde m_a(r,X,S)\\,|\\,X\\}$ scaled by the missingness odds, and surrogates observed only in the source data give no gain.","The coverage slack is a product of nuisance estimation errors, so any consistent estimators of the propensity scores and conditional score distributions—including flexible nonparametric learners—suffice for nominal coverage asymptotically.","In the Moderna COVE trial, Day 29 neutralizing and binding antibody surrogates shorten the intervals for the Day 57 individual vaccine effect on antibody titer, with both markers together giving the shortest intervals.","Because only a marginal contrast (e.g., $u\\{Y(1)\\}-u\\{Y(0)\\}$) is required, the same machinery covers ITEs, log-risk-ratio contrasts, and utility-weighted ordinal contrasts."],"supporting_citations":[{"why":"Supplies the base framework for conformal inference of counterfactuals and ITEs, the nested conformal procedure for target data, and the weighted-CQR baseline SCIENCE must beat.","marker":"Lei and Candès (2021)"},{"why":"Provides the doubly-robust reformulation of covariate-shift prediction as a missing-data problem and the sample-splitting arguments behind Theorem 3.","marker":"Yang et al. (2024)"},{"why":"Contributes the surrogate-assisted semi-supervised data configuration and the conditional-independence argument behind Lemma 1.","marker":"Kallus and Mao (2020)"},{"why":"Supplies cross-fitting and the product-of-errors ('rate double robustness') structure used for the coverage slack in Theorem 3.","marker":"Chernozhukov et al. (2018)"},{"why":"Defines conformalized quantile residuals, the non-conformity score used to build all intervals.","marker":"Romano et al. (2019)"},{"why":"Supplies the source/target data-combination identification template that Lemma 1 shows SCIENCE's assumptions recover.","marker":"Athey et al. (2020)"}],"fun_headline_variants":["Surrogates tighten causal effect intervals","Conformal inference gets a surrogate boost","Efficient intervals for individual causal effects","Surrogate-assisted intervals stay honest","Narrower ITE bounds with surrogates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the surrogate observed under the actual treatment is non-informative about the counterfactual outcome being predicted, conditional on baseline covariates; the paper's stated assumptions do not identify the distribution of the counterfactual non-conformity score evaluated at that observed surrogate, so this unstated cross-world independence must hold for the intervals to be calibrated.","fun_headline_variants_meta":{"raw":{"variants":["Surrogates tighten causal effect intervals","Conformal inference gets a surrogate boost","Efficient intervals for individual causal effects","Surrogate-assisted intervals stay honest","Narrower ITE bounds with surrogates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1398,"prompt_tokens":1036,"completion_tokens":362,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":299}},"tokens_in":652,"tokens_out":362,"duration_ms":3674,"temperature":1.0,"reasoning_tokens":299,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:10:53.420976+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate full potential outcomes with the surrogate under one treatment driving the outcome under the other treatment (say $Y(0)=\\beta_0+\\beta_1 X+\\beta_2 S(1)+\\epsilon$) while keeping Assumptions 3–5 satisfied, run SCIENCE, and measure the source-data coverage of $u\\{Y(0)\\}$ among treated units; coverage falling below $1-\\alpha$ as $\\beta_2$ grows would falsify the identification claim.","supporting_citations":[],"review_version":1}