REVIEW 3 major objections 4 minor 3 references
Treatment effect extrapolation in the presence of unmeasured confounding
T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper's central claim is that borrowing a second, larger, related trial's deconfounding function makes CATE extrapolation accurate outside the support of both trials when unmeasured confounding is nonlinear.
desk verdict A promising extension to multi-RCT deconfounding, but the simulation evidence is only under exact model specification and lacks all estimation details. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The deconfounding function eta_k(X), defined as the true CATE minus the biased observational estimate, together with the hierarchical decomposition eta_k(X) = eta(X) + epsilon_k(X) and the model E[eta_k(X)] = beta f(X) + gamma_k g(X). This allows the shared shape of unmeasured confounding, eta(X), to be estimated across trials and extrapolated beyond any single trial's covariate support. The biased observational estimate omega_hat_k(X) then anchors the debiased CATE at the target population, while the trial-based eta_hat_k(X) corrects it.
What would settle it
Simulate the same design but with deconfounding functions that are related but not additively shared, such as eta_1(X) = 2 - X and eta_2(X) = 1 - X^2, and compare one-RCT and two-RCT estimates outside both supports; the approach's advantage should disappear or reverse if the shared-structure assumption is doing the work.
Extended reading notes
Core claim
The paper claims that extrapolation of conditional average treatment effects for a treatment studied in a small RCT can be improved by pooling that RCT with a larger RCT of a different but related treatment, provided unmeasured confounders act through a shared function eta(X) plus treatment-specific deviations. It defines eta_k(X) = tau_k(X) - omega_k(X), where omega_k is the biased observational estimate, and models E[eta_k(X)] = beta f(X) + gamma_k g(X). The debiased CATE in the target population is tau_hat_k(X) = omega_hat_k(X) + eta_hat_k(X). Simulations vary the smaller RCT's sample size (100, 1000, 2000) and use linear or quadratic deconfounding functions. For quadratic deconfounding f
Load-bearing premise
The shared-structure assumption that all related treatments have deconfounding functions of the form eta_k(X) = eta(X) + epsilon_k(X) with the hierarchical model correctly specified; if true unmeasured confounding operates differently across treatments, borrowing from the second RCT can distort rather than improve the estimate.
Editorial extensions
If this is right
- When deconfounding functions are linear, a single RCT already extrapolates accurately, so the multi-RCT borrowing is not needed.
- With quadratic deconfounding functions, the two-RCT estimate stays accurate outside the union of both RCT supports, while the one-RCT estimate drifts; the gain is largest for the smallest smaller-RCT sample size.
- The debiased CATE for the target population is obtained by adding the estimated deconfounding function to the observational biased estimate, so extrapolation only requires observational data on the target population.
- Increasing the smaller RCT's sample size improves one-RCT extrapolation but does not eliminate the gap that borrowing closes.
- The approach is specified for any number of related treatments sharing the same outcome, suggesting the two-RCT simulation is a special case of a general multi-RCT estimator.
Reading between the lines
- If the hierarchical random-effects assumption is misspecified, borrowing could shrink eta_1 toward the wrong shared shape; a diagnostic comparing the fitted eta_hat to a nonparametric single-RCT estimate in the overlapping region would reveal when borrowing is safe.
- The same reasoning suggests borrowing helps most when the auxiliary trial is larger, has wider covariate support, and shares the outcome definition; the paper's simulations fix these conditions at favorable levels.
- A natural extension is to more than two treatments, where the shared eta(X) could be estimated with more precision but also risks over-smoothing across heterogeneous confounding mechanisms; that trade-off is not explored in the paper.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript extends the RCT-debiasing approach of Kallus et al. (2018) to settings where multiple RCTs for related treatments are available. The key idea is to define a deconfounding function eta_k(X) = tau_k(X) - omega_k(X) for each treatment, assume an additive decomposition eta_k(X) = eta(X) + epsilon_k(X), and estimate eta_k via a hierarchical model E[eta_k(X)] = beta f(X) + gamma_k g(X). The debiased CATE in the target population is then tau_hat_k(X) = omega_hat_k(X) + eta_hat_k(X). A simulation study with two treatments compares estimating the CATE of treatment 1 using only its small RCT versus using both RCTs. Results are reported for linear and quadratic deconfounding functions and for three sample sizes of the smaller RCT; the authors conclude that borrowing from a second RCT improves extrapolation when the deconfounding function is nonlinear. The paper is a two-page symposium contribution.
Significance. If the proposed multi-RCT borrowing strategy works in realistic settings, it would be a useful extension of the deconfounding-function literature, enabling extrapolation of CATEs to target populations when a single RCT has narrow support. The paper addresses an important problem and the conceptual mechanism is clear: related treatments may share components of unmeasured confounding, and additional RCTs can help identify those shared components. A strength is that the method is not circular in the sense that eta_k is defined from RCT and observational quantities and the hierarchical model is estimated from RCT data. However, the current evidence is preliminary and in-sample: the simulation generates data under exactly the additive shared-structure assumption the method requires, and the manuscript omits details of the estimation procedure, simulation repetitions, and uncertainty quantification. Thus the significance depends on whether the approach is robust to misspecification of the shared structure and of f and g.
major comments (3)
- [Section 4.1] The quadratic simulation sets eta_1(X)=2-X-0.75X^2 and eta_2(X)=1-2X-0.75X^2, so the two deconfounding functions share exactly the same quadratic coefficient -0.75X^2. This is precisely the additive shared structure eta_k(X)=eta(X)+epsilon_k(X) assumed in Section 3, with epsilon_k linear. The simulation therefore demonstrates the method when its key assumption holds, but it provides no evidence about performance when that assumption is violated, e.g., when eta_2 has a different X^2 coefficient or when epsilon_k is not additive. Since the central claim is that borrowing from a second RCT improves out-of-support extrapolation, the absence of any misspecification scenario leaves open the possibility that borrowing hurts under realistic deviations. The paper needs a misspecification analysis before the claim can be supported.
- [Section 4.2 and Figure 1] The estimation procedure is underspecified. The text says 'we used two approaches to estimate the deconfounding function' but does not state how eta_k is computed from the RCT and observational estimates, what priors or optimizer are used for the hierarchical model, how many simulation repetitions were run, what performance metric is reported, or how uncertainty is quantified. Figure 1 shows point summaries without error bars. Without these details, the 'RCT1 only' baseline is difficult to interpret; in particular, with K=1 the random-effects variance in the hierarchical model is weakly identified and the baseline may depend heavily on unspecified priors. The absence of this information makes the simulation results non-reproducible and weakens the empirical claim.
- [Section 3] The core estimation step is not fully defined. The manuscript defines eta_k(X) = tau_k(X) - omega_k(X), where tau_k is estimated from an RCT and omega_k from observational data, but eta_k is not directly observed. It then says 'Using the data from the RCTs to learn eta_k(X), we propose a hierarchical model...' without specifying how the RCT-based estimate of tau_k and the observational estimate of omega_k are combined, whether the hierarchical model is fit in one stage or two, or how f(X) and g(X) are chosen. These choices are load-bearing for the extrapolation claim, especially because the model is used outside the support of the RCTs.
minor comments (4)
- [Section 2] The notation for RCT data is confusing: 'Y_k, T_k^k, X_k' appears to have a typo; the treatment indicator for RCT k should be T_k, not T_k^k.
- [Section 3] The assumption 'It seems reasonable to assume that there will be some shared unmeasured confounding, such as underlying frailty or disease progression' is informal. It would help to state explicitly that the assumption is about the deconfounding functions eta_k(X) sharing an additive component, and to discuss conditions under which this might hold or fail.
- [Figure 1] The figure caption says the pink shaded region is the support of RCT1 and the blue region is the support of RCT2. Because RCT1's support [1.5, 2] is nested inside RCT2's support [0, 2.5], the shading may be visually unclear; consider using distinct outlines or labels. Adding error bars or credible intervals would also improve interpretability.
- [References] The only methodological reference is Kallus et al. (2018). The literature on deconfounding and transportability has grown considerably; citing a few more recent works would place the contribution in context.
Circularity Check
No significant circularity: the deconfounding function is estimated from RCT data and combined with the observational CATE, with no fitted parameter relabeled as a prediction.
full rationale
The paper's derivation chain is: define the deconfounding function η_k(X) = τ_k(X) − ω_k(X), model it hierarchically as η_k(X) = η(X) + ε_k(X) with E[η_k(X)] = βf(X) + γ_k g(X), estimate β and γ_k from RCT data, and then reconstruct τhat_k(X) = ωhat_k(X) + ηhat_k(X). The identity τhat = ωhat + ηhat is definitional, but ηhat is not fitted from the observational ωhat; it is learned from the RCT outcomes, which are unconfounded and separate from the observational data. No fitted parameter is renamed as a prediction: extrapolation outside RCT support is genuine model-based extrapolation using f(X) and g(X), not a re-statement of the fit. The simulation generates data under the same additive shared-structure assumption the method uses, so the demonstration is in-sample; this is a robustness and external-validity concern, not circularity. The paper contains no self-citation load-bearing argument: reference [2] is external prior work, and the hierarchical model is proposed in this paper rather than imported from the authors' own earlier results. No uniqueness theorem or ansatz is smuggled in via citation. Thus no circular step can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
free parameters (2)
- Functional forms f(X) and g(X)
- Random-effects distribution and variance
assumptions (4)
- domain assumption RCTs are unconfounded for their randomized treatments
- ad hoc to paper Deconfounding functions decompose additively across treatments: eta_k(X) = eta(X) + epsilon_k(X)
- domain assumption Observational biased CATE omega_k(X) can be estimated consistently from confounded data
- ad hoc to paper Hierarchical model E[eta_k] = beta f(X) + gamma_k g(X) is correctly specified
Cite this review
Pith. "Pith review of Treatment effect extrapolation in the presence of unmeasured confounding." pith.science (2026). https://pith.science/paper/GBN5XAOJ
@misc{pith2026250906045,
author = {Pith},
title = {Pith review of: Treatment effect extrapolation in the presence of unmeasured confounding},
year = {2026},
howpublished = {\url{https://pith.science/paper/GBN5XAOJ}},
note = {Machine review of arXiv:2509.06045}
}
read the original abstract
While randomised controlled trials (RCTs) are the gold standard for estimating causal treatment effects, their limited sample sizes and restrictive criteria make it difficult to extrapolate to a broader population. Observational data, while larger, suffer from unmeasured confounding. Therefore, we can combine the strengths of both data sources for more accurate results. This work extends existing methods that use RCTs to debias conditional average treatment effects (CATEs) estimated in observational data by defining a deconfounding function. Our proposed approach borrows information from RCTs of multiple related treatments to improve the extrapolation of CATEs. Simulation results showed that, for non-linear deconfounding functions, using only one RCT poorly estimates the CATE outside of the support of that RCT. This is emphasised for smaller RCTs. Borrowing information from a second RCT provided more accurate estimates of the CATE outside of the support of both RCTs.
Figures
Reference graph
Works this paper leans on
-
[1]
Irina Degtiar and Sherri Rose. 2023. A Review of Generalizability and Trans- portability.Annual Review of Statistics and Its Application(2023), 501–524. doi:10.1146/annurev-statistics-042522-103837
-
[2]
Nathan Kallus, Aahlad Manas Puli, and Uri Shalit. 2018. Removing Hidden Con- founding by Experimental Grounding. (2018). http://arxiv.org/abs/1810.11646 arXiv:1810.11646
work page Pith review arXiv 2018
-
[3]
To whom do the results of this trial apply?
Peter M. Rothwell. 2005. External validity of randomised controlled trials: “To whom do the results of this trial apply?”.The Lancet365, 9453 (2005), 82–93. doi:10.1016/S0140-6736(04)17670-8
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.