Pith. sign in

REVIEW 3 major objections 4 minor 3 references

Treatment effect extrapolation in the presence of unmeasured confounding

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The paper's central claim is that borrowing a second, larger, related trial's deconfounding function makes CATE extrapolation accurate outside the support of both trials when unmeasured confounding is nonlinear.

desk verdict A promising extension to multi-RCT deconfounding, but the simulation evidence is only under exact model specification and lacks all estimation details. read the letter →

arxiv 2509.06045 v1 pith:GBN5XAOJ submitted 2025-09-07 stat.ME

classification stat.ME MSC 62D20
keywords causalinferenceunmeasuredconfoundingconditionalaveragetreatmenteffectCATEextrapolationdeconfoundingfunctionrandomizedcontrolledtrialshierarchicalmodeltransportability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish a practical way to transport treatment-effect estimates from trials into a broader target population when observational data are confounded and no single trial covers the whole target range. It takes the existing idea that a trial can reveal a deconfounding function—the gap between a biased observational estimate and the true conditional average treatment effect—and adds a hierarchical structure that pools trials of different but related treatments. The shared part of unmeasured confounding is learned across trials, so the small trial's correction can be extrapolated through the larger trial's region. Simulation results show the two-trial strategy tracks the true CATE outside the support of both trials when the deconfounding function is quadratic, where the single-trial strategy fails. If this holds beyond the simulations, it offers a route to credible treatment-effect estimates in populations excluded from the trials.

What carries the argument

The deconfounding function eta_k(X), defined as the true CATE minus the biased observational estimate, together with the hierarchical decomposition eta_k(X) = eta(X) + epsilon_k(X) and the model E[eta_k(X)] = beta f(X) + gamma_k g(X). This allows the shared shape of unmeasured confounding, eta(X), to be estimated across trials and extrapolated beyond any single trial's covariate support. The biased observational estimate omega_hat_k(X) then anchors the debiased CATE at the target population, while the trial-based eta_hat_k(X) corrects it.

What would settle it

Simulate the same design but with deconfounding functions that are related but not additively shared, such as eta_1(X) = 2 - X and eta_2(X) = 1 - X^2, and compare one-RCT and two-RCT estimates outside both supports; the approach's advantage should disappear or reverse if the shared-structure assumption is doing the work.

Watch

Extended reading notes

Core claim

The paper claims that extrapolation of conditional average treatment effects for a treatment studied in a small RCT can be improved by pooling that RCT with a larger RCT of a different but related treatment, provided unmeasured confounders act through a shared function eta(X) plus treatment-specific deviations. It defines eta_k(X) = tau_k(X) - omega_k(X), where omega_k is the biased observational estimate, and models E[eta_k(X)] = beta f(X) + gamma_k g(X). The debiased CATE in the target population is tau_hat_k(X) = omega_hat_k(X) + eta_hat_k(X). Simulations vary the smaller RCT's sample size (100, 1000, 2000) and use linear or quadratic deconfounding functions. For quadratic deconfounding f

Load-bearing premise

The shared-structure assumption that all related treatments have deconfounding functions of the form eta_k(X) = eta(X) + epsilon_k(X) with the hierarchical model correctly specified; if true unmeasured confounding operates differently across treatments, borrowing from the second RCT can distort rather than improve the estimate.

Editorial extensions

If this is right

  • When deconfounding functions are linear, a single RCT already extrapolates accurately, so the multi-RCT borrowing is not needed.
  • With quadratic deconfounding functions, the two-RCT estimate stays accurate outside the union of both RCT supports, while the one-RCT estimate drifts; the gain is largest for the smallest smaller-RCT sample size.
  • The debiased CATE for the target population is obtained by adding the estimated deconfounding function to the observational biased estimate, so extrapolation only requires observational data on the target population.
  • Increasing the smaller RCT's sample size improves one-RCT extrapolation but does not eliminate the gap that borrowing closes.
  • The approach is specified for any number of related treatments sharing the same outcome, suggesting the two-RCT simulation is a special case of a general multi-RCT estimator.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the hierarchical random-effects assumption is misspecified, borrowing could shrink eta_1 toward the wrong shared shape; a diagnostic comparing the fitted eta_hat to a nonparametric single-RCT estimate in the overlapping region would reveal when borrowing is safe.
  • The same reasoning suggests borrowing helps most when the auxiliary trial is larger, has wider covariate support, and shares the outcome definition; the paper's simulations fix these conditions at favorable levels.
  • A natural extension is to more than two treatments, where the shared eta(X) could be estimated with more precision but also risks over-smoothing across heterogeneous confounding mechanisms; that trade-off is not explored in the paper.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript extends the RCT-debiasing approach of Kallus et al. (2018) to settings where multiple RCTs for related treatments are available. The key idea is to define a deconfounding function eta_k(X) = tau_k(X) - omega_k(X) for each treatment, assume an additive decomposition eta_k(X) = eta(X) + epsilon_k(X), and estimate eta_k via a hierarchical model E[eta_k(X)] = beta f(X) + gamma_k g(X). The debiased CATE in the target population is then tau_hat_k(X) = omega_hat_k(X) + eta_hat_k(X). A simulation study with two treatments compares estimating the CATE of treatment 1 using only its small RCT versus using both RCTs. Results are reported for linear and quadratic deconfounding functions and for three sample sizes of the smaller RCT; the authors conclude that borrowing from a second RCT improves extrapolation when the deconfounding function is nonlinear. The paper is a two-page symposium contribution.

Significance. If the proposed multi-RCT borrowing strategy works in realistic settings, it would be a useful extension of the deconfounding-function literature, enabling extrapolation of CATEs to target populations when a single RCT has narrow support. The paper addresses an important problem and the conceptual mechanism is clear: related treatments may share components of unmeasured confounding, and additional RCTs can help identify those shared components. A strength is that the method is not circular in the sense that eta_k is defined from RCT and observational quantities and the hierarchical model is estimated from RCT data. However, the current evidence is preliminary and in-sample: the simulation generates data under exactly the additive shared-structure assumption the method requires, and the manuscript omits details of the estimation procedure, simulation repetitions, and uncertainty quantification. Thus the significance depends on whether the approach is robust to misspecification of the shared structure and of f and g.

major comments (3)
  1. [Section 4.1] The quadratic simulation sets eta_1(X)=2-X-0.75X^2 and eta_2(X)=1-2X-0.75X^2, so the two deconfounding functions share exactly the same quadratic coefficient -0.75X^2. This is precisely the additive shared structure eta_k(X)=eta(X)+epsilon_k(X) assumed in Section 3, with epsilon_k linear. The simulation therefore demonstrates the method when its key assumption holds, but it provides no evidence about performance when that assumption is violated, e.g., when eta_2 has a different X^2 coefficient or when epsilon_k is not additive. Since the central claim is that borrowing from a second RCT improves out-of-support extrapolation, the absence of any misspecification scenario leaves open the possibility that borrowing hurts under realistic deviations. The paper needs a misspecification analysis before the claim can be supported.
  2. [Section 4.2 and Figure 1] The estimation procedure is underspecified. The text says 'we used two approaches to estimate the deconfounding function' but does not state how eta_k is computed from the RCT and observational estimates, what priors or optimizer are used for the hierarchical model, how many simulation repetitions were run, what performance metric is reported, or how uncertainty is quantified. Figure 1 shows point summaries without error bars. Without these details, the 'RCT1 only' baseline is difficult to interpret; in particular, with K=1 the random-effects variance in the hierarchical model is weakly identified and the baseline may depend heavily on unspecified priors. The absence of this information makes the simulation results non-reproducible and weakens the empirical claim.
  3. [Section 3] The core estimation step is not fully defined. The manuscript defines eta_k(X) = tau_k(X) - omega_k(X), where tau_k is estimated from an RCT and omega_k from observational data, but eta_k is not directly observed. It then says 'Using the data from the RCTs to learn eta_k(X), we propose a hierarchical model...' without specifying how the RCT-based estimate of tau_k and the observational estimate of omega_k are combined, whether the hierarchical model is fit in one stage or two, or how f(X) and g(X) are chosen. These choices are load-bearing for the extrapolation claim, especially because the model is used outside the support of the RCTs.
minor comments (4)
  1. [Section 2] The notation for RCT data is confusing: 'Y_k, T_k^k, X_k' appears to have a typo; the treatment indicator for RCT k should be T_k, not T_k^k.
  2. [Section 3] The assumption 'It seems reasonable to assume that there will be some shared unmeasured confounding, such as underlying frailty or disease progression' is informal. It would help to state explicitly that the assumption is about the deconfounding functions eta_k(X) sharing an additive component, and to discuss conditions under which this might hold or fail.
  3. [Figure 1] The figure caption says the pink shaded region is the support of RCT1 and the blue region is the support of RCT2. Because RCT1's support [1.5, 2] is nested inside RCT2's support [0, 2.5], the shading may be visually unclear; consider using distinct outlines or labels. Adding error bars or credible intervals would also improve interpretability.
  4. [References] The only methodological reference is Kallus et al. (2018). The literature on deconfounding and transportability has grown considerably; citing a few more recent works would place the contribution in context.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the deconfounding function is estimated from RCT data and combined with the observational CATE, with no fitted parameter relabeled as a prediction.

full rationale

The paper's derivation chain is: define the deconfounding function η_k(X) = τ_k(X) − ω_k(X), model it hierarchically as η_k(X) = η(X) + ε_k(X) with E[η_k(X)] = βf(X) + γ_k g(X), estimate β and γ_k from RCT data, and then reconstruct τhat_k(X) = ωhat_k(X) + ηhat_k(X). The identity τhat = ωhat + ηhat is definitional, but ηhat is not fitted from the observational ωhat; it is learned from the RCT outcomes, which are unconfounded and separate from the observational data. No fitted parameter is renamed as a prediction: extrapolation outside RCT support is genuine model-based extrapolation using f(X) and g(X), not a re-statement of the fit. The simulation generates data under the same additive shared-structure assumption the method uses, so the demonstration is in-sample; this is a robustness and external-validity concern, not circularity. The paper contains no self-citation load-bearing argument: reference [2] is external prior work, and the hierarchical model is proposed in this paper rather than imported from the authors' own earlier results. No uniqueness theorem or ansatz is smuggled in via citation. Thus no circular step can be exhibited, and the appropriate score is 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The deconfounding function is adopted from prior work (Kallus et al.), not invented here. The main modeling commitments are the additive shared structure and the hierarchical random-effects specification, both stated but not validated outside the simulation.

free parameters (2)
  • Functional forms f(X) and g(X)
    The shared and treatment-specific contributions to the deconfounding function are specified as f and g, but their forms (degree of polynomial, basis) are not stated; in the simulation they likely match the true generating functions, which is a modeling choice.
  • Random-effects distribution and variance
    The hierarchical model E[eta_k] = beta f(X) + gamma_k g(X) requires a distribution for gamma_k; not specified, but required to estimate and shrink the treatment-specific deviations.
assumptions (4)
  • domain assumption RCTs are unconfounded for their randomized treatments
    Section 2: 'We assume that the RCTs are unconfounded for their randomised treatments.' Standard in causal inference.
  • ad hoc to paper Deconfounding functions decompose additively across treatments: eta_k(X) = eta(X) + epsilon_k(X)
    Section 3 introduces this shared plus treatment-specific structure; it is the key assumption enabling borrowing between RCTs.
  • domain assumption Observational biased CATE omega_k(X) can be estimated consistently from confounded data
    The method relies on a 'biased estimate of the CATE ... learnt in the confounded observational data' (Section 3); estimation method and consistency conditions are not discussed.
  • ad hoc to paper Hierarchical model E[eta_k] = beta f(X) + gamma_k g(X) is correctly specified
    The borrowing benefit depends on the parametric forms f,g and random effects being adequate for the true deconfounding functions; this is not tested under misspecification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Treatment effect extrapolation in the presence of unmeasured confounding." pith.science (2026). https://pith.science/paper/GBN5XAOJ

@misc{pith2026250906045,
  author       = {Pith},
  title        = {Pith review of: Treatment effect extrapolation in the presence of unmeasured confounding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GBN5XAOJ}},
  note         = {Machine review of arXiv:2509.06045}
}
read the original abstract

While randomised controlled trials (RCTs) are the gold standard for estimating causal treatment effects, their limited sample sizes and restrictive criteria make it difficult to extrapolate to a broader population. Observational data, while larger, suffer from unmeasured confounding. Therefore, we can combine the strengths of both data sources for more accurate results. This work extends existing methods that use RCTs to debias conditional average treatment effects (CATEs) estimated in observational data by defining a deconfounding function. Our proposed approach borrows information from RCTs of multiple related treatments to improve the extrapolation of CATEs. Simulation results showed that, for non-linear deconfounding functions, using only one RCT poorly estimates the CATE outside of the support of that RCT. This is emphasised for smaller RCTs. Borrowing information from a second RCT provided more accurate estimates of the CATE outside of the support of both RCTs.

Figures

Figures reproduced from arXiv: 2509.06045 by the authors.

Figure 1
Figure 1. Summary of estimates of the conditional average treatment effect for treatment 1 (𝜏ˆ1 (𝑋 )). Panels are separated by the sample size of the smaller RCT (RCT1) and the true deconfounding function. The pink shaded region indicates the support of RCT1, and the blue shaded region the support of RCT2. also have data from a larger RCT exploring the effect of another treatment (𝑇2) on the same outcome 𝑌. We wish to utilise… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

3 extracted references · 1 canonical work pages

  1. [1]

    Irina Degtiar and Sherri Rose. 2023. A Review of Generalizability and Trans- portability.Annual Review of Statistics and Its Application(2023), 501–524. doi:10.1146/annurev-statistics-042522-103837

  2. [2]

    Nathan Kallus, Aahlad Manas Puli, and Uri Shalit. 2018. Removing Hidden Con- founding by Experimental Grounding. (2018). http://arxiv.org/abs/1810.11646 arXiv:1810.11646

  3. [3]

    To whom do the results of this trial apply?

    Peter M. Rothwell. 2005. External validity of randomised controlled trials: “To whom do the results of this trial apply?”.The Lancet365, 9453 (2005), 82–93. doi:10.1016/S0140-6736(04)17670-8

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.