{"id":"0c1fe47b-b582-4d01-997f-cd1e521e894b","arxiv_id":"2509.06045","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Pooling deconfounding functions across two related RCTs improves extrapolated CATE estimates in simulations with non-linear hidden confounding, compared with using only the smaller RCT.","lead":"A short paper proposes borrowing information from a second randomized trial of a related treatment to correct for hidden bias when extending a small trial's effect to a broader population. Simulations suggest this improves accuracy outside the trials' support when the bias correction is non-linear.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Multi-RCT borrowing is only shown to improve under exact shared-structure specification; the paper reports no misspecification analysis, so it is unknown whether the second RCT helps, hurts, or is neutral when the additive decomposition is wrong.","rationale":"The reader's weakest assumption is exactly the additive shared-structure assumption. The stress-test confirms this is the single most load-bearing concern: the simulation is constructed so the shared component is identical across treatments, and the claimed improvement is driven by that shared component. The paper provides no sensitivity analysis, no model checking, and no guidance on choosing f and g. The absence of estimation details further weakens the evidence, but the core logical gap is the untested assumption that the additive decomposition holds. A misspecification simulation that varies the shared quadratic coefficient would directly test whether the central claim generalizes. If the test shows that borrowing can hurt, the conditional verdict is appropriate and perhaps should lean toward rejecting the general claim; if it shows robustness, the conditional verdict is confirmed. Since the current evidence is insufficient but not contradictory, the reader's CONDITIONAL verdict remains appropriate: the idea is plausible but requires additional validation. No change to the verdict is needed; the stress-test reinforces the reader's assessment.","tokens_in":3605,"tokens_out":9475,"duration_ms":98359,"concrete_test":"Run a misspecification simulation where η2(X)=1−2X−cX² with c ∈ {−1.0, −0.5, 0} while η1(X)=2−X−0.75X², using the same sample sizes (n0=50,000, n2=5,000, n1∈{100,1,000,2,000}) and fitting the same hierarchical model with f(X)=X² and g(X)=(1,X). Compare the MSE of the multi-RCT estimate of η1 outside [0,2.5] against the single-RCT estimate. If the multi-RCT estimate has larger MSE for some c, the claimed improvement is not robust; if it remains better for all c, the claim gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that borrowing from a second RCT improves out-of-support CATE extrapolation. The mechanism is the hierarchical decomposition η_k(X)=η(X)+ε_k(X), estimated via E[η_k(X)]=βf(X)+γ_k g(X) (Section 3). In the only simulation (Section 4), η1(X)=2−X−0.75X² and η2(X)=1−2X−0.75X² share exactly the same quadratic coefficient −0.75X², so the model is correctly specified. RCT2 helps precisely because it provides additional data on that shared quadratic term, which is the dominant source of extrapolation error. No misspecification scenario is reported: if the treatments' deconfounding functions do not actually share the same shape, or if f/g are misspecified (e.g., true η2 has a different X² coefficient), the random-effects shrinkage will pull β toward an average, biasing η1 outside support. The paper also omits all estimation details (priors, MCMC/optimizer, number of repetitions, error bars), so the 'RCT1 only' baseline may itself be an artifact of an unidentifiable K=1 hierarchical fit. Thus the evidence only supports the claim under exact model specification; it is unknown whether borrowing helps, hurts, or is neutral under realistic misspecification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript extends the RCT-debiasing approach of Kallus et al. (2018) to settings where multiple RCTs for related treatments are available. The key idea is to define a deconfounding function eta_k(X) = tau_k(X) - omega_k(X) for each treatment, assume an additive decomposition eta_k(X) = eta(X) + epsilon_k(X), and estimate eta_k via a hierarchical model E[eta_k(X)] = beta f(X) + gamma_k g(X). The debiased CATE in the target population is then tau_hat_k(X) = omega_hat_k(X) + eta_hat_k(X). A simulation study with two treatments compares estimating the CATE of treatment 1 using only its small RCT versus using both RCTs. Results are reported for linear and quadratic deconfounding functions and for three sample sizes of the smaller RCT; the authors conclude that borrowing from a second RCT improves extrapolation when the deconfounding function is nonlinear. The paper is a two-page symposium contribution.","tokens_in":3958,"tokens_out":6649,"duration_ms":73215,"significance":"If the proposed multi-RCT borrowing strategy works in realistic settings, it would be a useful extension of the deconfounding-function literature, enabling extrapolation of CATEs to target populations when a single RCT has narrow support. The paper addresses an important problem and the conceptual mechanism is clear: related treatments may share components of unmeasured confounding, and additional RCTs can help identify those shared components. A strength is that the method is not circular in the sense that eta_k is defined from RCT and observational quantities and the hierarchical model is estimated from RCT data. However, the current evidence is preliminary and in-sample: the simulation generates data under exactly the additive shared-structure assumption the method requires, and the manuscript omits details of the estimation procedure, simulation repetitions, and uncertainty quantification. Thus the significance depends on whether the approach is robust to misspecification of the shared structure and of f and g.","major_comments":[{"comment":"The quadratic simulation sets eta_1(X)=2-X-0.75X^2 and eta_2(X)=1-2X-0.75X^2, so the two deconfounding functions share exactly the same quadratic coefficient -0.75X^2. This is precisely the additive shared structure eta_k(X)=eta(X)+epsilon_k(X) assumed in Section 3, with epsilon_k linear. The simulation therefore demonstrates the method when its key assumption holds, but it provides no evidence about performance when that assumption is violated, e.g., when eta_2 has a different X^2 coefficient or when epsilon_k is not additive. Since the central claim is that borrowing from a second RCT improves out-of-support extrapolation, the absence of any misspecification scenario leaves open the possibility that borrowing hurts under realistic deviations. The paper needs a misspecification analysis before the claim can be supported.","section":"Section 4.1"},{"comment":"The estimation procedure is underspecified. The text says 'we used two approaches to estimate the deconfounding function' but does not state how eta_k is computed from the RCT and observational estimates, what priors or optimizer are used for the hierarchical model, how many simulation repetitions were run, what performance metric is reported, or how uncertainty is quantified. Figure 1 shows point summaries without error bars. Without these details, the 'RCT1 only' baseline is difficult to interpret; in particular, with K=1 the random-effects variance in the hierarchical model is weakly identified and the baseline may depend heavily on unspecified priors. The absence of this information makes the simulation results non-reproducible and weakens the empirical claim.","section":"Section 4.2 and Figure 1"},{"comment":"The core estimation step is not fully defined. The manuscript defines eta_k(X) = tau_k(X) - omega_k(X), where tau_k is estimated from an RCT and omega_k from observational data, but eta_k is not directly observed. It then says 'Using the data from the RCTs to learn eta_k(X), we propose a hierarchical model...' without specifying how the RCT-based estimate of tau_k and the observational estimate of omega_k are combined, whether the hierarchical model is fit in one stage or two, or how f(X) and g(X) are chosen. These choices are load-bearing for the extrapolation claim, especially because the model is used outside the support of the RCTs.","section":"Section 3"}],"minor_comments":[{"comment":"The notation for RCT data is confusing: 'Y_k, T_k^k, X_k' appears to have a typo; the treatment indicator for RCT k should be T_k, not T_k^k.","section":"Section 2"},{"comment":"The assumption 'It seems reasonable to assume that there will be some shared unmeasured confounding, such as underlying frailty or disease progression' is informal. It would help to state explicitly that the assumption is about the deconfounding functions eta_k(X) sharing an additive component, and to discuss conditions under which this might hold or fail.","section":"Section 3"},{"comment":"The figure caption says the pink shaded region is the support of RCT1 and the blue region is the support of RCT2. Because RCT1's support [1.5, 2] is nested inside RCT2's support [0, 2.5], the shading may be visually unclear; consider using distinct outlines or labels. Adding error bars or credible intervals would also improve interpretability.","section":"Figure 1"},{"comment":"The only methodological reference is Kallus et al. (2018). The literature on deconfounding and transportability has grown considerably; citing a few more recent works would place the contribution in context.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a two-page symposium paper. The central idea is promising, but the current version lacks the details and robustness analysis needed for a journal publication. The stress-test concern about in-sample demonstration lands: the simulation is exactly aligned with the model assumptions, so the favorable result is not surprising. I recommend major revision requiring a fully specified estimation algorithm, a reproducibility-oriented simulation description, and a misspecification analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe one thing to know: this is a two-page workshop paper with a nice idea — borrowing information from RCTs of different but related treatments to improve CATE extrapolation in a target population — but the evidence is a single simulation generated under exactly the model's assumptions. The result is plausible but not demonstrated with the rigor it would need for a full paper.\n\nWhat's new: the hierarchical extension of Kallus et al.'s deconfounding function, writing η_k(X)=η(X)+ε_k(X), with E[η_k]=βf(X)+γ_k g(X). That allows information to be shared across RCTs, which is genuinely absent from the cited literature. The simulation design is sensible: small RCT for treatment 1, larger RCT for related treatment 2, and extrapolation outside the support of both. The quadratic case shows the single-RCT estimate is poor, and borrowing helps. The paper is clearly written and doesn't overclaim beyond its evidence.\n\nSoft spots: the paper omits nearly all estimation details. How are β and γ_k estimated? What priors or optimizer? How many simulation repetitions? There are no error bars or uncertainty measures in Figure 1, and no code. So the \"RCT1 only\" baseline is also underspecified; if it's a K=1 hierarchical fit, the decomposition is unidentifiable, and if it's something else, we need to know. The bigger issue is that the simulation tests only the correctly specified case: the shared quadratic coefficient is exactly the same in both treatments, so borrowing helps because the second RCT provides data on that shared term. No misspecification is reported — different shapes, wrong f or g, or misspecified random-effects distribution could all make the second RCT actively harmful. The authors don't state these limitations or give any analysis of robustness.\n\nWho this is for: anyone working on data fusion or transportability with unmeasured confounding will be interested in the idea. As it stands, it's a promising abstract, not a complete paper. I'd send it to peer review for a workshop, but for a journal it would need substantial additional work: code, uncertainty quantification, and tests under misspecification.\n\nRecommendation: if it comes across your desk, don't desk-reject; the idea is worth checking. But ask for the missing details and a misspecification analysis before publishing.","headline":"A promising extension to multi-RCT deconfounding, but the simulation evidence is only under exact model specification and lacks all estimation details.","tokens_in":4372,"tokens_out":3001,"would_cite":false,"duration_ms":33270,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's central claim is that borrowing a second, larger, related trial's deconfounding function makes CATE extrapolation accurate outside the support of both trials when unmeasured confounding is nonlinear.","keywords":["causal inference","unmeasured confounding","conditional average treatment effect","CATE extrapolation","deconfounding function","randomized controlled trials","hierarchical model","transportability"],"falsifier":"Simulate the same design but with deconfounding functions that are related but not additively shared, such as eta_1(X) = 2 - X and eta_2(X) = 1 - X^2, and compare one-RCT and two-RCT estimates outside both supports; the approach's advantage should disappear or reverse if the shared-structure assumption is doing the work.","tokens_in":3508,"feed_emoji":"📊","tokens_out":5915,"duration_ms":62018,"temperature":0.7,"pith_summary":"This paper tries to establish a practical way to transport treatment-effect estimates from trials into a broader target population when observational data are confounded and no single trial covers the whole target range. It takes the existing idea that a trial can reveal a deconfounding function—the gap between a biased observational estimate and the true conditional average treatment effect—and adds a hierarchical structure that pools trials of different but related treatments. The shared part of unmeasured confounding is learned across trials, so the small trial's correction can be extrapolated through the larger trial's region. Simulation results show the two-trial strategy tracks the true CATE outside the support of both trials when the deconfounding function is quadratic, where the single-trial strategy fails. If this holds beyond the simulations, it offers a route to credible treatment-effect estimates in populations excluded from the trials.","feed_headline":"A second clinical trial fixes effect estimates beyond both trials","feed_subtitle":"Sharing a common deconfounding shape lets a larger related trial rescue extrapolation a small trial misses.","key_machinery":"The deconfounding function eta_k(X), defined as the true CATE minus the biased observational estimate, together with the hierarchical decomposition eta_k(X) = eta(X) + epsilon_k(X) and the model E[eta_k(X)] = beta f(X) + gamma_k g(X). This allows the shared shape of unmeasured confounding, eta(X), to be estimated across trials and extrapolated beyond any single trial's covariate support. The biased observational estimate omega_hat_k(X) then anchors the debiased CATE at the target population, while the trial-based eta_hat_k(X) corrects it.","core_discovery":"The paper claims that extrapolation of conditional average treatment effects for a treatment studied in a small RCT can be improved by pooling that RCT with a larger RCT of a different but related treatment, provided unmeasured confounders act through a shared function eta(X) plus treatment-specific deviations. It defines eta_k(X) = tau_k(X) - omega_k(X), where omega_k is the biased observational estimate, and models E[eta_k(X)] = beta f(X) + gamma_k g(X). The debiased CATE in the target population is tau_hat_k(X) = omega_hat_k(X) + eta_hat_k(X). Simulations vary the smaller RCT's sample size (100, 1000, 2000) and use linear or quadratic deconfounding functions. For quadratic deconfounding f","pith_inferences":["If the hierarchical random-effects assumption is misspecified, borrowing could shrink eta_1 toward the wrong shared shape; a diagnostic comparing the fitted eta_hat to a nonparametric single-RCT estimate in the overlapping region would reveal when borrowing is safe.","The same reasoning suggests borrowing helps most when the auxiliary trial is larger, has wider covariate support, and shares the outcome definition; the paper's simulations fix these conditions at favorable levels.","A natural extension is to more than two treatments, where the shared eta(X) could be estimated with more precision but also risks over-smoothing across heterogeneous confounding mechanisms; that trade-off is not explored in the paper."],"forward_implications":["When deconfounding functions are linear, a single RCT already extrapolates accurately, so the multi-RCT borrowing is not needed.","With quadratic deconfounding functions, the two-RCT estimate stays accurate outside the union of both RCT supports, while the one-RCT estimate drifts; the gain is largest for the smallest smaller-RCT sample size.","The debiased CATE for the target population is obtained by adding the estimated deconfounding function to the observational biased estimate, so extrapolation only requires observational data on the target population.","Increasing the smaller RCT's sample size improves one-RCT extrapolation but does not eliminate the gap that borrowing closes.","The approach is specified for any number of related treatments sharing the same outcome, suggesting the two-RCT simulation is a special case of a general multi-RCT estimator."],"supporting_citations":[{"why":"Supplies the generalizability and transportability framing that motivates combining RCT and observational data.","marker":"[1]"},{"why":"Defines the deconfounding-function debiasing method that this paper extends to multiple RCTs.","marker":"[2]"},{"why":"Documents the external-validity limits of RCTs that motivate the extrapolation problem.","marker":"[3]"}],"fun_headline_variants":["Two RCTs beat one for extrapolating treatment effects","Borrowing from a second trial sharpens effect extrapolation","Related trial data rescues extrapolation beyond trial support","Extra RCT fixes effect estimates outside trial range","Shared confounding shape lets second RCT improve estimates"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The shared-structure assumption that all related treatments have deconfounding functions of the form eta_k(X) = eta(X) + epsilon_k(X) with the hierarchical model correctly specified; if true unmeasured confounding operates differently across treatments, borrowing from the second RCT can distort rather than improve the estimate.","fun_headline_variants_meta":{"raw":{"variants":["Two RCTs beat one for extrapolating treatment effects","Borrowing from a second trial sharpens effect extrapolation","Related trial data rescues extrapolation beyond trial support","Extra RCT fixes effect estimates outside trial range","Shared confounding shape lets second RCT improve estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000143,"raw_usage":{"total_tokens":988,"prompt_tokens":705,"completion_tokens":283,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":208}},"tokens_in":449,"tokens_out":283,"duration_ms":3461,"temperature":1.0,"reasoning_tokens":208,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T04:32:39.202821+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the same design but with deconfounding functions that are related but not additively shared, such as eta_1(X) = 2 - X and eta_2(X) = 1 - X^2, and compare one-RCT and two-RCT estimates outside both supports; the approach's advantage should disappear or reverse if the shared-structure assumption is doing the work.","supporting_citations":[{"cited_title":"Removing Hidden Confounding by Experimental Grounding","cited_arxiv_id":"1810.11646","evidence_quote":"Defines the deconfounding-function debiasing method that this paper extends to multiple RCTs."}],"review_version":1}