{"id":"36acdeb6-75ea-4bc2-9a42-b034c84f01cb","arxiv_id":"2607.29295","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fusing a randomized trial with real-world data under an assumed transportable treatment effect, Bayesian fusion forests estimate heterogeneous survival treatment effects in right- and interval-censored data using BART and a confounding function.","lead":"Bayesian fusion forests combine a randomized trial with real-world data to estimate how survival treatment effects vary across patients, absorbing confounding in the observational source with an explicit confounding function. The method handles right- and interval-censored outcomes and, on HIV data, reports subject-level certainty of benefit that a trial-only analysis lacks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported RWD precision gain may be an artifact of unmatched prior shrinkage, not borrowed information","rationale":"The reader's Assumption 4 concern is real and untestable, but a stress-test should also look for places where the central claim could fail even under the paper's own assumptions. The unmatched prior is such a place. The fusion model's τ prior is explicitly stronger than the BART default, while the trial-only comparator is not documented. This makes the headline efficiency-gain comparison potentially apples-to-oranges. It is concrete, testable, and would settle whether the RWD actually contributes. Until this is addressed, the conditional verdict stands: the method is coherent and the simulation is extensive, but the application's precision and 'benefit for nearly everyone' conclusion are not yet supported as presented.","tokens_in":32487,"tokens_out":6864,"duration_ms":82700,"concrete_test":"Re-run the trial-only AFT Bayesian causal forest on the ACTG175/MACS aligned cohort using the fusion's exact τ prior (BART(100, 1/2, 0.95, 3)), the same HDPM error prior, the same outcome standardization, and the same number of posterior samples. Also re-run the simulation at λd=λu=1 with the same matched settings. Compute the average 95% interval width and posterior variance of the subject-level acceleration factor. If the 62%/88% reductions fall to near zero, the claimed RWD precision gain is prior regularization, not data fusion. Report the original trial-only forest's hyperparameters to confirm the mismatch.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim is that fusing the RWD cuts subject-level posterior variance by about 88% and interval width by 62% relative to a trial-only analysis (Supplementary S.5.1, Table S4). The comparator is described only as the 'single-source counterpart' AFT Bayesian causal forest (Jacobs, 2026), and its hyperparameters are never specified. The fusion model deliberately shrinks the treatment-effect forest beyond the BART default: Jτ=100, kτ=1/2, ατ=0.95, βτ=3 (§2.5). If the trial-only forest uses the default BART settings (J=200, k=1, α=0.95, β=2), part or all of the reported precision reduction could be a prior-shrinkage effect rather than information borrowed from the RWD. The same unmatched comparison underlies the simulation's variance-ratio and RMSE findings (Figures 1–2, S1–S2). This threatens the central claim that the fusion estimates τ with lower variance than a trial-only analysis, even under Assumptions 1–4. Unlike transportability, which the paper explicitly acknowledges as fragile, this is a concrete, checkable internal-validity issue. The application's headline conclusion—'benefit for nearly every patient'—relies on the precision gain, so this concern is load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes the Bayesian fusion forest, a BART-based AFT model for heterogeneous treatment effects on survival that combines an RCT with real-world data. The key identifying device is a confounding function c(x) that absorbs RWD bias, while τ(x) is identified from the RCT under cross-source transportability (Assumptions 1–4). Each component of the decomposition Eq. (3) receives its own regularised tree ensemble, with a hierarchical Dirichlet process mixture for the error distribution. Simulation studies compare the method with a trial-only Bayesian causal forest and several machine-learning baselines, and claim large precision gains and robustness to confounding and covariate dimensionality. The method is applied to ACTG 175 plus MACS, giving an average acceleration factor of 1.65 [1.43;1.90] and a posterior probability of benefit above 0.95 for 96.4% of patients, versus 38% for the trial-only analysis.","tokens_in":32731,"tokens_out":5517,"duration_ms":67466,"significance":"The paper addresses an important problem: combining randomised and real-world evidence for heterogeneous treatment effects on censored survival outcomes without assuming the RWD is unconfounded. The causal decomposition is clean, Propositions 1 and 2 are proved transparently in the supplement, and the implementation is reproducible with public code and 1000-replication simulations. If the efficiency claim is established, this would be a useful contribution. However, the central quantitative claim—that fusion reduces posterior variance by about 88% and interval width by 62% relative to a trial-only analysis—rests on a comparator whose hyperparameters are not specified, so part or all of the apparent gain could be a prior-shrinkage artifact. The application's headline conclusion also depends on the untestable Assumption 4, and the reported certainty should be conditioned on that assumption or accompanied by a sensitivity analysis.","major_comments":[{"comment":"The comparator for the main precision claim is the 'single-source counterpart' AFT Bayesian causal forest (Jacobs, 2026), but its hyperparameters are never given. The fusion model deliberately uses a strongly shrunk treatment-effect forest: τ∼BART(100, 1/2, 0.95, 3) with kτ=1/2 (§2.5). If the trial-only causal forest uses standard BART defaults (e.g., J=200, k=1, α=0.95, β=2), the reported reductions in posterior variance (88%) and credible-interval width (62%) in Table S4, and the RMSE/variance ratios in Figures 1–2 and S2, may reflect prior shrinkage rather than information borrowed from the RWD. This is load-bearing for the paper's central claim that fusion estimates τ with lower variance than a trial-only analysis. Please re-run the trial-only causal forest with exactly the same τ-forest prior (J=100, k=1/2, α=0.95, β=3), and also, as a robustness check, the fusion with the trial-onl","section":"§2.5 and §3; Table S4 in Supplement S.5.1"},{"comment":"The application's conclusion—'benefit for nearly every patient' and the precision gain at the individual level—is conditional on cross-source transportability, Assumption 4: E[logT(1)−logT(0)|X,S=1] = E[logT(1)−logT(0)|X,S=0]. Proposition 2(iii) shows the RWD identifies only the composite τ+c, so if Assumption 4 fails, the fusion's conditional benefit statements inherit the RWD bias. Restricting both sources to men with CD4 200–500 is a reasonable design choice but does not verify the assumption. The manuscript should either include a sensitivity analysis that perturbs the transportability assumption (e.g., by allowing a shift δ(x) in the CATE transport equation and re-examining the posterior probability of benefit) or temper the conclusion to explicitly state that the certainty holds only under Assumption 4. This is not a circularity concern—Proposition 2(i) identifies τ from the RCT in","section":"§2.1 (Assumption 4) and §4 (Table 2, Figure 4)"}],"minor_comments":[{"comment":"The statement that the hyperparameters are 'tuned further by cross-validation' is never operationalised. State what is tuned (e.g., kτ, kc, or tree-structure parameters), on which data, and with what criterion. The simulation and application both use 'default parameters', so the role of cross-validation is unclear.","section":"§2.5"},{"comment":"The trial-only comparator is described only via a software-package reference. Provide the exact model specification, including the BART prior hyperparameters and any treatment-effect shrinkage, so the comparison is reproducible.","section":"§3; reference Jacobs (2026)"},{"comment":"The comparison with the four machine-learning methods is reasonable but the paper should be careful when saying it 'improves on all' of them: those methods are not designed for data fusion and are adapted in ways that may be suboptimal. The supplementary discussion appropriately notes the lack of ground truth in the application, but the main text should carry that caveat more explicitly.","section":"Table 1 and Supplement S.5.2"},{"comment":"The limitation section correctly notes that the acceleration factor is assumed time-invariant. It would be helpful to also mention that the full model has no posterior-concentration or consistency theorem for the four-forest HDPM specification; the empirical calibration in the simulation is the current support. If the authors intend a theoretical claim, a proof or reference is needed.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The main issue is the unmatched comparator: the reported efficiency gain could be an artifact of the deliberately stronger shrinkage on the fusion treatment-effect forest. This is fixable with a matched-prior re-analysis and should be the primary request. The transportability concern is real but the authors already restrict the population to make it plausible; a sensitivity analysis would make the application's certainty statements more defensible. I see no evidence of circularity in the identification argument, and the simulation infrastructure is solid. The revision is substantial but within scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper is worth engaging with. The method is a real first — fusion for interval-censored survival with unmeasured confounding in the RWD — but the central precision-gain claim is not yet backed up, because the trial-only comparator is not hyperparameter-matched.\n\nThe identification setup is clean. Proposition 1 is just a decomposition, and Proposition 2 shows tau is identified from the RCT alone while the RWD identifies only tau+c. That’s exactly the right structure for the confounding-function approach, and it’s stated without hand-waving. The HDPM error model is a sensible choice for source-specific but linked residual distributions. The simulation is extensive: 1000 replications, bias/coverage/variance reported across confounding and heterogeneity grids, and the code is public. Those parts are solid.\n\nThe soft spot is the stress-test concern: the fusion uses a deliberately stronger shrinkage prior on tau (J=100, k=1/2, beta=3) while the 'single-source counterpart' is presumably a default BART forest (J=200, k=1, beta=2). The paper never specifies the comparator’s hyperparameters. The reported 62% reduction in interval width and 88% reduction in variance in the application, and the variance ratios in the simulations, may be prior shrinkage rather than borrowed information. That threatens the central claim that fusion helps because it borrows from the RWD. It’s checkable: refit the trial-only forest with the same tau prior, and see how much of the gain remains. The same applies to the simulation comparisons.\n\nThe transportability assumption is acknowledged as fragile, and the paper restricts both sources to a common population to make it plausible. That’s a reasonable attempt, but it remains untestable; a sensitivity analysis that varies the confounding function or down-weights the RWD would make the application claims more robust.\n\nVerdict: conditional. The method is coherent, the theory is sound as far as it goes, and the empirical work is thorough. But the headline efficiency gain needs to survive an hyperparameter-matched comparison before I trust it. The authors should also add a sensitivity analysis for transportability. With those fixes, this is a serious methodological contribution.\n\nWould I send it to reviewers? Yes. It deserves referee time. The core idea is novel and the execution is mostly careful — the flaw is in the evaluation, not the model. After a revision that addresses the prior-matching issue, I’d be happy to cite it.","headline":"A genuinely useful fusion framework for survival HTE, but the claimed efficiency gain is compromised by an un-matched prior comparison.","tokens_in":33312,"tokens_out":2477,"would_cite":true,"duration_ms":26015,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62G08","62N01","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that combining a randomised trial with confounded real-world survival data yields per-patient treatment effects with trial-level validity and roughly half the uncertainty, provided the true effect transports across sources.","keywords":["Bayesian fusion forest","heterogeneous treatment effects","survival analysis","data fusion","confounding function","accelerated failure time","Bayesian additive regression trees","interval censoring"],"falsifier":"Simulate two sources where the true CATE differs by covariates (e.g., τ_RWD = τ_RCT + g(X) with nonzero g), then fit the fusion forest: the posterior mean of τ will show bias that grows with the RWD sample size, while a trial-only fit stays unbiased. A real-data check would compare fusion estimates against an independent large randomised trial in a population where the RWD composition differs.","tokens_in":32302,"feed_emoji":"🩺","tokens_out":4449,"duration_ms":44182,"temperature":0.7,"pith_summary":"The paper develops a Bayesian nonparametric model for estimating heterogeneous treatment effects on censored survival times by fusing a randomised controlled trial with real-world data. It argues that an accelerated-failure-time decomposition — baseline prognosis, source-specific deviation, treatment effect, and a confounding function — lets the real-world data contribute power and follow-up while the trial pins down the causal effect. The confounding function absorbs whatever bias the real-world treatment assignment carries, so the model never needs to assume the observational source is unconfounded. In simulations the fusion stays unbiased and roughly halves posterior variance compared with trial-only analysis. Applied to HIV antiretroviral therapy, it estimates an average acceleration factor of 1.65 and a posterior probability of benefit above 0.95 for 96.4% of patients, versus 38% from the trial alone.","feed_headline":"Fusing trial and real-world data halves per-patient uncertainty","feed_subtitle":"Keeps effects unbiased while borrowing real-world power; HIV benefit certainty rises from 38% to 96%.","key_machinery":"The load-bearing object is the accelerated-failure-time decomposition (Eq. 3) with four separate Bayesian additive regression tree priors: a shared baseline prognosis, a source-specific deviation centred at zero, a depth-penalised treatment-effect forest, and a strongly regularised confounding function. The confounding function is the mechanism that absorbs RWD bias; the shared baseline with zero-centred deviation encodes borrowing between sources; the hierarchical Dirichlet process mixture lets the error law differ across sources while sharing mixture components. A blocked Gibbs sampler with data augmentation handles right- and interval-censored event times.","core_discovery":"Assume consistency, RCT unconfoundedness, positivity, non-informative censoring, and cross-source transportability of the conditional treatment effect. Then the conditional mean log survival time decomposes as E[logT|A,X,S] = m0(X,S) + τ(X)A + (1-S)Ac(X), where τ is the CATE and c is a confounding function active only in the real-world data. Proposition 2 shows τ is identified from the RCT alone and c is identified given τ, while the RWD alone identifies only the sum τ+c. The Bayesian fusion forest places separate tree-ensemble priors on each component, with a hierarchical Dirichlet process mixture for the error distribution, and thereby estimates τ(x) with lower variance than a trial-only a","pith_inferences":["If transportability fails, the RWD contribution will pull the posterior of τ toward the confounded contrast; Assumption 4 is the fragile hinge, so sensitivity analyses comparing fusion with trial-only where the two disagree most would be valuable.","The time-invariance of the acceleration factor is an explicit limitation; a time-dependent variant would let delayed-onset or waning effects be captured.","The same decomposition could fuse multiple RWD sources, one confounding function per source, as long as τ transports to all of them.","The method's per-patient 'certainty to benefit' depends on the calibration established in simulations; in a single real cohort without ground truth, the 96.4% figure inherits that calibration assumption."],"forward_implications":["Fusion estimates remain unbiased under unmeasured confounding in the RWD and stay at nominal coverage.","Posterior variance is roughly half the trial-only variance across simulation settings; RMSE ratio stays below one up to p=500 covariates.","The method opens right- and interval-censored outcomes to data fusion.","In the HIV application, average acceleration factor 1.65 [1.43; 1.90], and 96.4% of patients have posterior probability of benefit >0.95 vs 38% trial-only.","A regression-tree projection of the treatment forest identifies CD8 and race as effect modifiers."],"fun_headline_variants":["Fusing trial and RWD raises HIV benefit certainty to 96%","Bayesian fusion forest finds HIV benefit for nearly all patients","Combining RCT and real-world data reduces survival treatment effect uncertainty","Data fusion sharpens HIV survival benefit estimates from mixed sources"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole edifice rests on Assumption 4: that for every covariate profile, the average treatment effect is the same in the trial and the real-world population. This cannot be tested from the data, and if it is wrong, the real-world data drag the estimate toward a biased contrast.","fun_headline_variants_meta":{"raw":{"variants":["Fusing trial and RWD raises HIV benefit certainty to 96%","Bayesian fusion forest finds HIV benefit for nearly all patients","Combining RCT and real-world data reduces survival treatment effect uncertainty","Data fusion sharpens HIV survival benefit estimates from mixed sources"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000592,"raw_usage":{"total_tokens":2609,"prompt_tokens":742,"completion_tokens":1867,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":1796}},"tokens_in":486,"tokens_out":1867,"duration_ms":15036,"temperature":1.0,"reasoning_tokens":1796,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T09:47:42.516019+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate two sources where the true CATE differs by covariates (e.g., τ_RWD = τ_RCT + g(X) with nonzero g), then fit the fusion forest: the posterior mean of τ will show bias that grows with the RWD sample size, while a trial-only fit stays unbiased. A real-data check would compare fusion estimates against an independent large randomised trial in a population where the RWD composition differs.","supporting_citations":[],"review_version":1}