{"id":"987620a5-9967-4152-8de6-2d6cbc1a6956","arxiv_id":"2607.02787","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Fusing trial and real-world data improves precision only for small real-world bias and small trials, and only a conservative block-jackknife interval gives honest uncertainty for that gain.","lead":"This paper asks how much precision trial-plus-real-world-data fusion really buys, using the adaptive-TMLE estimator. It delivers an audit card for the learned bias model, a simulation map of when fusion helps, and a conservative interval that decides whether fusion or the trial alone should be primary — a practical guardrail for analysts and regulators.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The block-jackknife recommendation is selected post hoc from ten candidates and validated on the same experiments; its status as the unique valid SE is not independently confirmed, so the guardrail decision rule may be overfit.","rationale":"The reader's verdict is already CONDITIONAL, citing the block-jackknife selection issue as one of three addressable concerns. My stress-test identifies this as the single most load-bearing issue because the practical guardrail — the rule that fusion should be primary only when the lower block-jackknife limit exceeds one — rests entirely on the block jackknife being calibrated outside the experiments used to choose it. The efficiency-map scope caveat (constant enrollment positivity, homogeneous CATE) is important but is explicitly disclosed in the paper and is partially probed by robustness slices; the block-jackknife selection is not similarly disclosed as a selection-on-evaluation-data problem. The proposed concrete test would settle whether the block jackknife's superiority is robust or an artifact of overfitting. Since this concern reinforces the reader's CONDITIONAL verdict rather than changing it, the verdict remains UNCHANGED.","tokens_in":36912,"tokens_out":8556,"duration_ms":102737,"concrete_test":"Run an independent out-of-sample validation on a new simulation family not used anywhere in the paper: e.g., B=500 datasets with a binary outcome, 8 correlated covariates, W-dependent enrollment Π(W)=expit(0.5W1+0.8W2^2), external size 3× trial, n_rct=300, wiggly bias at m=2 and m=4. Apply all ten SEs exactly as defined in Web Appendix C, and score fixed-truth coverage against a B=2000 locked truth computed on the same new DGP. If the block jackknife's coverage is ≥0.95 in all cells, the post-hoc-selection concern is resolved; if not, the recommendation must be re-scoped or a pre-registered validation is needed before the decision rule is adopted as a general reporting standard.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central decision rule — claim an efficiency improvement only when the lower block-jackknife limit exceeds one — depends on the block jackknife being at least nominally calibrated in the deployment setting. The evidence is Section 4.6 (Table 8, Figure 4): the block jackknife is the only one of ten SEs achieving fixed-truth coverage 0.984–0.998 across the 15-cell grid and robustness slices. But the same data were used to select the block jackknife from the ten candidates; reporting the best candidate's coverage on the selection data is subject to winner-selection inflation. The paper explicitly disclaims a theorem for this non-smooth statistic (Section 4.6), so the entire warrant is empirical. The robustness slices — W-dependent enrollment, 4× external arm, degree-2 HAL, n=2000, binary outcome (Web Appendix E) — are additional cells in the same DGP family and were evidently examined before finalizing the recommendation, so they do not break the selection loop. The real-data verdicts in Section 5 hinge on these intervals being calibrated (e.g., the one interval clearing one is [1.02,1.45]); if the jackknife's true coverage is lower outside the studied grid, the guardrail could either miss real gains or over-claim them. This is not an accusation of p-hacking; it is a structural limitation of an empirical calibration claim with no independent validation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses adaptive targeted maximum likelihood estimation (A-TMLE) as a worked example of adaptive trial-plus-real-world-data fusion and develops three tools: (1) a report card that audits the learned bias model via surface recovery, influence-curve variance attribution, and targeting drift; (2) a simulation-based efficiency map of the finite-sample variance gain of A-TMLE relative to a matched RCT-only AIPW, with the gain driven mainly by bias magnitude rather than functional complexity, crossing parity near a bias of about one residual SD and eroding as the trial grows; and (3) selection-aware inference for the efficiency gain, culminating in a delete-a-fold block jackknife as the only one of ten candidate standard errors with consistently near- or above-nominal coverage. The tools are applied to six real-data fusions from three openly available trials; in five of six the block-jackknife interval includes one, so the paper recommends keeping the RCT-only estimate primary, with only one marginal exception.","tokens_in":37158,"tokens_out":6930,"duration_ms":72657,"significance":"If correct, this is a valuable and timely contribution. The paper is unusually candid about its scope: the headline map is restricted to constant enrollment probability and constant CATE, the block-jackknife result is explicitly empirical rather than a theorem, and the real-data illustrations are partly based on constructed external arms. The simulation work is extensive (15,000 A-TMLE fits with zero failures; 40 fixed-truth cells for the standard-error head-to-head) and the analysis pipeline is reproducible from public code. The identification of a conservative selection-aware interval for the efficiency gain, if independently validated, would be an important guardrail for practice and a useful caution against naive influence-function standard errors after model selection. The paper is well within the scope of the journal and deserves serious consideration.","major_comments":[{"comment":"The block jackknife is recommended after examining its coverage on the same 15-cell grid and the robustness slices. Because it was selected as the best of ten candidates using those same results, the reported coverage (0.984–0.998) is a selected maximum and may be optimistic; the additional robustness cells do not break the selection loop because they were evidently examined before the recommendation was finalized. This matters because the Section 3.2 decision rule and the Section 5 primary-analysis verdicts rest entirely on this empirical calibration, and the paper explicitly provides no theorem for this non-smooth statistic. I ask for an independent validation strategy — e.g., a pre-specified holdout set of DGP configurations not examined during method selection, or a separate selection/validation split of the simulation grid — or, failing that, a clear re-labeling of the guardrail as","section":"Section 4.6, Table 8, Figure 4"},{"comment":"The abstract states the efficiency-map findings ('driven mainly by the magnitude of the real-world bias... crosses break-even near a moderate bias and erodes as the trial grows') without the scope caveat that the main grid assumes constant trial-enrollment probability and a constant within-trial effect (Section 2 scope caveat; Section 4.1). Under these restrictions the matched GLM reference is correctly specified on the trial arm, so the comparison is a pure variance contrast. The paper itself notes that under a heterogeneous effect the matched-GLM gain sits at parity even at zero bias, and under selective enrollment A-TMLE's own ATE interval degrades at the rough large-bias corner (Section 4.5). The abstract and the title's promise of a general 'when' map should carry the same caveat; otherwise readers may take a design-specific map as a general characterization.","section":"Abstract; Sections 2, 4.3, 4.5"},{"comment":"The population identity var(D_A)=a+b m^2 is derived under a forced intercept-only oracle working model (Φ≡1), which is not the estimator deployed in the simulations or real data. The authors are careful to call this an oracle restriction, but the abstract's phrase 'a dominance an exact population-oracle variance identity explains' could oversell the explanation: the identity does not cover the HAL-selected working model, and the finite-sample shape ordering is actually reversed. The paper addresses this, but the abstract and the 'takeaway' would be clearer if the identity were described as a diagnostic analogue under a restricted oracle model rather than 'the' explanation of the empirical magnitude-dominance finding.","section":"Section 4.4, Proposition 1"}],"minor_comments":[{"comment":"Please add a sentence stating that the efficiency map and the parity crossing are established under constant trial enrollment and a homogeneous within-trial effect, with secondary robustness checks for W-dependent enrollment and heterogeneous CATE.","section":"Abstract"},{"comment":"The column 'Cover., own mean' is explicitly circular; consider adding '(circular)' to the column header or footnote so that the diagnostic purpose is immediately clear and not misread as a valid coverage estimate.","section":"Table 8"},{"comment":"The sentence about the population shape ordering being the reverse of the empirical finite-sample ordering is easy to misread. Please spell out in one or two sentences that Proposition 1 concerns the oracle intercept-only influence curve, while the deployed HAL-selected estimator is affected by finite-sample basis selection and nuisance estimation error.","section":"Section 4.4"},{"comment":"The basis-count contrast between PSID (one basis) and CPS (seven bases) is a nice illustration, but the text already notes the count is not monotone in complexity and may reflect power; consider adding a one-sentence reference to Section 4.3 near the table so the reader does not over-interpret the count.","section":"Section 5.3, Table 12"},{"comment":"There are occasional tense shifts between 'we show' and 'we do not establish' that make it harder to track which claims are the paper's contributions versus caveats. A short 'evidence class' table (as in Table 13) for Section 4.6 would help.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid, honest, and reproducible methodological evaluation. My main concern for the editor is that the central guardrail — the block-jackknife decision rule — rests on an empirical calibration result that was selected post hoc from ten candidates and validated on the same experiments used for selection. This is fixable in revision with an independent validation strategy or a more provisional claim. The scope mismatch between the abstract and the strict design limits is also fixable. I would not recommend rejection; the simulation work and the real-data demonstrations are substantial and within the journal's remit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know first: this is a solid, honest paper about when RCT+RWD fusion with A-TMLE actually buys precision. The headline finding is that the gain is finite-sample and reference-relative: it crosses break-even near a bias of about one residual SD, erodes as the trial grows, and is driven more by bias magnitude than functional complexity. The paper also builds a report card for the learned bias model and shows that among ten candidate standard errors for the gain, only a delete-a-fold block jackknife gives near- or above-nominal coverage in their experiments.\n\nWhat's genuinely new: Proposition 1 (var(D_A)=a+bm^2) is a clean population result that explains why magnitude dominates, and the report card is a simple idea that could catch on. The simulations are extensive and reproducible — 15,000 fits, zero failures, multiple robustness slices — and the paper is unusually candid about what it does not establish: it explicitly disclaims a theorem for the jackknife and labels the main map as strictly constant-positivity.\n\nThe soft spots are real but not fatal. The biggest is the post-hoc selection of the block jackknife: it's the only one of ten that works in their grid, and the same experiments were used to pick it. The stress-test note is right that this is winner-selection inflation, and there is no independent validation. That said, the gap is large — no rival reaches 0.87 coverage against fixed truth while the jackknife sits at 0.98–1.00 across 40 cells — so the conclusion is unlikely to flip completely. Still, the recommendation would be stronger with an out-of-sample check or a theoretical argument.\n\nThe main efficiency map is also narrower than the abstract implies: it's established under constant enrollment probability and homogeneous CATE, with W-dependent enrollment only in secondary slices. The paper says this plainly, but the abstract's phrasing is broader. Similarly, the real-data evidence is thin: four of six fusions use constructed within-trial external arms, and only one interval clears one, marginally.\n\nWho this is for: anyone working on trial+RWD fusion or adaptive debiased ML. It's a genuine contribution to the practice of reporting such analyses. It deserves a serious referee — there are addressable issues but nothing load-bearing. I'd send it to review with a request for a more careful handling of the selection-inflation problem.\n\nRecommendation: yes, engage with it. It's worth citing and discussing.","headline":"A genuinely useful, unusually candid empirical-methods paper that maps when A-TMLE fusion actually pays; the block-jackknife guardrail is the most valuable and also the most fragile piece.","tokens_in":37807,"tokens_out":2355,"would_cite":true,"duration_ms":25331,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The efficiency gain from adding real-world data to a trial is real but narrow: with A-TMLE it crosses break-even near one residual SD of bias, erodes as the trial grows, and survives a block-jackknife interval in only one of six real fusion","keywords":["adaptive targeted maximum likelihood estimation","trial–real-world data fusion","efficiency gain","selection-aware inference","block jackknife","bias magnitude","highly adaptive lasso","real-world evidence"],"falsifier":"Re-run the main 15-cell grid with W-dependent trial enrollment and a heterogeneous within-trial effect (for example, CATE = 1.5 + 0.8W1 − 0.5W2) at n_rct = n_ext = 250. If the efficiency gain then crosses parity at m < 0.5, or if functional complexity rather than magnitude orders the cells, or if the block-jackknife interval's coverage drops below nominal in the low-bias cells, the paper's headline map and selection-aware verdict are contradicted.","tokens_in":36631,"feed_emoji":"📊","tokens_out":4874,"duration_ms":46356,"temperature":0.7,"pith_summary":"Using adaptive targeted maximum likelihood estimation (A-TMLE) as a worked example, this paper asks when fusing a randomized trial with real-world data actually narrows the confidence interval for the treatment effect, and how to state that gain honestly. It finds that the gain is governed mainly by the magnitude of the real-world bias, not by the complexity of the bias function: the variance of the estimator rises quadratically with bias magnitude, so the efficiency ratio crosses parity near a bias of about one residual standard deviation and falls badly at larger bias. The gain also shrinks as the trial sample size grows, so it is finite-sample rather than a super-efficiency result. The paper then treats the gain itself as a data-adaptive estimand and shows that, among ten candidate standard errors, only a block jackknife that re-selects the working model gives near- or above-nominal coverage; the naive standard error undercovers. In six real-data fusions, that conservative interval keeps the trial-only analysis primary in five of six — the toolkit acts as a guardrail, not a booster.","feed_headline":"Trial fusion gains fade once real-world bias tops one residual SD","feed_subtitle":"A block-jackknife interval keeps five of six real fusions with the RCT-only estimate; the naive standard error overclaims.","key_machinery":"Central object is the A-TMLE decomposition of the trial ATE into a pooled-projection estimand and a bias projection built from the learned enrollment-effect surface τ_S(W,A); the working model is a relaxed highly-adaptive-lasso (HAL) basis selected by cross-validation and then targeted. The efficiency gain R is the ratio of influence-curve variances, var(D_rct)/var(D_atmle), with D = D_A − D_S. The argument is carried by Proposition 1, an exact population-oracle identity var(D_A) = a + b m² showing bias magnitude enters quadratically as the leading variance driver, and by a delete-a-fold block jackknife that re-selects the working model on each leave-fold-out subsample, giving conservative n","core_discovery":"The paper's central claim is that the finite-sample efficiency of A-TMLE for RCT-plus-real-world fusion is modest and reference-relative. Writing the ATE as a pooled projection minus a learned bias-correction term, the estimator's influence curve is D = D_A − D_S, and the efficiency gain is R = var(D_rct)/var(D_atmle) relative to a matched, correctly-specified trial-only estimator. An exact population-oracle identity shows var(D_A) = a + b m² under a restricted working model: the bias magnitude m enters quadratically with no linear term and no shape dependence at that order, explaining why magnitude dominates complexity in the simulation map. The gain starts near 1.15 at zero bias, crosses o","pith_inferences":["If the quadratic magnitude-dominance identity extends beyond restricted oracle models, similar break-even maps should appear for other debiased estimators that borrow through a learned correction; that is a testable transfer, not something the paper establishes.","The one real fusion whose interval clears parity rests on a four-basis correction for a prognostically distant external arm, suggesting the method's value may be concentrated precisely in the large, strongly biased external cohorts the public examples do not contain.","A practical extension would be to convert the block-jackknife width ratio into a pre-study power or sample-size tool: the conservative interval implies that detecting a gain near 1.15 requires either very large trials or unusually small bias, which could inform whether to collect real-world data at all.","The report card's targeting drift, not the basis count, flagged the one coverage-failure corner; an analyst facing a single fusion with no comparison panel cannot yet use drift as a calibrated detector, so a calibration study on real data would be a natural next step."],"forward_implications":["Practitioners should report the efficiency gain with a block-jackknife interval and claim an efficiency improvement only when its lower bound exceeds one; the naive influence-function standard error is unsafe.","The break-even near one residual SD of real-world bias gives a concrete stopping rule: if the external cohort is expected to be biased at or beyond that scale, fusion is unlikely to buy precision at moderate trial sizes.","The asymptotic oracle super-efficiency guarantee of A-TMLE does not translate to a finite-sample advantage; the gain erodes as the trial grows and can be below one against an efficient trial-only reference.","The learned bias model should be reported as a stress diagnostic (effective basis count, variance attribution, targeting drift), not interpreted as a measure of confounding or a proxy for the gain.","When trial enrollment depends on covariates and the bias surface is rough and large, A-TMLE's own ATE interval can undercover with main-terms nuisances; flexible nuisance fits restore coverage, so the envelope matters."],"fun_headline_variants":["Fusion gain hinges on bias magnitude, not its shape","Block-jackknife interval keeps fusion honest","Only one of six real fusions beats RCT alone","Fusion edge is finite-sample, not super-efficiency","Naive SE overclaims fusion gain; jackknife covers"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline efficiency map assumes trial enrollment is completely random and the within-trial effect is homogeneous, so the matched trial-only estimator is correctly specified by construction and the comparison is a pure variance contrast; if enrollment depends on covariates or treatment effects vary, the break-even point and the magnitude-dominance ordering could shift.","fun_headline_variants_meta":{"raw":{"variants":["Fusion gain hinges on bias magnitude, not its shape","Block-jackknife interval keeps fusion honest","Only one of six real fusions beats RCT alone","Fusion edge is finite-sample, not super-efficiency","Naive SE overclaims fusion gain; jackknife covers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000695,"raw_usage":{"total_tokens":3059,"prompt_tokens":900,"completion_tokens":2159,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":644,"completion_tokens_details":{"reasoning_tokens":2080}},"tokens_in":644,"tokens_out":2159,"duration_ms":14313,"temperature":1.0,"reasoning_tokens":2080,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T08:55:58.220218+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the main 15-cell grid with W-dependent trial enrollment and a heterogeneous within-trial effect (for example, CATE = 1.5 + 0.8W1 − 0.5W2) at n_rct = n_ext = 250. If the efficiency gain then crosses parity at m < 0.5, or if functional complexity rather than magnitude orders the cells, or if the block-jackknife interval's coverage drops below nominal in the low-bias cells, the paper's headline map and selection-aware verdict are contradicted.","supporting_citations":[],"review_version":2}