{"id":"b264d745-5541-49d3-bdff-0563124cb47e","arxiv_id":"2510.19243","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A targeted federated estimator for heterogeneous treatment effects, combining doubly robust scores with density-ratio weighting and a bootstrap source-selection step.","lead":"This paper proposes a privacy-preserving federated learning method to estimate how treatment effects vary across patient subgroups using data from multiple hospitals without sharing patient records. It adds a bootstrap step to drop data sources whose results would pull the estimate away from the target population.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (4) as printed omits external-site OR centering and has a weight mismatch, so Theorem 2's double-robustness does not follow; the consistency/efficiency claims need corrected equations or a proof.","rationale":"The paper's central claim is the doubly robust consistency of the federated estimator in Theorem 2. The reader flagged Assumption 8 as the weakest assumption, and that is indeed a substantive transportability condition. But the more immediately load-bearing issue is internal: the estimating equations as displayed do not support the theorem even when Assumption 8 holds. Under branch (ii), external source contributions are pure IPW-tilted residuals without the outcome-regression centering that is present in the target-only estimator; their expectation is not zero under OR misspecification. Under branch (i), the external terms are zero-mean and beta-free, which is inconsistent with the reported efficiency gains. This mismatch means either the theorem, the equations, or the simulation implementation is incorrect as written. The simulations in Table 1 Setting III may appear to support branch (ii), but under the printed P_m the OR-correct Setting I could not produce the dramatically smaller MCSD shown for Fed-SS, so it is likely that the simulations implement a different estimator than Eq. (4). This is not a critique of the underlying idea, which may be salvageable with an augmented P_m, but a request for the corrected estimating equations and a proof. I therefore keep the reader's CONDITIONAL verdict: acceptance should require the authors to reconcile Eq. (4) with Theorem 2, provide the deferred proof, and ideally rerun the simulation with the exact printed estimator.","tokens_in":15668,"tokens_out":16563,"duration_ms":151516,"concrete_test":"Implement exactly the printed Eq. (4) in the paper's Setting III (OR misspecified, PS and density ratio correctly specified), with n1=100, K_s=10, and 1000 Monte Carlo repetitions. If Theorem 2 branch (ii) is correct, the Fed-SS bias for the A_i*X_{1i} coefficient should vanish. Our calculation predicts a nonzero asymptotic bias; if the realized |bias| exceeds 2*MCSE, the printed theorem is false. Alternatively, compute E[U(beta0)] analytically for a minimal two-site normal-linear example with an omitted covariate; if E[U(beta0)] != 0, Eq. (4) cannot identify beta0 under branch (ii). If the authors' code instead uses augmented external P_m and reproduces Table 1, then the printed equations are not the estimator being analyzed and must be corrected and reproven.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The source-site estimating functions in Eq. (4) are printed as P_m = (1/n_m) sum_i eta tau_m I(A=a)/pi_m (Y - g_m), with no (g_m - mu_beta) centering; the only centering term Q1 is for the target site and is not multiplied by w1. Consequences: (i) under Theorem 2 branch (i) (all OR models correct), E[P_m]=0 and P_m does not depend on beta, so external blocks are asymptotically zero-mean, beta-free noise; they cannot produce the variance reduction reported for Fed-SS in Table 1. (ii) under branch (ii) (PS and density-ratio models correct, OR misspecified), taking expectations under Assumptions 5-8 gives E[P_m] = E_{f1}[eta(Y(a) - g_m^a(X))], which is not zero when g_m is misspecified. The target block at beta0 has expectation E_{f1}[eta(Y(a)-mu_{beta0})]=0, so the root of (4) is shifted by sum_{m in S} w_m E_{f1}[eta(Y(a)-g_m^a(X))] unless the external ORs are also correct. Thus only the OR-correct branch is supported by the displayed equation. The proof of Theorem 2 is deferred to Supplementary Section 1.5 and is not available in the manuscript, so the required cancellation cannot be checked. If the intended estimator includes external centering terms (g_m - mu_beta) and correct weights, the manuscript must display those corrected estimating equations; as written, Theorem 2 and the associated efficiency claims are not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a targeted-federated learning framework for estimating heterogeneity of treatment effects (HTEs) in a prespecified target population. The target estimand is defined through a working structural model whose parameters solve a projection moment condition (Eq. 2). The authors develop a target-only doubly robust estimator (Eq. 3) and a federated estimator (Eq. 4) that combines target and external-site contributions, using density-ratio weighting to adjust for covariate shift and a bootstrap-based procedure to select transportable sources. The paper claims double robustness for the federated estimator (Theorem 2), reports variance reductions in simulations, and applies the method to Medicare hip-fracture data.","tokens_in":16104,"tokens_out":10115,"duration_ms":84405,"significance":"If the theoretical claims are correct, this is a potentially useful contribution: it provides a projection-based estimand for HTEs that accommodates binary/count outcomes, operates under federated privacy constraints with one round of communication, and includes a practical source-selection procedure. The target-only estimator is standard and the density-ratio approach is sensible. The simulation studies and real-data application are extensive and the proposal addresses a genuine gap in the federated causal inference literature. However, the central theorem's proof is deferred to unavailable supplementary material and, more importantly, the displayed estimating equation appears inconsistent with the stated double-robustness property, so the main claims are not currently verifiable.","major_comments":[{"comment":"The external-site estimating functions P_m for m∈S omit the centering term (g_m − l^{-1}{η^T β_F}) that is present for the target site in Q_1. Consequently, E[P_m] at the true projection parameter equals E_{f_1}[η(Ỹ,a){E[Y(a)|X] − g_m(a,X)}], which is nonzero under branch (ii) of Theorem 2 (PS and density-ratio correct, OR misspecified). Only branch (i) — all OR models correct — is supported by the printed equation. The proof is deferred to Supplementary Section 1.5, which is not included, so the cancellation central to double robustness cannot be checked. The equations must be corrected to include external centering (and appropriate weights), or Theorem 2 and the associated efficiency claims are unsupported.","section":"§2.3.2, Eq. (4)"},{"comment":"Under Setting I (all OR and PS correct), Eq. (4) implies the external P_m are zero-mean and do not depend on β_F. Adding independent zero-mean terms cannot produce the substantial MCSD reductions relative to Target-only reported in Table 1 (e.g., interaction MCSD 64.9 → 50.5 at n1=100). Under Setting III (OR misspecified, PS/DR correct), the printed equation predicts a bias equal to Σ_{m∈S} w_m E_{f_1}[η{Y(a)−g_m^a(X)}] at the federated root, yet the reported Fed-SS bias for A_i*X_1i is 0.9 (n1=100). These discrepancies indicate that the implemented estimator differs from the printed Eq. (4), so the simulation results cannot validate the method as presented.","section":"§3, Table 1"},{"comment":"The bootstrap selection procedure constructs β_m^{(b)} by setting w_m=1 in Eq. (4). For external sites this is not, in general, a consistent estimating equation for β_0 even under transportability unless the site's OR is correct (see the Eq. (4) comment). The paper acknowledges that the procedure's validity relies on the consistency of the target-only estimator and lists a formal theoretical assessment as future work, but then presents Fed-BS as a validated safeguard against negative transfer. Given the missing theory and the dependence on the problematic estimating equation, the selection procedure should be presented as a heuristic; its simulation performance cannot be interpreted as support without a proof or a corrected equation.","section":"§2.4"}],"minor_comments":[{"comment":"The notation is confusing: K is used both as the set of source indices and as the number of external sources (e.g., 'K+1 sources' vs. 'm∈K'). Consider using a script letter for the set and k for the count.","section":"§2.1"},{"comment":"The weights w_m are defined only after the estimating equation is displayed. State the normalization (Σ w_m = 1, w_m ≥ 0) before the equation, and note that the target-only case w_1=1 is one of many possible choices.","section":"§2.3.2, Eq. (4)"},{"comment":"The relationship between r(X) and ṟ(X) is not clearly stated. If they are the same (e.g., r(X)=X), say so; otherwise explain what distinguishes them. The description in Section 2.4 suggests only covariate means are shared, so clarify whether ṟ is a subvector.","section":"§2.3.3, Eq. (5)"},{"comment":"The definition of non-transportable sources is asymmetric: half the sources have misspecified PS and OR models, so it is unclear whether the Fed-BS improvement comes from detecting OR misspecification or PS misspecification. A cleaner experiment would vary one of the two at a time.","section":"§3, Table 2"},{"comment":"In the real-data application, all source datasets were identified as transportable, so the bootstrap selection procedure's utility in that example is not demonstrated. Consider a sensitivity analysis where some sources are excluded or artificially made non-transportable to illustrate the procedure.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The central issue is whether the estimating equations in the manuscript match what is actually implemented and proven. The discrepancy between Eq. (4) and the simulation outcomes (Setting III, Table 1) suggests a serious typographical or conceptual error. I would require the authors to provide the correct estimating equations, the full proof of Theorem 2 (not just a pointer to a missing supplement), and a reproducible check of Table 1 Setting III before the paper can be considered further. If the proof is genuinely unavailable, the double-robustness claim should be withdrawn."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper targets a real gap — federated estimators for effect modification rather than just ATE — and it ships a clean projection estimand, a one-round communication scheme, extensive simulations, and a Medicare data application. The target-only doubly robust estimator (Eq. 3) is correct and standard. The problem is the federated estimating equation (Eq. 4) as printed: external-site blocks P_m contain only τ I(A=a)/π (Y - g_m), with no (g_m - μ_β) centering. Consequently the claimed double robustness in Theorem 2 isn't supported. Under the OR-correct branch, each P_m is zero-mean and beta-free noise; under the PS/density-correct branch with misspecified OR, each external block has expectation E_{f1}[η(Y(a)-g_m)], which cancels only if all sites share the same misspecified OR limit. The theorem as stated allows arbitrary misspecified ORs, so this is a load-bearing gap — unless the supplementary proof shows something not visible in the manuscript. Also, the variance reductions reported in Table 1 don't follow from Eq. (4): with weights proportional to sample size and n1=100, the target block is nearly absent and external blocks carry no beta information, so the reported efficiency gain is unexplained. That suggests either the equation is misprinted or the implementation used something else.\n\nWhat's genuinely good: the projection estimand with link functions is a sensible way to handle binary/count HTEs; the bootstrap source-selection procedure is practical and does mitigate negative transfer in simulations; the real-data analysis is careful. The paper is clearly written and the simulation design is thorough.\n\nSofter spots: the selection rule depends on target-only consistency, which the paper acknowledges; Assumption 8 (mean exchangeability across sources) is strong, though the selection step is a reasonable attempt to cope. The main issue is the federated theorem, not the target-only part.\n\nBottom line: worth a serious referee, but I would not accept the central claim until the authors correct Eq. (4) (likely adding Q_m terms for external sites) or provide a proof in the main text. If that's fixed, this is a useful contribution. As is, it's a conditional accept / major revision.","headline":"Useful federated HTE proposal with a real gap: Eq. (4) as printed doesn't support the claimed double robustness, so the central theorem needs either corrected equations or a proof.","tokens_in":16551,"tokens_out":7423,"would_cite":false,"duration_ms":65093,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A federated estimator can recover heterogeneous treatment effects for a target population without sharing patient-level data, staying consistent under either of two model specifications.","keywords":["federated learning","heterogeneity of treatment effects","doubly robust estimation","density ratio model","propensity score","bootstrap selection","covariate shift","categorical outcomes"],"falsifier":"Simulate K=20 sites satisfying Assumptions 1–8 with shift in covariate distributions, misspecify all outcome regressions, specify all propensity scores and the density-ratio model correctly, and check that the estimator's bias decreases with sample size; then misspecify the density-ratio model while keeping propensity scores correct and confirm bias appears. This isolates the second robustness branch and tests whether the density-ratio calibration is doing the claimed work.","tokens_in":15600,"feed_emoji":"🔐","tokens_out":6646,"duration_ms":55689,"temperature":0.7,"pith_summary":"This paper tries to establish that heterogeneity of treatment effects (HTE) can be estimated in a federated setting—where patient-level data never leave their source institutions—with one round of communication and without bias from covariate distribution shifts. The authors define HTE through a projection-based working structural model, then construct an estimator that is consistent if either the outcome regressions are correct at every site, or the propensity scores and density-ratio models are correct at every site: a doubly robust property. A bootstrap-based selection step is added to detect and exclude non-transportable data sources, preventing negative transfer. The method is demonstrated in simulations and on Medicare hip-fracture cohorts, where it recovers known effect modification (e.g., sex) with narrower intervals than using the target site alone. A sympathetic reader would care because it makes multi-site precision-medicine analyses feasible under privacy constraints. The paper also candidly flags that the selection step's validity depends on a consistent target-only estimator (Section 2.4) and that formal theory for the bootstrap selector is not developed (Section 5).","feed_headline":"Federated HTE estimator stays consistent when either model branch is right","feed_subtitle":"One round of site-level summaries plus a bootstrap screen that drops non-transportable sources.","key_machinery":"The argument is carried by three interacting pieces: (1) a working structural model l{E(Y(a)|X~,M=1)} = η(X~,a)^T β whose coefficients are defined as a projection via moment condition (2), giving a target-population-specific, interpretable HTE estimand; (2) a density-tilting term τ_m(X;α_m)=exp(α_m^T r(X)) that reweights each source's covariate distribution to match the target's, estimated by moment matching in Equation (5); and (3) the doubly robust estimating equation (4) that combines tilted source-specific score contributions with the target's own augmented score. A bootstrap selection procedure around this estimator screens out non-transportable sources by comparing each source's estima","core_discovery":"The central claim is Theorem 2: the targeted-federated estimator obtained by solving Equation (4) is consistent for the projection parameter that best approximates the conditional treatment effect in the target population, provided either every site's outcome regression is correctly specified, or every site's propensity score and every source's density-ratio model is correctly specified. This makes the estimator doubly robust in a federated setting where covariate distributions differ across sites, and it is achieved with a single round of communication in which no individual-level data is exchanged.","pith_inferences":["The projection estimand is working-model-dependent: if the chosen η(X~,a) is far from the truth, the target parameter itself changes. A natural extension is to let the federated protocol compare multiple working models or use a more flexible basis for η.","The bootstrap selection procedure can only catch non-transportability that shifts the federated estimate away from a consistent target-only estimate; if the target site itself is confounded or its model is misspecified, the screen may retain biased sources. A sensitivity analysis that assumes a range of target-only biases would be a direct extension.","The exponential tilting form of the density ratio is a parametric restriction; when it is misspecified, the second robustness branch loses its guarantee. A diagnostic based on a richer moment set (e.g., second moments) would make the shift calibration testable in practice.","The single-round design could be extended to adaptive re-weighting across rounds if an initial pass reveals some sources near the decision boundary, trading extra communication for more precise selection."],"forward_implications":["If correct, multi-site HTE studies can pool evidence across institutions without sharing patient-level data or requiring more than one round of communication.","The projection estimand gives interpretable effect-modification parameters (e.g., log-odds ratios for binary outcomes) even when the true conditional effect is not exactly linear in the working model.","The bootstrap selection step provides a practical guard against negative transfer; in the Medicare application it identified all 2010–2016 sources as transportable to the 2017 target while improving precision over target-only analysis.","The double robustness means a site need only get one branch right (outcome regression, or propensity plus density ratio) for its data to contribute without bias.","Communication cost is one round: a target summary (covariate means) goes out, each source returns a scalar or vector score term, and the target solves the joined estimating equation."],"fun_headline_variants":["Doubly robust federated HTE without raw data sharing","One-round federated learning for effect heterogeneity","Bootstrap screen removes non-transportable sites for HTE","Targeted federated estimator tolerates model misspecification","Privacy-safe HTE that works when either model is right"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The treatment effect conditional on the measured covariates must be identical across all data sources (Assumption 8); the paper's own selection procedure also assumes the target-only estimate is consistent, a point it flags in Section 2.4.","fun_headline_variants_meta":{"raw":{"variants":["Doubly robust federated HTE without raw data sharing","One-round federated learning for effect heterogeneity","Bootstrap screen removes non-transportable sites for HTE","Targeted federated estimator tolerates model misspecification","Privacy-safe HTE that works when either model is right"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000157,"raw_usage":{"total_tokens":1038,"prompt_tokens":704,"completion_tokens":334,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":448,"completion_tokens_details":{"reasoning_tokens":269}},"tokens_in":448,"tokens_out":334,"duration_ms":3931,"temperature":1.0,"reasoning_tokens":269,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T08:44:06.610807+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate K=20 sites satisfying Assumptions 1–8 with shift in covariate distributions, misspecify all outcome regressions, specify all propensity scores and the density-ratio model correctly, and check that the estimator's bias decreases with sample size; then misspecify the density-ratio model while keeping propensity scores correct and confirm bias appears. This isolates the second robustness branch and tests whether the density-ratio calibration is doing the claimed work.","supporting_citations":[],"review_version":1}