{"id":"61005d24-81d0-456c-8dd0-9438e8cad3f0","arxiv_id":"2506.11482","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A doubly robust entropy-balancing estimator integrates external summary means with internal individual data, with data-adaptive entropy selection and efficiency diagnostics.","lead":"This paper develops a two-stage entropy-balancing estimator that borrows external summary statistics to estimate a population mean while guarding against bias from a mismatched external sample. The method includes a data-adaptive choice of balancing model and simple diagnostics that tell when borrowing is guaranteed to improve precision.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 1's existence proof relies on Lemma A.1, which fails when the external covariate mean lies outside the internal support; without a support-overlap condition, the C3-only arm of the double-robustness claim is unestablished.","rationale":"The reader's weakest_assumption was Condition (C1) in the form of unmeasured effect modifiers, which is a standard and acknowledged limitation. My concern is different and more directly tied to the correctness of the main theorem: even when (C1) holds, it does not guarantee that the external covariate mean is in the convex hull of the internal covariate sample. This affects whether the entropy-balancing estimator exists, not merely whether it is unbiased. The proof of Proposition 1 is the place where this is assumed, and the cited Lemma A.1 is false in the second-step application because the variables have a degenerate first coordinate; it is also insufficient for the first step because (C1) alone permits disjoint covariate supports. This is an internal gap in a stated proof, not a disagreement with external consensus. The missing bootstrap MSE selection criterion flagged by the reader is a real completeness issue, but it concerns a secondary contribution rather than the central consistency theorem. The variance estimator's acknowledged inconsistency is a limitation of inference, not of consistency. My proposed check is a small, concrete simulation that would settle whether the existence claim fails; if it fails, the paper needs a support-overlap condition or a restriction of Theorem 1 to the density-ratio arm. Because the estimator and theorem remain plausible under such a repair, the reader's CONDITIONAL verdict is appropriate; I therefore recommend UNCHANGED.","tokens_in":26968,"tokens_out":24070,"duration_ms":219190,"concrete_test":"Run a Monte Carlo check with the daisy package: draw internal X ~ Unif(0,1), n=200; external X ~ Unif(2,3), n1=2000, so the external summary mean is about 2.5; set Y = X + ε with ε ~ N(0,0.1), and define S|X with P(S=1|X=x)=0.5 for x∈(0,1) and P(S=1|X=x)=1 for x∈(2,3), so (C1) holds. Attempt the first-step entropy balancing in Equation (1) for each replicate. If the optimizer reports infeasibility (or fails to converge) in nearly every replicate, Proposition 1 is contradicted and Theorem 1 requires a support condition. Separately, analytically verify that Lemma A.1 fails for the second-step variables: for ε < 1/2, boxes centered at 1 ± 1.5ε contain no points whose first coordinate is 1, so the lemma's conclusion cannot hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Proposition 1, which underpins Theorem 1, asserts that the two-stage entropy-balancing weights exist with probability tending to one under (C1), (C2), (C4), and (C3) or (C3)'. Its proof applies Lemma A.1 to the second-stage variables \\tilde{H}^*_j = (1, H^*_j)^\\top. Lemma A.1 requires the sample to contain a point in each of 3^p boxes centered at μ*_1 + 3ε b/2. For the second step, the first coordinate of \\tilde{H}^*_j is identically 1, while boxes with b_1 = ±1 have first-coordinate intervals bounded away from 1 for small ε; the event then has probability 0. Thus the proof of Proposition 1 is invalid as written. The underlying existence question is also substantive: Condition (C1) does not imply that the external covariate mean lies in the support or convex hull of the internal X distribution. Example: internal X ~ Unif(0,1), external X ~ Unif(2,3), with P(S=1|X)=0.5 on (0,1) and P(S=1|X)=1 on (2,3). Here (C1) and (C3) hold (e.g., Y = X + ε), but no nonnegative weights can satisfy n^{-1}∑ w_i X_i = 2.5, so the first-step balancing problem (1) is infeasible. Hence Proposition 1's existence claim is false under the stated conditions, and Theorem 1's 'C3 only' side of double robustness is not established without an additional support-overlap condition (or the density-ratio assumption (C3)').","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage generalized entropy-balancing estimator for the internal-population outcome mean that borrows external summary statistics (means of X and Y) while explicitly allowing the external sample to be biased relative to the target. The first round of balancing calibrates internal covariates to the external covariate mean; the second round calibrates an outcome-model-based quantity to the external outcome mean, and the final estimator is the second-stage weighted mean of Y. The authors establish consistency under either a correctly specified linear outcome model (C3) or a correctly specified density-ratio model (C3)', prove asymptotic normality with an explicit variance formula, propose a data-adaptive selection rule among candidate entropy functions, and derive estimable efficiency criteria based on the Mahalanobis distance and Pearson chi-squared divergence. Simulations and a Japan Utstein registry application illustrate the method, and an R package daisy is provided.","tokens_in":27361,"tokens_out":5230,"duration_ms":57369,"significance":"If the main theorem holds, this is a practically valuable contribution: it uses only moment summaries from the external source, permits covariate shift, and provides double robustness plus easy-to-compute applicability diagnostics. The paper also ships reproducible software, gives detailed proofs of the asymptotic variance and efficiency criteria, and includes a substantive real-data analysis. These are genuine strengths that make the paper potentially publishable. However, the double-robustness claim is not established under the stated conditions because the existence proof of the balancing weights is flawed and the feasibility of the balancing equations is not guaranteed by Conditions (C1)-(C4). This gap affects the central consistency theorem, so the result needs substantial revision before the manuscript can be accepted.","major_comments":[{"comment":"The proof of Proposition 1 applies Lemma A.1 to the augmented variables \\tilde H^*_j = (1, H^*_j)^\\top with p=2. This application is invalid: the first coordinate of \\tilde H^*_j is identically 1, while Lemma A.1 requires the sample to contain a point in each of 3^p boxes centered at \\mu^*_1 + 3\\epsilon b/2 with b_1 \\in \\{-1,0,1\\}. For b_1 = \\pm 1 and sufficiently small \\epsilon, the first-coordinate interval is bounded away from 1, so the required event has probability zero. Thus the proof of Proposition 1 fails as written, and the existence of \\hat\\lambda_2 needs a different argument that handles the deterministic intercept coordinate.","section":"Appendix A, proof of Proposition 1"},{"comment":"The consistency claim under the C3-only arm is not established because Conditions (C1), (C2), and (C3) do not imply that the first-step balancing problem (1) is feasible. For example, let internal X ~ Unif(0,1) and external X ~ Unif(2,3), with P(S=1|X)=0.5 on (0,1) and P(S=1|X)=1 on (2,3). Then (C1) and (C3) hold (e.g., Y = X + \\epsilon), but the external covariate mean is 2.5, which lies outside the convex hull of the internal support [0,1]; no nonnegative weights can satisfy n^{-1}\\sum_i w_i X_i = 2.5. Hence Proposition 1's existence claim is false under the stated conditions, and the 'C3 only' side of double robustness requires an additional support-overlap condition, or at minimum a feasibility condition on the balancing constraints, before Theorem 1 can be regarded as correct.","section":"Theorem 1 and Conditions (C1)-(C4)"},{"comment":"The proposed estimator of the asymptotic variance is explicitly acknowledged to be inconsistent due to heterogeneity between the internal and external populations, and the bootstrap algorithm in Appendix B estimates the variability of the external summaries from bootstrap resamples of the internal data. Since the internal and external covariate distributions may be arbitrarily different under (C1), this step has no clear theoretical justification. Because the coverage claims and the practical variance estimates in Section 6 depend on this estimator, the manuscript should either provide regularity conditions under which the bootstrap is valid, or clearly label the variance estimator as heuristic and temper the inferential claims accordingly.","section":"Section 4, paragraph after Theorem 2; Appendix B"}],"minor_comments":[{"comment":"The sentence 'when the regression model is misspecified ... the choice of G1 does not affect consistency, so any G1 can be used without compromising first-order validity' is misleading: if (C3) fails, consistency relies on (C3)', so a misspecified G1 generally breaks consistency unless (C3) actually holds. The subsequent discussion of selection among G1 candidates implicitly assumes the correct density-ratio model is in the candidate set, which is fine, but the earlier sentence should be qualified.","section":"Section 3.4"},{"comment":"In Proposition 2, the assumption that G2 is 'strictly convex and uniquely minimized at 1' should also require that the candidate set is not empty and that the minimizer j0 is unique almost surely; the statement is otherwise clear.","section":"Equation (9) and Proposition 2"},{"comment":"In Theorem 3, the phrase 'both (C3) and (C3)' hold' is correct, but the proof of Theorem 3 in Appendix A silently uses the density-ratio representation E[\\rho_1(\\lambda_1^{*\\top}\\tilde X)\\tilde X] = \\tilde\\mu_{x|ex}; this should be stated explicitly before the block-inverse calculation.","section":"Section 5.1"},{"comment":"The notation M^{-1} in the block-inverse identity is correct, but the definition of \\Sigma_{ex} is introduced only implicitly; adding a sentence defining \\Sigma_{ex} = E_{ex}[(X-\\mu_{ex})(X-\\mu_{ex})^\\top] would improve readability.","section":"Appendix A, proof of Theorem 4"},{"comment":"Some display equations, such as (18), use the same norm symbol for vectors and matrices without clarification; this is not a substantive issue but should be cleaned up.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The support-overlap flaw is substantial but likely fixable by adding a feasibility/support condition and revising the proof of Proposition 1. The rest of the manuscript is well developed, with careful derivations, useful diagnostics, and a reproducible software package. I do not see a novelty disclosure concern; the contribution is clearly distinguished from the cited MAIC and empirical-likelihood literature. The paper fits the scope of a statistical methodology journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The two-stage entropy-balancing estimator is a real extension of entropy balancing to summary-data integration, and the efficiency diagnostics based on Mahalanobis distance and chi-square divergence are practical and easy to compute. But Theorem 1's \"C3 only\" arm is not established: without a support-overlap condition, the first-step balancing problem can be infeasible, and Proposition 1's proof is invalid as written.\n\nWhat is actually new: reweighting the internal data to external covariate moments, then applying a second entropy-balancing step using predicted outcomes, with data-adaptive selection among entropy families. The chi-square/Mahalanobis criteria for when borrowing beats the sample mean are clean and estimable. The R package daisy is a useful deliverable, and the simulations and the defibrillation application are appropriate checks. The asymptotic variance calculation is careful and the paper mostly engages the literature fairly.\n\nNow the soft spots, in proportion. The support-overlap flaw is load-bearing. Conditions (C1) and (C3) do not imply that the external covariate mean lies in the support or convex hull of the internal covariate distribution. If internal X is uniform on (0,1), external X is uniform on (2,3), and Y|X is linear, then all stated assumptions hold but no nonnegative weights can satisfy the first-step balancing equation. Lemma A.1 cannot rescue the proof because the second-step variable has a constant first coordinate, so the boxes with b_1 = ±1 have probability zero. This is not a cosmetic gap; it removes the C3-only side of the double robustness. Adding an explicit support-overlap condition, or proving Proposition 1 under C3' with support containment, should repair the claim, but the paper needs to do that work and restate Theorem 1 accordingly.\n\nThe variance estimator issue is real but secondary: Section 4 openly says the estimator is not consistent when the external sample is not large, and the bootstrap fallback is heuristic. That is addressable. Also, the abstract promises a bootstrap-based MSE selection criterion, but the body only formally proves consistency for entropy-family selection, not the bootstrap MSE criterion. That is an overstatement.\n\nThis paper is for statisticians working on data integration in clinical trials, surveys, and observational studies. It should go to peer review, and a good referee should require fixing Proposition 1 and revising Theorem 1's scope. I would not cite the double-robustness theorem as it stands, but I would use the estimator in practice only when the support condition is credible.","headline":"Two-stage entropy balancing is a genuinely useful estimator, but the paper's central double-robustness claim needs a support-overlap condition that is missing.","tokens_in":27836,"tokens_out":3966,"would_cite":false,"duration_ms":44702,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62G05","62D05"],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-stage entropy-balancing estimator consistently recovers the internal outcome mean from external summaries, even biased ones, under double-robustness conditions.","keywords":["entropy balancing","external summary statistics","outcome mean estimation","double robustness","transportability","covariate shift","data integration","mean squared error"],"falsifier":"Simulate internal data with outcome model $Y = X^\\top \\beta + U$, where $U$ is an unmeasured effect modifier whose distribution differs across the internal and external sources, and supply external summary means drawn from the shifted external distribution; if the estimator's bias does not vanish as the internal sample size grows while the balancing constraints are satisfied exactly, the claim that only measured covariates need transportability would be falsified.","tokens_in":26789,"feed_emoji":"📊","tokens_out":6420,"duration_ms":62097,"temperature":0.7,"pith_summary":"This paper argues that the mean outcome of an internal target population can be estimated more precisely by borrowing only a few external summary statistics—the sample means of the covariates and the outcome—even when the external sample is biased relative to the target. It proposes a two-stage generalized entropy-balancing estimator: first reweight internal individuals to match external covariate means, then reweight again using a working outcome model to absorb external outcome information. The headline theoretical result is double robustness: the estimator remains consistent if either the outcome-regression model or the density-ratio model implied by the first reweighting is correctly specified. The paper also provides a data-adaptive rule for choosing among candidate entropy functions, simple estimable efficiency criteria, and a bootstrap-based mean-squared-error criterion for deciding whether to borrow at all. Simulations and a cardiac-arrest registry application support the claim that the estimator often reduces mean squared error while controlling bias under distributional shift.","feed_headline":"External summary data can sharpen estimates without inheriting bias","feed_subtitle":"Doubly robust to misspecification, it reverts to the internal mean when external bias is too large.","key_machinery":"The engine is a pair of entropy-balancing steps linked through the Fenchel dual of a convex entropy function $G$. In the first step, weights $w_i^{(1)}=\\rho_1(\\lambda_1^\\top \\tilde{X}_i)$ are chosen to satisfy mean constraints $n^{-1}\\sum_i w_i^{(1)} X_i = \\hat{\\mu}_{x|\\mathrm{ex}}$, calibrating internal covariates to external moments. In the second step, a working outcome model $\\eta(X;\\beta)$ is combined with the first weights into $H_i = w_i^{(1)} \\eta_i|_{\\mathrm{in}}$, and weights $w_i^{(2)}=\\rho_2(\\lambda_2^\\top \\tilde{H}_i)$ are chosen to match the external outcome mean; the final estimator is the weighted mean of $Y$. Because the map from entropy function to weight model is the inverse derivative $\\rho = (G')^{-1}$, choosing $G_1$ is equivalent to choosing a density-ratio model, and the Lambert-W tempered family provides a regularized entropy whose inverse link uses the principal Lambert $W$ function. The selection rule picks the $G_1$ candidate whose second-step weights are closest to uniform, and the paper proves this selector consistently recovers a correctly specified density ratio when one is in the candidate set.","core_discovery":"The central claim is that internal-versus-external distribution shift can be exploited rather than feared: under transportability (the conditional distribution of Y given X is the same across sources), the two-stage entropy-balancing estimator $\\hat{\\theta}_{\\mathrm{EBW}}$ converges to the true internal mean $\\theta^*$ even when the external population's marginal covariate distribution and outcome mean are shifted. The first balancing step learns a density ratio between external and internal covariates from external covariate means; the second step uses the fitted outcome model and the external outcome mean to produce weights that are asymptotically uniform, so the final weighted mean is consistent. Double robustness means consistency survives if the outcome model is wrong as long as the density-ratio model is right, and vice versa. Under a linear homoscedastic model the efficiency gain is characterized in closed form: the estimator beats the internal sample mean exactly when the squared Mahalanobis distance between covariate means is at most 1, and in the weighted version exactly when Pearson's chi-squared divergence between the two covariate distributions is at most 1.","pith_inferences":["A natural but undeveloped extension is to apply the same two-stage calibration to any estimand defined by a moment condition, not just the mean, whenever the external summaries are moments of the same estimating function.","The paper's remark on aggregating multiple external sources together with its borrowing rule suggests a forward-selection algorithm over subsets of sources, where each candidate subset is evaluated by the bootstrap mean-squared-error criterion; the paper does not implement this.","The chi-squared criterion $D_2 \\le 1$ is equivalent to requiring the density-ratio weights to have second moment at most $2$; external covariate distributions with heavy tails will quickly violate this, so the method is most useful for modest distributional shifts.","Since Condition (C1) is not testable from internal data and external summary means alone, a sensitivity analysis that varies which covariates enter the balancing constraints would strengthen practical deployment; the paper does not provide such an analysis."],"forward_implications":["Practitioners can integrate external control arms from historical trials without needing to know whether their outcome regression model is correct, provided the density-ratio model implied by the first balancing step is credible.","When the external sample is large relative to the internal one, the asymptotic variance contribution of external summary noise becomes negligible, simplifying variance estimation; a bootstrap algorithm is provided for the situation where it does not.","The estimable thresholds—squared Mahalanobis distance at most $1$ for the plain estimator and Pearson chi-squared divergence at most $1$ for the weighted version—give an operational diagnostic for whether integration is guaranteed to improve on the internal sample mean.","The bootstrap mean-squared-error comparison is selection consistent under fixed alternatives, so the procedure asymptotically borrows when borrowing helps and reverts to the internal sample mean when external bias is nonvanishing.","The same two-stage construction extends to outcome regression coefficients from an external generalized linear model, though the efficiency condition in that case no longer reduces to a simple distance. "],"supporting_citations":[{"why":"Supplies entropy balancing as the reweighting method for matching covariate moments, which the first balancing step directly builds on.","marker":"(Hainmueller 2012)"},{"why":"Provides the double-robustness result for entropy balancing and the existence lemma used to prove the balancing weights are well defined.","marker":"(Zhao and Percival 2017)"},{"why":"Establishes the covariate-shift setting and density-ratio-weighted regression, motivating Condition (C3)' and the weighted outcome regression.","marker":"(Shimodaira 2000)"},{"why":"Frames integration of summary information through estimating equations and empirical likelihood, the broad framework the two-stage construction extends.","marker":"(Qin and Lawless 1994)"},{"why":"Demonstrates entropy balancing for transporting results across populations, the transportability context the method operates in.","marker":"(Josey et al. 2021)"},{"why":"Articulates the transportability and missing-at-random condition (C1) for combining probability and non-probability samples.","marker":"(Kim et al. 2021)"},{"why":"Provides matching-adjusted indirect comparison, the prior approach that reweights internal individuals to match external summary characteristics.","marker":"(Signorovitch et al. 2010)"},{"why":"Represents the constrained-maximum-likelihood alternative for model calibration using external summary information, a comparison point for the proposed approach.","marker":"(Chatterjee et al. 2016)"}],"fun_headline_variants":["Doubly robust external-summary integration for outcome means","Exploit external summaries without inheriting their bias","A data-adaptive rule for borrowing external summary data","Borrow external summary data only when bias is small","Doubly robust when to borrow external summary data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is transportability: the conditional distribution of the outcome given the measured covariates is the same in the internal and external populations; if an unmeasured effect modifier differs across sources, the external summaries cannot be recalibrated and the estimator is biased no matter how well the weights balance.","fun_headline_variants_meta":{"raw":{"variants":["Doubly robust external-summary integration for outcome means","Exploit external summaries without inheriting their bias","A data-adaptive rule for borrowing external summary data","Borrow external summary data only when bias is small","Doubly robust when to borrow external summary data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000673,"raw_usage":{"total_tokens":3102,"prompt_tokens":1019,"completion_tokens":2083,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":2009}},"tokens_in":635,"tokens_out":2083,"duration_ms":17050,"temperature":1.0,"reasoning_tokens":2009,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:05:07.574995+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate internal data with outcome model $Y = X^\\top \\beta + U$, where $U$ is an unmeasured effect modifier whose distribution differs across the internal and external sources, and supply external summary means drawn from the shifted external distribution; if the estimator's bias does not vanish as the internal sample size grows while the balancing constraints are satisfied exactly, the claim that only measured covariates need transportability would be falsified.","supporting_citations":[],"review_version":1}