REVIEW 3 major objections 5 minor 31 references
Data-Adaptive Integration with External Summary Data for Outcome Mean Estimation
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A two-stage entropy-balancing estimator consistently recovers the internal outcome mean from external summaries, even biased ones, under double-robustness conditions.
desk verdict Two-stage entropy balancing is a genuinely useful estimator, but the paper's central double-robustness claim needs a support-overlap condition that is missing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is a pair of entropy-balancing steps linked through the Fenchel dual of a convex entropy function $G$. In the first step, weights $w_i^{(1)}=\rho_1(\lambda_1^\top \tilde{X}_i)$ are chosen to satisfy mean constraints $n^{-1}\sum_i w_i^{(1)} X_i = \hat{\mu}_{x|\mathrm{ex}}$, calibrating internal covariates to external moments. In the second step, a working outcome model $\eta(X;\beta)$ is combined with the first weights into $H_i = w_i^{(1)} \eta_i|_{\mathrm{in}}$, and weights $w_i^{(2)}=\rho_2(\lambda_2^\top \tilde{H}_i)$ are chosen to match the external outcome mean; the final estimator is the weighted mean of $Y$. Because the map from entropy function to weight model is the inverse derivative $\rho = (G')^{-1}$, choosing $G_1$ is equivalent to choosing a density-ratio model, and the Lambert-W tempered family provides a regularized entropy whose inverse link uses the principal Lambert $W$ function. The selection rule picks the $G_1$ candidate whose second-step weights are closest to uniform, and the paper proves this selector consistently recovers a correctly specified density ratio when one is in the candidate set.
What would settle it
Simulate internal data with outcome model $Y = X^\top \beta + U$, where $U$ is an unmeasured effect modifier whose distribution differs across the internal and external sources, and supply external summary means drawn from the shifted external distribution; if the estimator's bias does not vanish as the internal sample size grows while the balancing constraints are satisfied exactly, the claim that only measured covariates need transportability would be falsified.
Extended reading notes
Core claim
The central claim is that internal-versus-external distribution shift can be exploited rather than feared: under transportability (the conditional distribution of Y given X is the same across sources), the two-stage entropy-balancing estimator $\hat{\theta}_{\mathrm{EBW}}$ converges to the true internal mean $\theta^*$ even when the external population's marginal covariate distribution and outcome mean are shifted. The first balancing step learns a density ratio between external and internal covariates from external covariate means; the second step uses the fitted outcome model and the external outcome mean to produce weights that are asymptotically uniform, so the final weighted mean is consistent. Double robustness means consistency survives if the outcome model is wrong as long as the density-ratio model is right, and vice versa. Under a linear homoscedastic model the efficiency gain is characterized in closed form: the estimator beats the internal sample mean exactly when the squared Mahalanobis distance between covariate means is at most 1, and in the weighted version exactly when Pearson's chi-squared divergence between the two covariate distributions is at most 1.
Load-bearing premise
The load-bearing premise is transportability: the conditional distribution of the outcome given the measured covariates is the same in the internal and external populations; if an unmeasured effect modifier differs across sources, the external summaries cannot be recalibrated and the estimator is biased no matter how well the weights balance.
Editorial extensions
If this is right
- Practitioners can integrate external control arms from historical trials without needing to know whether their outcome regression model is correct, provided the density-ratio model implied by the first balancing step is credible.
- When the external sample is large relative to the internal one, the asymptotic variance contribution of external summary noise becomes negligible, simplifying variance estimation; a bootstrap algorithm is provided for the situation where it does not.
- The estimable thresholds—squared Mahalanobis distance at most $1$ for the plain estimator and Pearson chi-squared divergence at most $1$ for the weighted version—give an operational diagnostic for whether integration is guaranteed to improve on the internal sample mean.
- The bootstrap mean-squared-error comparison is selection consistent under fixed alternatives, so the procedure asymptotically borrows when borrowing helps and reverts to the internal sample mean when external bias is nonvanishing.
- The same two-stage construction extends to outcome regression coefficients from an external generalized linear model, though the efficiency condition in that case no longer reduces to a simple distance.
Reading between the lines
- A natural but undeveloped extension is to apply the same two-stage calibration to any estimand defined by a moment condition, not just the mean, whenever the external summaries are moments of the same estimating function.
- The paper's remark on aggregating multiple external sources together with its borrowing rule suggests a forward-selection algorithm over subsets of sources, where each candidate subset is evaluated by the bootstrap mean-squared-error criterion; the paper does not implement this.
- The chi-squared criterion $D_2 \le 1$ is equivalent to requiring the density-ratio weights to have second moment at most $2$; external covariate distributions with heavy tails will quickly violate this, so the method is most useful for modest distributional shifts.
- Since Condition (C1) is not testable from internal data and external summary means alone, a sensitivity analysis that varies which covariates enter the balancing constraints would strengthen practical deployment; the paper does not provide such an analysis.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a two-stage generalized entropy-balancing estimator for the internal-population outcome mean that borrows external summary statistics (means of X and Y) while explicitly allowing the external sample to be biased relative to the target. The first round of balancing calibrates internal covariates to the external covariate mean; the second round calibrates an outcome-model-based quantity to the external outcome mean, and the final estimator is the second-stage weighted mean of Y. The authors establish consistency under either a correctly specified linear outcome model (C3) or a correctly specified density-ratio model (C3)', prove asymptotic normality with an explicit variance formula, propose a data-adaptive selection rule among candidate entropy functions, and derive estimable efficiency criteria based on the Mahalanobis distance and Pearson chi-squared divergence. Simulations and a Japan Utstein registry application illustrate the method, and an R package daisy is provided.
Significance. If the main theorem holds, this is a practically valuable contribution: it uses only moment summaries from the external source, permits covariate shift, and provides double robustness plus easy-to-compute applicability diagnostics. The paper also ships reproducible software, gives detailed proofs of the asymptotic variance and efficiency criteria, and includes a substantive real-data analysis. These are genuine strengths that make the paper potentially publishable. However, the double-robustness claim is not established under the stated conditions because the existence proof of the balancing weights is flawed and the feasibility of the balancing equations is not guaranteed by Conditions (C1)-(C4). This gap affects the central consistency theorem, so the result needs substantial revision before the manuscript can be accepted.
major comments (3)
- [Appendix A, proof of Proposition 1] The proof of Proposition 1 applies Lemma A.1 to the augmented variables \tilde H^*_j = (1, H^*_j)^\top with p=2. This application is invalid: the first coordinate of \tilde H^*_j is identically 1, while Lemma A.1 requires the sample to contain a point in each of 3^p boxes centered at \mu^*_1 + 3\epsilon b/2 with b_1 \in \{-1,0,1\}. For b_1 = \pm 1 and sufficiently small \epsilon, the first-coordinate interval is bounded away from 1, so the required event has probability zero. Thus the proof of Proposition 1 fails as written, and the existence of \hat\lambda_2 needs a different argument that handles the deterministic intercept coordinate.
- [Theorem 1 and Conditions (C1)-(C4)] The consistency claim under the C3-only arm is not established because Conditions (C1), (C2), and (C3) do not imply that the first-step balancing problem (1) is feasible. For example, let internal X ~ Unif(0,1) and external X ~ Unif(2,3), with P(S=1|X)=0.5 on (0,1) and P(S=1|X)=1 on (2,3). Then (C1) and (C3) hold (e.g., Y = X + \epsilon), but the external covariate mean is 2.5, which lies outside the convex hull of the internal support [0,1]; no nonnegative weights can satisfy n^{-1}\sum_i w_i X_i = 2.5. Hence Proposition 1's existence claim is false under the stated conditions, and the 'C3 only' side of double robustness requires an additional support-overlap condition, or at minimum a feasibility condition on the balancing constraints, before Theorem 1 can be regarded as correct.
- [Section 4, paragraph after Theorem 2; Appendix B] The proposed estimator of the asymptotic variance is explicitly acknowledged to be inconsistent due to heterogeneity between the internal and external populations, and the bootstrap algorithm in Appendix B estimates the variability of the external summaries from bootstrap resamples of the internal data. Since the internal and external covariate distributions may be arbitrarily different under (C1), this step has no clear theoretical justification. Because the coverage claims and the practical variance estimates in Section 6 depend on this estimator, the manuscript should either provide regularity conditions under which the bootstrap is valid, or clearly label the variance estimator as heuristic and temper the inferential claims accordingly.
minor comments (5)
- [Section 3.4] The sentence 'when the regression model is misspecified ... the choice of G1 does not affect consistency, so any G1 can be used without compromising first-order validity' is misleading: if (C3) fails, consistency relies on (C3)', so a misspecified G1 generally breaks consistency unless (C3) actually holds. The subsequent discussion of selection among G1 candidates implicitly assumes the correct density-ratio model is in the candidate set, which is fine, but the earlier sentence should be qualified.
- [Equation (9) and Proposition 2] In Proposition 2, the assumption that G2 is 'strictly convex and uniquely minimized at 1' should also require that the candidate set is not empty and that the minimizer j0 is unique almost surely; the statement is otherwise clear.
- [Section 5.1] In Theorem 3, the phrase 'both (C3) and (C3)' hold' is correct, but the proof of Theorem 3 in Appendix A silently uses the density-ratio representation E[\rho_1(\lambda_1^{*\top}\tilde X)\tilde X] = \tilde\mu_{x|ex}; this should be stated explicitly before the block-inverse calculation.
- [Appendix A, proof of Theorem 4] The notation M^{-1} in the block-inverse identity is correct, but the definition of \Sigma_{ex} is introduced only implicitly; adding a sentence defining \Sigma_{ex} = E_{ex}[(X-\mu_{ex})(X-\mu_{ex})^\top] would improve readability.
- [General] Some display equations, such as (18), use the same norm symbol for vectors and matrices without clarification; this is not a substantive issue but should be cleaned up.
Circularity Check
No circularity: the estimator, selection rule, and efficiency criteria are derived from stated assumptions rather than fitted to the target result.
full rationale
The paper's claimed derivation chain is self-contained and does not reduce any prediction to its inputs by construction. The target parameter is the internal-population outcome mean theta* = E(Y), and the estimator theta_hat_EBW is a weighted internal sample mean. Consistency (Theorem 1) is shown by solving the estimating equations (6), (7), and (10), then verifying that the limiting weight rho2(lambda2*^T \tilde H*) equals 1 almost surely when lambda2* = (rho2^{-1}(1),0). This is not circular: it depends on (C1), (C3) or (C3)', the first-step balance equation E[rho1(lambda1*^T \tilde X)\tilde X] = tilde mu*_{x|ex}, and the algebraic identity E(H*) = beta_ex*^T mu*_{x|ex}. The external outcome summary enters through the constraint eta_ex = mu_{y|ex}, but theta* is never inserted into the constraints or the estimating equations. The data-adaptive selection rule in Section 3.4 and Proposition 2 is a model-selection consistency result: candidates G1[j] are fixed before estimation, the selector minimizes a strictly convex criterion, and consistency requires the candidate set to contain the true density-ratio model. No fitted parameter is renamed as a prediction. The efficiency criteria in Theorems 3, 4, and C.1 are algebraic consequences of the influence-function variance formulas under stated linearity, homoscedasticity, and density-ratio assumptions; the Mahalanobis and chi-square diagnostics are fully estimable and are not used to define the estimand. There are no load-bearing self-citations: the only imported technical result, Lemma A.1, is attributed to Zhao and Percival (2017), an external source, and the paper's main theorems are proved from the stated regularity conditions rather than from a uniqueness theorem of the authors. Two mathematical-validity concerns are worth flagging separately, but they are not circularity: Proposition 1's proof applies Lemma A.1 to \tilde H*_j = (1, H*_j)^top, whose first coordinate is constant and therefore cannot fall in the 3^2 boxes with first-coordinate intervals bounded away from 1, so the existence proof appears incomplete; and the 'C3 only' arm of double robustness may require support overlap or a correctly specified outcome model for the equality E(H*) = eta_ex to hold. These are correctness risks, and Section 8's own limitations statement acknowledges that the method does not guarantee uniform improvement.
Assumptions & free parameters
free parameters (4)
- Lambert-W entropy tempering parameter a =
candidate set {0.1, 0.5, 1.0, 1.5}, data-adaptively selected
- Quadratic log-sum entropy parameter b =
candidate set {0.1, 0.5, 1.0, 1.5}, data-adaptively selected
- Tempered softplus entropy parameter c =
candidate set {0.1, 0.5, 1.0, 1.5}, data-adaptively selected
- Bootstrap resample sizes B1, B2 =
200
assumptions (6)
- domain assumption C1: Transportability, P(S=1|Z)=P(S=1|X)>0
- domain assumption C2: External means of X and Y are available
- domain assumption C3: Linear outcome regression E(Y|X=x)=beta^T x_tilde
- domain assumption C3': Density ratio has single-index form fex/fin=rho1(lambda^T x)
- standard math C4-C5: Standard regularity conditions for ULLN and CLT
- domain assumption C6: Homoscedasticity var(Y|X)=sigma^2
Cite this review
Pith. "Pith review of Data-Adaptive Integration with External Summary Data for Outcome Mean Estimation." pith.science (2026). https://pith.science/paper/APRBTWGH
@misc{pith2026250611482,
author = {Pith},
title = {Pith review of: Data-Adaptive Integration with External Summary Data for Outcome Mean Estimation},
year = {2026},
howpublished = {\url{https://pith.science/paper/APRBTWGH}},
note = {Machine review of arXiv:2506.11482}
}
read the original abstract
Combining an internal individual-level study with readily available external summary statistics promises major efficiency gains at minimal additional cost, yet heterogeneity between sources can bias estimates for the internal target population. We develop a generalized entropy-balancing integration strategy that calibrates the internal individual-level sample to externally reported moments while retaining the internal population as the target, explicitly permitting a biased external sample. The weighted-regression version of our estimator is doubly robust: it remains consistent when either the outcome-regression model or the entropy-balancing model is correctly specified. When multiple balancing specifications are plausible, we introduce a data-adaptive entropy-family selection rule. For the final borrowing decision, we propose a bootstrap-based criterion comparing stabilized mean squared error (MSE) estimates for the selected entropy-balancing estimator and the internal sample mean. This criterion is selection consistent under fixed alternatives and reverts to the internal estimator when a nonvanishing bias is detected. Separately, under a linear homoscedastic benchmark, the asymptotic efficiency criteria admit geometric interpretations through the Mahalanobis distance and Pearson chi-squared divergence. The entropy-balancing estimators and numerical-experiment routines are implemented in the R package daisy. Simulations show stable MSE reductions for the weighted-regression estimator across calibrated distributional shifts and the predicted reversion toward the internal estimator under fixed simultaneous misspecification as the sample size increases. An application to nationwide public-access defibrillation records in Japan illustrates the resulting MSE-based borrowing decision.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
M. Borenstein, L. V. Hedges, J. P. T. Higgins, and H. R. Rothstein. Introduction to Meta ‐ Analysis . Wiley, 1 edition, Mar. 2009
work page 2009
-
[2]
N. Chatterjee, Y.-H. Chen, P. Maas, and R. J. Carroll. Constrained Maximum Likelihood Estimation for Model Calibration Using Summary - Level Information From External Big Data Sources . Journal of the American Statistical Association , 111(513):107--117, Jan. 2016
work page 2016
-
[3]
C. Chen, P. Han, S. Chen, M. Shardell, and J. Qin. Integrating external summary information in the presence of prior probability shift: an application to assessing essential hypertension. Biometrics , 80(3):ujae090, July 2024
work page 2024
-
[4]
J. Chu, W. Lu, and S. Yang. Targeted optimal treatment regime learning using summary statistics. Biometrika , 110(4):913--931, Dec. 2023
work page 2023
-
[5]
S. R. Cole and E. A. Stuart. Generalizing Evidence From Randomized Clinical Trials to Target Populations . American Journal of Epidemiology , 172(1):107--115, July 2010
work page 2010
-
[6]
J. Hainmueller. Entropy Balancing for Causal Effects : A Multivariate Reweighting Method to Produce Balanced Samples in Observational Studies . Polit. anal. , 20(1):25--46, 2012
work page 2012
- [7]
-
[8]
I. Jacobs, V. Nadkarni, J. Bahr, R. A. Berg, J. E. Billi, L. Bossaert, P. Cassan, A. Coovadia, K. D’Este, J. Finn, H. Halperin, A. Handley, J. Herlitz, R. Hickey, A. Idris, W. Kloeck, G. L. Larkin, M. E. Mancini, P. Mason, G. Mears, K. Monsieurs, W. Montgomery, P. Morley, G. Nichol, J. Nolan, K. Okada, J. Perlman, M. Shuster, P. A. Steen, F. Sterz, J. Tib...
work page 2004
Show all 31 references
-
[9]
K. P. Josey, S. A. Berkowitz, D. Ghosh, and S. Raghavan. Transporting experimental results with entropy balancing. Statistics in Medicine , 40(19):4310--4326, Aug. 2021
2021
-
[10]
J. K. Kim, S. Park, Y. Chen, and C. Wu. Combining Non - Probability and Probability Survey Samples Through Mass Imputation . Journal of the Royal Statistical Society Series A: Statistics in Society , 184(3):941--963, July 2021
2021
-
[11]
Kitamura, T
T. Kitamura, T. Iwami, T. Kawamura, K. Nagao, H. Tanaka, and A. Hiraide. Nationwide Public - Access Defibrillation in Japan . N Engl J Med , 362(11):994--1004, Mar. 2010
2010
-
[12]
Kundu, R
P. Kundu, R. Tang, and N. Chatterjee. Generalized meta-analysis for multiple regression models across studies with disparate covariate information. Biometrika , 106(3):567--585, Sept. 2019
2019
-
[13]
Y. Kwon, J. K. Kim, and Y. Qiu. Debiased calibration estimation using generalized entropy in survey sampling, Sept. 2024. arXiv:2404.01076 [stat]
2024 arXiv
-
[14]
S. Y. Lee, B. Lei, and B. Mallick. Estimation of COVID -19 spread curves integrating global data and borrowing information. PLoS ONE , 15(7):e0236860, July 2020
2020
-
[15]
Li, , and Y
X. Li, , and Y. Song. Target Population Statistical Inference With Data Integration Across Multiple Sources — An Approach to Mitigate Information Shortage in Rare Disease Clinical Trials . Statistics in Biopharmaceutical Research , 12(3):322--333, July 2020
2020
-
[16]
R. J. A. Little and D. B. Rubin. Statistical Analysis with Missing Data . John Wiley & Sons, Apr. 2019
2019
-
[17]
Lumley, P
T. Lumley, P. A. Shaw, and J. Y. Dai. Connections between Survey Calibration Estimators and Semiparametric Models for Incomplete Data . International Statistical Review , 79(2):200--220, 2011
2011
-
[18]
S. J. Pocock. The combination of randomized and historical controls in clinical trials. Journal of Chronic Diseases , 29(3):175--188, Mar. 1976
1976
-
[19]
J. Qin. Combining Parametric and Empirical Likelihoods . Biometrika , 87(2):484--490, 2000
2000
-
[20]
Qin and J
J. Qin and J. Lawless. Empirical Likelihood and General Estimating Equations . Ann. Statist. , 22(1), Mar. 1994
1994
-
[21]
J. Qin, H. Zhang, P. Li, D. Albanes, and K. Yu. Using covariate-specific disease prevalence information to increase the power of case-control studies. Biometrika , 102(1):169--180, 2015
2015
-
[22]
R. D. Riley, P. C. Lambert, and G. Abo-Zaid. Meta-analysis of individual participant data: rationale, conduct, and reporting. BMJ , 340(feb05 1):c221--c221, Aug. 2010
2010
-
[23]
R. D. Riley, L. A. Stewart, and J. F. Tierney. Individual Participant Data Meta - Analysis for Healthcare Research . John Wiley & Sons, Ltd, 2021
2021
-
[24]
D. B. Rubin. Inference and missing data. Biometrika , 63(3):581--592, 1976
1976
-
[25]
J. L. Schafer and J. W. Graham. Missing data: Our view of the state of the art. Psychological Methods , 7(2):147--177, 2002
2002
-
[26]
Shimodaira
H. Shimodaira. Improving predictive inference under covariate shift by weighting the log-likelihood function. Journal of Statistical Planning and Inference , 90(2):227--244, Oct. 2000
2000
-
[27]
J. E. Signorovitch, E. Q. Wu, A. P. Yu, C. M. Gerrits, E. Kantor, Y. Bao, S. R. Gupta, and P. M. Mulani. Comparative Effectiveness Without Head -to- Head Trials . PharmacoEconomics , 28(10):935--945, Oct. 2010
2010
-
[28]
Stewart and M
L. Stewart and M. Parmar. Meta-analysis of the literature or of individual patient data: is there a difference? The Lancet , 341(8842):418--422, Feb. 1993
1993
-
[29]
Viele, S
K. Viele, S. Berry, B. Neuenschwander, B. Amzal, F. Chen, N. Enas, B. Hobbs, J. G. Ibrahim, N. Kinnersley, S. Lindborg, S. Micallef, S. Roychoudhury, and L. Thompson. Use of historical control data for assessing treatment effects in clinical trials. Pharmaceutical statistics ,...
2014
-
[30]
Yang and P
S. Yang and P. Ding. Combining Multiple Observational Data Sources to Estimate Causal Effects . Journal of the American Statistical Association , 115(531):1540--1554, 2020
2020
-
[31]
Zhao and D
Q. Zhao and D. Percival. Entropy Balancing is Doubly Robust . Journal of Causal Inference , 5(1), Mar. 2017
2017
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.