REVIEW 2 major objections 2 minor 46 references
Doubly Robust Quadratic Inference Functions for Causal Inference in Cluster Randomized Trials
T0 review · 2 major / 2 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read The DR-QIF estimator is consistent for the average treatment effect in cluster randomized trials if either the propensity score model or the outcome regression model is correctly specified.
desk verdict DR-QIF matches DR-GEE exactly in cross-section and adds only a small efficiency edge in longitudinal CRTs under correlation misspecification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Doubly robust pseudo-outcomes constructed from fitted propensity and outcome models and inserted into the QIF extended score equations
What would settle it
A simulation study in which the propensity score model is correctly specified yet the DR-QIF estimator fails to converge to the true average treatment effect would falsify the consistency claim.
Extended reading notes
Core claim
By forming doubly robust pseudo-outcomes from a propensity score model and an outcome regression model and then substituting those pseudo-outcomes into the quadratic inference function estimating equations, the DR-QIF estimator is consistent for the marginal average treatment effect whenever either working model is correct. The paper establishes that this estimator is asymptotically more efficient than its doubly robust GEE counterpart under misspecification of the working correlation matrix and supplies an analytic characterization of the efficiency difference. The two estimators coincide algebraically in cross-sectional cluster randomized trials but diverge in longitudinal settings with st
Load-bearing premise
The doubly robust pseudo-outcomes retain their marginal causal interpretation and double-robustness property when substituted into the QIF estimating equations under cluster sampling.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a doubly robust quadratic inference function (DR-QIF) estimator for the average treatment effect in cluster randomized trials (CRTs) that may involve confounding. It combines doubly robust pseudo-outcomes (from propensity score and outcome regression models) with the QIF extended score equations. The central claims are that DR-QIF is consistent for the ATE if either the propensity or outcome model is correct (but not necessarily both), that it is asymptotically more efficient than doubly robust GEE (DR-GEE) when the working correlation is misspecified, and that this efficiency gain can be characterized analytically; gains are reported to reach 3.5% in longitudinal settings with N=120 and T=8. Finite-sample behavior is assessed via Monte Carlo simulation and the method is illustrated on the WASH Benefits Kenya CRT data. For cross-sectional CRTs the estimators coincide.
Significance. If the double-robustness property transfers to the QIF estimating function under cluster sampling, the work supplies a more efficient marginal estimator than DR-GEE for longitudinal CRTs in which the working correlation is typically misspecified. The analytical derivation of the efficiency gain (rather than purely numerical comparison) and the explicit simulation checks are strengths that would strengthen the contribution if the consistency argument is placed on firmer footing.
major comments (2)
- [§3] §3 (consistency argument): the claim that DR pseudo-outcomes inserted into the QIF extended score vector preserve E[extended score] = 0 at the true marginal ATE whenever either the propensity or outcome model is correct does not follow immediately from the standard DR-GEE argument, because QIF forms a stacked vector from multiple basis matrices and minimizes a quadratic form; an explicit verification that the zero-mean property survives cluster-level treatment assignment and within-cluster dependence is required for the consistency result to be load-bearing.
- [§4] §4 (efficiency comparison): the analytical efficiency gain over DR-GEE is stated to hold when the working correlation is misspecified, but the derivation should explicitly display the asymptotic variance expressions for both estimators under the cluster sampling measure and confirm that the gain is not an artifact of the particular basis-matrix choice or the longitudinal design parameters used in the 3.5% calculation.
minor comments (2)
- [Simulations] The Monte Carlo section should report the precise cluster-size distribution, the exact form of the data-generating propensity and outcome models, and the rule used to exclude any simulated replicates (if any).
- [Methods] Notation for the extended score vector and the basis matrices should be restated once the DR pseudo-outcomes are substituted, to avoid ambiguity when readers compare the construction to standard QIF.
Simulated Author's Rebuttal
We thank the referee for the careful reading and constructive comments on the consistency and efficiency arguments. We address each major comment below and will revise the manuscript to place both results on firmer footing.
read point-by-point responses
-
Referee: [§3] §3 (consistency argument): the claim that DR pseudo-outcomes inserted into the QIF extended score vector preserve E[extended score] = 0 at the true marginal ATE whenever either the propensity or outcome model is correct does not follow immediately from the standard DR-GEE argument, because QIF forms a stacked vector from multiple basis matrices and minimizes a quadratic form; an explicit verification that the zero-mean property survives cluster-level treatment assignment and within-cluster dependence is required for the consistency result to be load-bearing.
Authors: We agree that an explicit verification is needed rather than relying on the DR-GEE argument by analogy. In the revision we will insert a dedicated lemma that directly computes the expectation of each component of the stacked extended score vector under cluster-level randomization. The argument will use the law of total expectation, conditioning first on the cluster-level treatment indicator and then on the within-cluster covariates, to show that the zero-mean property holds for every basis matrix whenever either the propensity-score or outcome-regression model is correct. This will be placed immediately after the definition of the DR-QIF estimator. revision: yes
-
Referee: [§4] §4 (efficiency comparison): the analytical efficiency gain over DR-GEE is stated to hold when the working correlation is misspecified, but the derivation should explicitly display the asymptotic variance expressions for both estimators under the cluster sampling measure and confirm that the gain is not an artifact of the particular basis-matrix choice or the longitudinal design parameters used in the 3.5% calculation.
Authors: We will revise §4 to display the full asymptotic variance formulas for both DR-QIF and DR-GEE under the cluster sampling measure, written in terms of the cluster-level influence functions and the sandwich form that accounts for within-cluster dependence. The comparison will be carried out at the level of the Godambe information matrices, showing that the difference is nonnegative whenever the working correlation differs from the true one, and that the sign of the difference does not depend on the specific choice of basis matrices or on the particular values of N and T used in the numerical illustration. The 3.5% figure will be retained only as an example; the general analytic result will be stated first. revision: yes
Circularity Check
No circularity; claims rest on imported DR properties and analytical derivations
full rationale
The paper's consistency claim for DR-QIF imports standard double-robustness from prior literature rather than deriving it from its own fitted quantities or self-citations. The efficiency comparison and asymptotic gain are characterized analytically from the QIF extended score and DR pseudo-outcome constructions. No step reduces by construction to its inputs, renames a known result, or relies on a load-bearing self-citation chain. The derivation is self-contained against external benchmarks for the marginal ATE under cluster sampling.
Assumptions & free parameters
assumptions (2)
- domain assumption Either the propensity score model or the outcome regression model is correctly specified
- domain assumption The marginal causal model remains valid after insertion of the pseudo-outcomes into the QIF score equations under cluster sampling
Cite this review
Pith. "Pith review of Doubly Robust Quadratic Inference Functions for Causal Inference in Cluster Randomized Trials." pith.science (2026). https://pith.science/paper/AI3LFHYU
@misc{pith2026260626630,
author = {Pith},
title = {Pith review of: Doubly Robust Quadratic Inference Functions for Causal Inference in Cluster Randomized Trials},
year = {2026},
howpublished = {\url{https://pith.science/paper/AI3LFHYU}},
note = {Machine review of arXiv:2606.26630}
}
read the original abstract
Quadratic inference functions (QIF) provide a robust and efficient alternative to generalized estimating equations (GEE) for marginal regression with correlated data, particularly in cluster randomized trials (CRTs). However, existing QIF methodology does not account for confounding due to covariate imbalance between treatment arms, a common concern in observational CRTs or CRTs with prognostic covariate adjustment. We propose a doubly robust QIF (DR-QIF) estimator that combines doubly robust pseudo-outcomes, constructed from propensity score and outcome regression models, with the QIF extended score equations. The DR-QIF estimator is consistent for the average treatment effect when either the propensity score model or the outcome regression model is correctly specified, but not necessarily both. We show that DR-QIF is more efficient than doubly robust GEE (DR-GEE) when the working correlation structure is misspecified, and we characterize the asymptotic efficiency gain analytically. For cross-sectional CRTs the two estimators are algebraically identical; efficiency gains emerge in longitudinal CRTs with strong temporal correlation, reaching 3.5% at N=120 and T=8 repeated measures. Finite-sample properties are evaluated via Monte Carlo simulation, and the method is illustrated using data from the WASH Benefits Kenya cluster randomized trial.
Reference graph
Works this paper leans on
-
[1]
and Robins, J
Bang, H. and Robins, J. M. , title =. Biometrics , year =
-
[2]
and Chetverikov, D
Chernozhukov, V. and Chetverikov, D. and Demirer, M. and Duflo, E. and Hansen, C. and Newey, W. and Robins, J. , title =. The Econometrics Journal , year =
-
[3]
Fay, M. P. and Graubard, B. I. , title =. Biometrics , year =
-
[4]
Hansen, L. P. , title =. Econometrica , year =
-
[5]
Hayes, R. J. and Moulton, L. H. , title =
-
[6]
and Carroll, R
Kauermann, G. and Carroll, R. J. , title =. Journal of the American Statistical Association , year =
-
[7]
and Zeger, S
Liang, K.-Y. and Zeger, S. L. , title =. Biometrika , year =
-
[8]
and Preisser, J
Lu, B. and Preisser, J. S. and Qaqish, B. F. and Suchindran, C. and Bangdiwala, S. I. and Wolfson, M. , title =. Biometrics , year =
Show all 46 references
-
[9]
and Groenwold, R
Luijken, K. and Groenwold, R. H. H. and Van Calster, B. and Steyerberg, E. W. and van Smeden, M. , title =. Statistics in Medicine , year =
-
[10]
Mancl, L. A. and DeRouen, T. A. , title =. Biometrics , year =
-
[11]
and Wang, Rui , title =
Rabideau, Dustin J. and Wang, Rui , title =. Statistics in Medicine , year =
-
[12]
Preisser, J. S. and Young, M. L. and Zaccaro, D. J. and Wolfson, M. , title =. Statistics in Medicine , year =
-
[13]
Preisser, J. S. and Lu, B. and Qaqish, B. F. , title =. Statistics in Medicine , year =
-
[14]
and Lindsay, B
Qu, A. and Lindsay, B. G. and Li, B. , title =. Biometrika , year =
-
[15]
Robins, J. M. and Rotnitzky, A. and Zhao, L. P. , title =. Journal of the American Statistical Association , year =
-
[16]
Rubin, D. B. , title =. Journal of Educational Psychology , year =
-
[17]
Seaman, S. R. and Copas, A. J. , title =. Statistics in Medicine , year =
-
[18]
and Li, F
Yu, H. and Li, F. and Turner, E. L. , title =. Contemporary Clinical Trials Communications , year =
-
[19]
and Tong, G
Yu, H. and Tong, G. and Li, F. , title =. Communications in Statistics -- Simulation and Computation , year =
-
[20]
, title =
Wang, N. , title =. Biometrika , year =
-
[21]
and Rahman, Mahbubur and Arnold, Benjamin F
Luby, Stephen P. and Rahman, Mahbubur and Arnold, Benjamin F. and Unicomb, Leanne and Ashraf, Sania and Winch, Peter J. and Stewart, Christine P. and Begum, Farjana and Hussain, Faruq and Yunus, Mohammad and Chakraborty, Joy and Clasen, Thomas F. and Leontsini, Elli and Naser,...
-
[22]
and Null, Clair and Luby, Stephen P
Arnold, Benjamin F. and Null, Clair and Luby, Stephen P. and Unicomb, Leanne and Stewart, Christine P. and Dewey, Kathryn G. and Ahmed, Tahmeed and Ashraf, Sania and Christensen, Gwen and Clasen, Thomas and Dentz, Holly N. and Fernald, Lia C.H. and Haque, Rashidul and Hubbard,...
-
[23]
and Jiang, Zhiling and Park, Eden and Qu, Annie , title =
Song, Peter X.-K. and Jiang, Zhiling and Park, Eden and Qu, Annie , title =. Statistics in Medicine , year =
-
[24]
Arnold, Clair Null, Stephen P
Benjamin F. Arnold, Clair Null, Stephen P. Luby, Leanne Unicomb, Christine P. Stewart, Kathryn G. Dewey, Tahmeed Ahmed, Sania Ashraf, Gwen Christensen, Thomas Clasen, Holly N. Dentz, Lia C.H. Fernald, Rashidul Haque, Alan E. Hubbard, Patricia Kariger, Elli Leontsini, Audrie Li...
2013 doi
-
[25]
Bang and J
H. Bang and J. M. Robins. Doubly robust estimation in missing data and causal inference models. Biometrics, 61: 0 962--973, 2005
2005
-
[26]
Chernozhukov, D
V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins. Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal, 21: 0 C1--C68, 2018
2018
-
[27]
M. P. Fay and B. I. Graubard. Small-sample adjustments for W ald-type tests using sandwich estimators. Biometrics, 57: 0 1198--1206, 2001
2001
-
[28]
L. P. Hansen. Large sample properties of generalized method of moments estimators. Econometrica, 50: 0 1029--1054, 1982
1982
-
[29]
R. J. Hayes and L. H. Moulton. Cluster Randomised Trials. Chapman & Hall/CRC, 2009
2009
-
[30]
Kauermann and R
G. Kauermann and R. J. Carroll. A note on the efficiency of sandwich covariance matrix estimation. Journal of the American Statistical Association, 96: 0 1387--1396, 2001
2001
-
[31]
Liang and S
K.-Y. Liang and S. L. Zeger. Longitudinal data analysis using generalized linear models. Biometrika, 73: 0 13--22, 1986
1986
-
[32]
B. Lu, J. S. Preisser, B. F. Qaqish, C. Suchindran, S. I. Bangdiwala, and M. Wolfson. A comparison of two bias-corrected covariance estimators for generalized estimating equations. Biometrics, 63: 0 935--941, 2007
2007
-
[33]
Luby, Mahbubur Rahman, Benjamin F
Stephen P. Luby, Mahbubur Rahman, Benjamin F. Arnold, Leanne Unicomb, Sania Ashraf, Peter J. Winch, Christine P. Stewart, Farjana Begum, Faruq Hussain, Mohammad Yunus, Joy Chakraborty, Thomas F. Clasen, Elli Leontsini, Abu Mohd. Naser, Sarker Parvez, Mahbubur Rahman, Rubhana R...
2018 doi
-
[34]
Luijken, R
K. Luijken, R. H. H. Groenwold, B. Van Calster, E. W. Steyerberg, and M. van Smeden. Impact of predictor measurement heterogeneity across settings on the performance of prediction models: A measurement error perspective. Statistics in Medicine, 38: 0 3444--3459, 2019
2019
-
[35]
L. A. Mancl and T. A. DeRouen. A covariance estimator for GEE with improved small-sample properties. Biometrics, 57: 0 126--134, 2001
2001
-
[36]
J. S. Preisser, M. L. Young, D. J. Zaccaro, and M. Wolfson. An integrated population-averaged approach to the design, analysis and sample size determination of cluster-unit trials. Statistics in Medicine, 22: 0 1235--1254, 2003
2003
-
[37]
J. S. Preisser, B. Lu, and B. F. Qaqish. Finite sample adjustments in estimating equations and covariance estimators for intracluster correlations. Statistics in Medicine, 27: 0 5764--5785, 2008
2008
-
[38]
A. Qu, B. G. Lindsay, and B. Li. Improving generalised estimating equations using quadratic inference functions. Biometrika, 87: 0 823--836, 2000
2000
-
[39]
Rabideau and Rui Wang
Dustin J. Rabideau and Rui Wang. Multiply robust generalized estimating equations for cluster randomized trials with missing outcomes. Statistics in Medicine, 43 0 (7): 0 1458--1474, 2024. doi:10.1002/sim.10027
2024 doi
-
[40]
J. M. Robins, A. Rotnitzky, and L. P. Zhao. Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association, 89: 0 846--866, 1994
1994
-
[41]
D. B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66: 0 688--701, 1974
1974
-
[42]
S. R. Seaman and A. J. Copas. Doubly robust generalised estimating equations for longitudinal data. Statistics in Medicine, 28: 0 937--955, 2009
2009
-
[43]
Song, Zhiling Jiang, Eden Park, and Annie Qu
Peter X.-K. Song, Zhiling Jiang, Eden Park, and Annie Qu. Quadratic inference functions in marginal models for longitudinal data. Statistics in Medicine, 28 0 (29): 0 3683--3696, 2009. doi:10.1002/sim.3719
2009 doi
-
[44]
N. Wang. Marginal nonparametric kernel regression accounting for within-subject correlation. Biometrika, 90: 0 43--52, 2003
2003
-
[45]
H. Yu, F. Li, and E. L. Turner. An evaluation of quadratic inference functions for estimating intervention effects in cluster randomized trials. Contemporary Clinical Trials Communications, 19: 0 100605, 2020
2020
-
[46]
H. Yu, G. Tong, and F. Li. A note on the estimation and inference with quadratic inference functions for correlated outcomes. Communications in Statistics -- Simulation and Computation, 51: 0 6525--6536, 2022
2022
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.