Pith. sign in

REVIEW 3 major objections 4 minor 54 references

Marginal generalized raking with parametric working models

T0 review · 3 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read Generalized raking can be aimed directly at marginal estimands such as average treatment effects and relative risks, and when calibrated on the efficient influence function it matches the asymptotic performance of the best augmented inverse

desk verdict MGR is a useful practical twist on raking for marginal estimands, but Theorem 1's efficiency claim under partial misspecification does not follow from the proof — the remainder is only O_P(n^{-1/2}). read the letter →

arxiv 2607.29629 v1 pith:34IIVILM submitted 2026-07-31 stat.ME

classification stat.ME MSC 62D0562F12
keywords marginalgeneralizedrakingefficientinfluencefunctionmissingdatatwo-phasestudiesaveragetreatmenteffectmultiplerobustnessAIPCW
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Generalized raking, a survey-sampling device that recalibrates inverse-probability weights to match known or estimated totals, has until now been used mainly to estimate regression coefficients. This paper claims that the same machinery can be aimed directly at marginal targets—mean outcomes under a treatment, average treatment effects, relative risks—by calibrating the weights on the projection of the marginal estimand's efficient influence function onto the always-observed variables. The resulting estimator, marginal generalized raking (MGR), is claimed to be asymptotically equivalent to the best augmented inverse-probability-weighted estimator and to be efficient when the parametric working models are right, while staying multiply robust: consistent if one model at the missing-data level and one model at the outcome/exposure level is correct. This matters because marginal quantities are the natural targets in causal analyses with missing or error-prone data, and raking is already implemented in standard software, so optimal efficiency becomes available without a bespoke estimator. Simulations and a cohort of people living with HIV are used to show that MGR delivers the promised efficiency and beats the naive strategy of fitting a conditional raked regression and then marginalizing.

What carries the argument

The optimal raking variable η(v) = E{φ1,P(Y, X, L) | V = v}, the projection of the parametric efficient influence function for the marginal estimand onto the always-observed variables V. This is the auxiliary variable used in the calibration constraint; it plays the same role in the missing-data raking as the optimal augmentation term in AIPCW, and it is the object whose estimation (via regression, multiple imputation, or error-prone proxies) determines whether the calibration removes the influence of the missing-data mechanism. The efficient influence function itself, built from the treatment-specific-mean influence function and its projection, carries the argument: calibrating weights to i

What would settle it

Estimate the empirical influence function of MGR in a simulated two-phase sample where the treatment-propensity model is misspecified but the outcome regression is correct, and compare its variance to the theoretical efficient-variance formula; Theorem 1 predicts exact agreement (up to Monte Carlo error) regardless of which single model is correct at each level. A statistically significant mismatch would falsify the asymptotic-linearity claim under partial misspecification.

Watch

Extended reading notes

Core claim

The central claim is Theorem 1: when the missing-data probability and the outcome/treatment models are estimated parametrically, the MGR estimator is asymptotically linear with influence function equal to the efficient influence function φ1,obs,P, provided at least one of the missing-data model or the EIF projection is consistent and at least one of the treatment-propensity model or outcome-regression model is consistent. Consequently, if all four nuisance models are correct, MGR is semiparametrically efficient. The proof proceeds by showing MGR is asymptotically equivalent to the AIPCW-AIPTW estimator and that the remainder is a product of errors from the two levels, so it vanishes whenever

Load-bearing premise

The argument leans on the assumption that at least one model in each of the two pairs—missingness (π or η) and outcome/exposure (g or Q)—is correctly specified, because the error term that drives the theory is a product of one error from each level; if both models in either pair are wrong, the asymptotic-linearity and efficiency statement collapses.

Editorial extensions

If this is right

  • MGR gives efficient estimation of average treatment effects and relative risks in two-phase or error-prone studies whenever one model at each of two levels is correct, with the same asymptotic variance as the optimal AIPCW estimator.
  • Because the weights are calibrated, the estimates stay inside the parameter space (e.g., risk differences respect bounds), unlike some AIPCW implementations that can fall outside the range of the estimand.
  • Directly raking on the marginal EIF removes the need to fit regression parameters first and marginalize, so MGR is asymptotically at least as efficient as conditional GR followed by the delta method, and more efficient when the propensity model is misspecified.
  • The multiply robust consistency property—correct specification of either the missing-data model or the EIF projection, plus either the outcome regression or the propensity score—generalizes double robustness to two levels of nuisance models.
  • The approach can be implemented with standard raking software, making semiparametric efficient missing-data estimation accessible in routine practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next step, which the paper flags but does not take, is a nonparametric version where the outcome regression, propensity score, and missing-data model are estimated by flexible machine-learning tools; the parametric result here sets the efficiency baseline for such an extension.
  • The calibration constraint equates the weighted sum of η over the observed subsample to the full-cohort sum of η, so comparing calibrated and uncalibrated inverse-probability weights in a given dataset could serve as a practical diagnostic for how informative—or how misspecified—the EIF projection is.
  • Because MGR is asymptotically equivalent to AIPCW-AIPTW, it likely inherits known finite-sample sensitivities of augmented weighting; raking's nonnegative, bounded weights may soften extreme-weight problems, but that is a testable conjecture rather than a claim of this paper.
  • The multiple-imputation route to estimating η suggests that in error-prone electronic health record data, standard MI software plus a raking step could replace bespoke measurement-error estimators, provided the imputation model is rich enough to capture the EIF projection.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes marginal generalized raking (MGR) for estimating marginal estimands (e.g., treatment-specific means, ATE, RR) in two-phase and coarsened-data settings. MGR calibrates inverse-probability-of-coarsening weights using a projection of the efficient influence function (EIF) for the marginal target, and then evaluates an augmented inverse-probability-of-treatment estimator with the calibrated weights. The paper claims that MGR is asymptotically equivalent to the AIPCW-AIPTW estimator, is multiply robust (consistent if one model from each of two pairs is correct), and is asymptotically linear with the efficient influence function when one model from each pair is correctly specified; efficiency is claimed when all four nuisance models are correct. The claims are supported by simulation studies comparing MGR, conditional generalized raking (CGR), and AIPCW-AIPTW across five misspecification scenarios, and by a validation-sampling analysis of an HIV cohort.

Significance. If the central theorem is correct, MGR offers a practical improvement over existing doubly robust estimators: it is implemented with standard survey raking software, respects the parameter-space bounds, and avoids the need to marginalize conditional regression estimates. The paper provides reproducible code and an extensive simulation study that demonstrates close numerical agreement between MGR and AIPCW-AIPTW in correctly specified settings. However, the proof of the key efficiency claim under partial misspecification is flawed: the first-order remainder term does not vanish at the claimed rate, and the equivalence lemma contains an algebraic error. The simulation results in Scenario 4 actually corroborate the breakdown of the variance estimator based on the efficient influence function. The paper's central claim is defensible only in a weaker form, so substantial revision is needed.

major comments (3)
  1. [Theorem 1; S3.1, Lemma S3] Theorem 1 asserts that if (i) one of (B3) or (B4) and (ii) one of (B1) or (B2) hold, then μ1,n,MGR is asymptotically linear with influence function equal to the efficient influence function φ1,obs,P. This is not supported by the proof. In Lemma S3, the remainder R(P,P0) is the sum of E0{(πP−π0)/πP [E0{φF_P|V}−EP{φF_P|V}]} and E0{(gP−g0)/gP [QP−Q0]}. Under partial misspecification, one factor in each product is O_P(n^{-1/2}) while the other converges to a nonzero limit. For example, if π is consistent but η is misspecified, then (πP−π0)=O_P(n^{-1/2}) while the bracket converges to a nonzero constant, so the product is O_P(n^{-1/2}) — not o_P(n^{-1/2}) as the proof claims. The same occurs if g is consistent but Q is misspecified. Consequently, R_n contributes to the first-order asymptotic distribution, so the asymptotic linear representation with φ1,obs,P fails and the estimator is not eff
  2. [S3.1, Lemma S2] The proof of asymptotic equivalence between MGR and AIPCW-AIPTW contains an algebraic error in the centering. The text defines ηu_n(v)=η_n(v)+μ, but then uses the term {ηu_n(V_i)+μ} in the estimating equation, which equals η_n(V_i)+2μ, not η_n(V_i)+μ. More importantly, solving the displayed estimating equation does not yield μ = (1/n)Σ R_i/π_n [I(X_i=1)/g_n {Y_i−Q_n}+Q_n] + o_P(n^{-1/2}). The μ terms on the right-hand side do not cancel: after rearrangement, the coefficient of μ involves 1−2R_i/π_n (up to the exact form of the centering), so the claimed cancellation is unjustified. Since Lemma S2 is used directly in the proof of Theorem 1, the derivation as written does not establish the equivalence. A corrected derivation should start from the correctly centered EIF, e.g., φ = (r/π)(A−μ) + (1−r/π)(η−μ) with η=E[A|V], and then verify the calibration constraint properly.
  3. [Section 3.4, Tables 5 and S10–S12] The simulation results in Scenario 4 provide a clean empirical falsification of the theorem's partial-misspecification efficiency claim. In Scenario 4 the outcome regression and the optimal raking variable are misspecified, while the propensity score and missing-data model are correctly specified. Under Theorem 1, MGR should be asymptotically linear with the efficient influence function and the ASE-based coverage should approach 0.95. Instead, coverage is 0.891, 0.844, 0.804, and 0.780 at n=1000, 2000, 4000, and 8000 in Table 5, with similar patterns in Tables S10–S12. The systematic under-coverage shows that the variance estimator based on φ1,obs,P is not valid when either Q or η is misspecified, consistent with the O_P(n^{-1/2}) remainder term. The simulation section should be revised to acknowledge this limitation explicitly, rather than presenting MGR as fully efficient in this setti
minor comments (4)
  1. [Tables 2–5 and S2–S15] The abbreviations in the table footnotes contain repeated entries (e.g., 'RR: relative risk' appears several times in the same footnote). Please clean up the footnote formatting for consistency.
  2. [Algorithm 2, Step 7] The variance formula in Step 7 is written as a double sum over i and j of a matrix product, but the notation is unclear and likely should be a single sum of outer products of the estimated influence function. Please clarify the expression.
  3. [S3.1, Proof of Theorem 1] In the proof, the text refers to '[(A2) or (B2)]' and '[(A1) or (B1)]' when the relevant assumptions for MGR are B1–B4. This inconsistent use of the CGR assumptions (A1–A3) is confusing and should be corrected.
  4. [Table S1] The column labeled 'Y' uses '—' for some rows without explanation, and the table formatting for the continuous-outcome rows is not fully aligned. Please add a footnote clarifying the '—' entries.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the marginal raking equivalence is derived from calibration constraints and standard semiparametric theory, not restated from inputs.

full rationale

The central claim (Theorem 1) is not circular. MGR is defined by calibrating weights with an estimated raking variable ηn, while AIPCW-AIPTW is defined by an explicit augmentation term. The proof (Lemmas S2–S3) derives their asymptotic equivalence from the calibration equation Σηn = ΣR q π^{-1} ηn and from the standard second-order remainder decomposition of AIPCW estimators. The influence-function representation is a consequence of the expansion, not an input assumption. The choice of the optimal raking variable η(v)=E{φ1,P(Y,X,L)|V=v} is imported from external semiparametric efficiency theory (Robins & Rotnitzky 1995; Lumley et al. 2011), and the paper does not define MGR in terms of the efficient influence function so as to make Theorem 1 true by definition. Self-citations to Williamson et al. (2026) are motivational/comparative and do not carry the proof; Lumley et al. (2011) is foundational and the paper supplies its own proof for the marginal extension. Simulations and the VCCC analysis are checked against externally specified data-generating mechanisms and validated outcomes, so no fitted parameter is relabeled as a prediction. I note two technical concerns that are correctness issues, not circularity: the algebra in Lemma S2 involving ηu and μ appears to double-count μ, and under partial misspecification Lemma S3's remainder is O_P(n^{-1/2}) rather than o_P(n^{-1/2}), so Theorem 1's strong efficiency claim under partial consistency is likely unsupported. These concerns do not make the derivation circular.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No ad hoc fitted constants or new entities are introduced. The method's parameters are standard nuisance functions estimated from data. The main load-bearing assumptions are MAR, positivity, correct parametric/nuisance specification for efficiency, and standard semiparametric theory.

assumptions (5)
  • domain assumption Missing at random: R ⊥ D | V
    Section 2.1; required for the observed-data identification formulas and for the form of the observed-data EIF.
  • domain assumption Positivity: π(v) > ε > 0 and 0 < g(l) < 1
    Section 2.2 and Assumptions A2/B3; inverse weighting and raking calibration require bounded weights.
  • domain assumption The outcome regression follows a parametric GLM Q(x,l) = h^{-1}{β0 + (x,l)β}
    Section 2.1; the parametric EIF projection and efficiency results are defined relative to this model.
  • domain assumption Nuisance estimators converge to their targets at n^{-1/2} rates under correct specification (B1-B4)
    Theorem 1 conditions; the multiply-robust claims trade off these consistency assumptions.
  • standard math Standard semiparametric efficiency theory: the observed-data EIF is r/π φF - (r-π)/π E(φF|V)
    Used in Sections 2.3-2.5 and in the supplementary proofs; from Bickel et al. (1993) and Tsiatis (2006).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Marginal generalized raking with parametric working models." pith.science (2026). https://pith.science/paper/34IIVILM

@misc{pith2026260729629,
  author       = {Pith},
  title        = {Pith review of: Marginal generalized raking with parametric working models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/34IIVILM}},
  note         = {Machine review of arXiv:2607.29629}
}
read the original abstract

Generalized raking (GR) was originally developed in the survey statistics literature to incorporate auxiliary information in estimation. Recently, it has been used in the biostatistical and epidemiological literature to estimate regression coefficients in parametric models in cases with missing data, including missing data by design (e.g., two-phase studies). In the regression parameter context, the optimal GR estimator has been shown to be equivalent to the optimal augmented inverse probability weighted estimator. In this paper, we generalize the influence function-based theory for GR to marginal estimands; we call our approach \textit{marginal generalized raking}. We compare our approach to a naive procedure that marginalizes a conditional GR estimator of regression parameters in both fully-synthetic simulations and in an application using data from an observational cohort of persons living with HIV.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 1 linked inside Pith

  1. [1]

    Biometrics , volume=

    Doubly robust estimation in missing data and causal inference models , author=. Biometrics , volume=

  2. [2]

    1993 , publisher=

    Efficient and adaptive estimation for semiparametric models , author=. 1993 , publisher=

  3. [3]

    Machine Learning , volume=

    Random forests , author=. Machine Learning , volume=

  4. [4]

    Improved

    Breslow, Norman E and Lumley, Thomas and Ballantyne, Christie M and Chambless, Lloyd E and Kulich, Michal , journal=. Improved

  5. [5]

    American Journal of Epidemiology , volume=

    Using the whole cohort in the analysis of case-cohort data , author=. American Journal of Epidemiology , volume=

  6. [6]

    Biometrika , volume=

    Improving efficiency and robustness of the doubly robust estimator for a population mean with incomplete data , author=. Biometrika , volume=

  7. [7]

    2018 , journal=

    Double/debiased machine learning for treatment and structural parameters , author=. 2018 , journal=

  8. [8]

    Epidemiology , volume=

    The consistency statement in causal inference: a definition or an assumption? , author=. Epidemiology , volume=

Show all 54 references
  1. [9]

    Journal of the American Statistical Association , volume=

    Calibration estimators in survey sampling , author=. Journal of the American Statistical Association , volume=

  2. [10]

    Journal of the American Statistical Association , volume=

    Generalized raking procedures in survey sampling , author=. Journal of the American Statistical Association , volume=

  3. [11]

    Biostatistics , volume=

    Machine learning in the estimation of causal effects: targeted minimum loss-based estimation and double/debiased machine learning , author=. Biostatistics , volume=

  4. [12]

    arXiv preprint arXiv:2405.15242 , year=

    Causal machine learning methods and use of sample splitting in settings with high-dimensional confounding , author=. arXiv preprint arXiv:2405.15242 , year=

  5. [13]

    Annals of Statistics , volume=

    Greedy function approximation: a gradient boosting machine , author=. Annals of Statistics , volume=

  6. [14]

    The Annals of Applied Statistics , volume=

    Accounting for dependent errors in predictors and time-to-event outcomes using electronic health records, validation samples, and multiple imputation , author=. The Annals of Applied Statistics , volume=

  7. [15]

    Scandinavian Journal of Statistics , volume=

    Combining inverse probability weighting and multiple imputation to improve robustness of estimation , author=. Scandinavian Journal of Statistics , volume=

  8. [16]

    Statistics in Medicine , volume=

    Combining multiple imputation with raking of weights: An efficient and robust approach in the setting of nearly true models , author=. Statistics in Medicine , volume=

  9. [17]

    The Annals of Statistics , pages=

    Ignorability and coarse data , author=. The Annals of Statistics , pages=

  10. [18]

    Biometrics , volume=

    Efficient nonparametric inference on the effects of stochastic interventions under two-phase sampling, with applications to vaccine efficacy trials , author=. Biometrics , volume=

  11. [19]

    Journal of the American Statistical Association , volume=

    A generalization of sampling without replacement from a finite universe , author=. Journal of the American Statistical Association , volume=

  12. [20]

    Sociological Methods & Research , volume=

    A comparison of three popular methods for handling missing data: complete-case analysis, inverse probability weighting, and multiple imputation , author=. Sociological Methods & Research , volume=

  13. [21]

    International Statistical Review , volume=

    Connections between survey calibration estimators and semiparametric models for incomplete data , author=. International Statistical Review , volume=

  14. [22]

    2007 , journal=

    Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data , author=. 2007 , journal=

  15. [23]

    Journal of the American Statistical Association , volume=

    Improving the efficiency of relative-risk estimation in case-cohort studies , author=. Journal of the American Statistical Association , volume=

  16. [24]

    Considerations for analysis of time-to-event outcomes measured with error: bias and correction with

    Oh, Eric J and Shepherd, Bryan E and Lumley, Thomas and Shaw, Pamela A , journal=. Considerations for analysis of time-to-event outcomes measured with error: bias and correction with

  17. [25]

    Biometrical Journal , volume=

    Improved generalized raking estimators to address dependent covariate and failure-time outcome error , author=. Biometrical Journal , volume=

  18. [26]

    International Journal of Epidemiology , volume=

    Practical considerations for specifying a super learner , author=. International Journal of Epidemiology , volume=

  19. [27]

    Biometrics , volume=

    Addressing confounding and continuous exposure measurement error using corrected score functions , author=. Biometrics , volume=

  20. [28]

    Journal of the American Statistical Association , volume=

    Estimation of regression coefficients when some regressors are not always observed , author=. Journal of the American Statistical Association , volume=

  21. [29]

    Journal of the American Statistical Association , volume=

    Semiparametric efficiency in multivariate regression models with missing data , author=. Journal of the American Statistical Association , volume=

  22. [30]

    Journal of the American Statistical Association , volume=

    Analysis of semiparametric regression models for repeated outcomes in the presence of missing data , author=. Journal of the American Statistical Association , volume=

  23. [31]

    Biometrika , volume=

    Inference for imputation estimators , author=. Biometrika , volume=

  24. [32]

    The International Journal of Biostatistics , volume=

    A targeted maximum likelihood estimator for two-stage designs , author=. The International Journal of Biostatistics , volume=

  25. [33]

    Flexible imputation of missing data, second edition , pages=

    Multiple imputation , author=. Flexible imputation of missing data, second edition , pages=. 2018 , publisher=

  26. [34]

    Statistical Science , volume=

    Introduction to double robust methods for incomplete data , author=. Statistical Science , volume=

  27. [35]

    Journal of the American Statistical Association , volume=

    Adjusting for nonignorable drop-out using semiparametric nonresponse models , author=. Journal of the American Statistical Association , volume=

  28. [36]

    Biometrics , volume=

    Multiwave validation sampling for error-prone electronic health records , author=. Biometrics , volume=

  29. [37]

    Biometrics , volume=

    Double robust variance estimation with parametric working models , author=. Biometrics , volume=

  30. [38]

    Annals of Statistics , volume=

    Semiparametric theory for causal mediation analysis: efficiency bounds, multiple robustness, and sensitivity analysis , author=. Annals of Statistics , volume=

  31. [39]

    Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume=

    Regression shrinkage and selection via the lasso , author=. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume=

  32. [40]

    2006 , publisher=

    Semiparametric theory and missing data , author=. 2006 , publisher=

  33. [41]

    2018 , publisher=

    Flexible Imputation of Missing Data , author=. 2018 , publisher=

  34. [42]

    2003 , publisher=

    Unified methods for censored longitudinal data and causality , author=. 2003 , publisher=

  35. [43]

    The International Journal of Biostatistics , volume=

    Targeted Maximum Likelihood Learning , author=. The International Journal of Biostatistics , volume=

  36. [44]

    Statistical Applications in Genetics and Molecular Biology , volume=

    Super Learner , author=. Statistical Applications in Genetics and Molecular Biology , volume=

  37. [45]

    2011 , publisher=

    Targeted learning: causal inference for observational and experimental data , author=. 2011 , publisher=

  38. [46]

    2000 , publisher=

    Asymptotic statistics , author=. 2000 , publisher=

  39. [47]

    Statistics in Medicine , year=

    Assessing treatment effects in observational data with missing confounders: a comparative study of practical doubly-robust and traditional missing-data methods , author=. Statistics in Medicine , year=

  40. [48]

    2019 , note =

    gam: Generalized Additive Models , author=. 2019 , note =

  41. [49]

    Journal of Statistical Software , year =

    Regularization Paths for Generalized Linear Models via Coordinate Descent , author =. Journal of Statistical Software , year =

  42. [50]

    2024 , organization =

    marginaleffects: Predictions, Comparisons, Slopes, Marginal Means, and Hypothesis Tests , author =. 2024 , organization =

  43. [51]

    Journal of Statistical Software , year =

    Stef. Journal of Statistical Software , year =

  44. [52]

    Wright and Andreas Ziegler , journal =

    Marvin N. Wright and Andreas Ziegler , journal =. 2017 , volume =

  45. [53]

    survey: Analysis of Complex Survey Samples , note =

    Thomas Lumley , year =. survey: Analysis of Complex Survey Samples , note =

  46. [54]

    Tianqi Chen and Tong He and Michael Benesty and Vadim Khotilovich and Yuan Tang and Hyunsu Cho and others , year =

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.