{"id":"1c0c2d83-6be5-4c62-b55e-8dca148c523d","arxiv_id":"2607.29629","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Marginal generalized raking directly targets treatment-specific means, ATE, and RR by raking weights on the marginal efficient influence function, and is asymptotically equivalent to optimal AIPCW estimators.","lead":"Researchers propose a new estimator, marginal generalized raking, for treatment effects in studies with missing or error-prone data. It calibrates survey-style weights using the efficient influence function of the marginal target, and the authors show it matches the efficiency of augmented inverse probability weighting with simpler implementation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1 overclaims efficient influence-function representation under partial misspecification; the remainder term in the proof is O_P(n^{-1/2}) when exactly one nuisance in each pair is misspecified, so efficiency fails.","rationale":"The reader's weakest assumption identified the same load-bearing concern: the efficient-influence-function statement under partial misspecification is not established, and the actual influence function is not efficient unless all four nuisance parameters are consistently estimated. My analysis of the supplementary proof confirms this: the remainder R(P,P0) in Lemma S3 is first-order when exactly one of each pair is misspecified, so the proof of Theorem 1 is invalid as written. This is the central theoretical claim of the paper, and it has practical consequences: the reported coverage in Scenario 4 suggests the efficiency-based variance estimator is not valid under partial misspecification. However, the paper's practical contributions — consistency of MGR, equivalence with AIPCW-AIPTW in correctly specified settings, and the simulation/application evidence — remain largely intact, and the theorem could be corrected by weakening the conclusion to consistency (not efficiency) under partial consistency, and efficiency only when (B1)–(B4) all hold. Therefore the reader's CONDITIONAL verdict remains appropriate; I do not see a reason to change it to ACCEPT or REJECT. The proposed concrete test would settle whether the influence-function claim fails, as I expect it would.","tokens_in":37652,"tokens_out":7735,"duration_ms":83787,"concrete_test":"Independently re-derive the first-order influence function of MGR under the partial-misspecification scenario (B3 and B2 hold; B1 and B4 fail), e.g., with Qn converging to a fixed misspecified limit Q*. Compute the asymptotic variance via the empirical sandwich or influence-function expansion, and compare it to E{φ1,obs,P^2}. If the two differ, Theorem 1 is false. A complementary simulation: in a simple DGP with binary L, fit Qn misspecified by omitting L while keeping π and g correctly specified, run a large-n (n ≥ 10^5) Monte Carlo, and compare the sampling variance of MGR to the efficient-influence-function variance. If the sampling variance exceeds the efficient bound substantially, the efficiency claim under partial misspecification is refuted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1 asserts that if (i) one of B3/B4 and (ii) one of B1/B2 hold, then MGR is asymptotically linear with influence function equal to the efficient influence function φ1,obs,P. This is not supported by the proof. In Lemma S3, the remainder R(P,P0) is the sum of two product terms: E0{(πP−π0)/πP [E0{φF_P|V}−EP{φF_P|V}]} and E0{(gP−g0)/gP [QP−Q0]}. Under partial misspecification, one factor in each product converges to zero at rate n^{-1/2} while the other converges to a non-zero limit (e.g., QP→Q*≠Q0). The product is therefore O_P(n^{-1/2}), not o_P(n^{-1/2}), so R_n contributes to the first-order asymptotic distribution. Consequently, the asymptotic linear representation with influence function φ1,obs,P fails, and the estimator is not efficient under these partial-consistency conditions. Consistency may still hold, but the influence function contains additional first-order terms from the misspecified nuisance functions. This is consistent with the under-coverage observed for MGR in Scenario 4 (Tables 5 and S10–S12), where Q and η are misspecified: the ASE based on the efficient influence function underestimates the empirical variance. Additionally, the algebraic step in Lemma S2 double-counts μ: ηu is defined as η+μ but then used as {ηu+μ}; the subsequent cancellation to μ = A is therefore incorrect, further undermining the equivalence claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes marginal generalized raking (MGR) for estimating marginal estimands (e.g., treatment-specific means, ATE, RR) in two-phase and coarsened-data settings. MGR calibrates inverse-probability-of-coarsening weights using a projection of the efficient influence function (EIF) for the marginal target, and then evaluates an augmented inverse-probability-of-treatment estimator with the calibrated weights. The paper claims that MGR is asymptotically equivalent to the AIPCW-AIPTW estimator, is multiply robust (consistent if one model from each of two pairs is correct), and is asymptotically linear with the efficient influence function when one model from each pair is correctly specified; efficiency is claimed when all four nuisance models are correct. The claims are supported by simulation studies comparing MGR, conditional generalized raking (CGR), and AIPCW-AIPTW across five misspecification scenarios, and by a validation-sampling analysis of an HIV cohort.","tokens_in":38066,"tokens_out":7756,"duration_ms":86099,"significance":"If the central theorem is correct, MGR offers a practical improvement over existing doubly robust estimators: it is implemented with standard survey raking software, respects the parameter-space bounds, and avoids the need to marginalize conditional regression estimates. The paper provides reproducible code and an extensive simulation study that demonstrates close numerical agreement between MGR and AIPCW-AIPTW in correctly specified settings. However, the proof of the key efficiency claim under partial misspecification is flawed: the first-order remainder term does not vanish at the claimed rate, and the equivalence lemma contains an algebraic error. The simulation results in Scenario 4 actually corroborate the breakdown of the variance estimator based on the efficient influence function. The paper's central claim is defensible only in a weaker form, so substantial revision is needed.","major_comments":[{"comment":"Theorem 1 asserts that if (i) one of (B3) or (B4) and (ii) one of (B1) or (B2) hold, then μ1,n,MGR is asymptotically linear with influence function equal to the efficient influence function φ1,obs,P. This is not supported by the proof. In Lemma S3, the remainder R(P,P0) is the sum of E0{(πP−π0)/πP [E0{φF_P|V}−EP{φF_P|V}]} and E0{(gP−g0)/gP [QP−Q0]}. Under partial misspecification, one factor in each product is O_P(n^{-1/2}) while the other converges to a nonzero limit. For example, if π is consistent but η is misspecified, then (πP−π0)=O_P(n^{-1/2}) while the bracket converges to a nonzero constant, so the product is O_P(n^{-1/2}) — not o_P(n^{-1/2}) as the proof claims. The same occurs if g is consistent but Q is misspecified. Consequently, R_n contributes to the first-order asymptotic distribution, so the asymptotic linear representation with φ1,obs,P fails and the estimator is not eff","section":"Theorem 1; S3.1, Lemma S3"},{"comment":"The proof of asymptotic equivalence between MGR and AIPCW-AIPTW contains an algebraic error in the centering. The text defines ηu_n(v)=η_n(v)+μ, but then uses the term {ηu_n(V_i)+μ} in the estimating equation, which equals η_n(V_i)+2μ, not η_n(V_i)+μ. More importantly, solving the displayed estimating equation does not yield μ = (1/n)Σ R_i/π_n [I(X_i=1)/g_n {Y_i−Q_n}+Q_n] + o_P(n^{-1/2}). The μ terms on the right-hand side do not cancel: after rearrangement, the coefficient of μ involves 1−2R_i/π_n (up to the exact form of the centering), so the claimed cancellation is unjustified. Since Lemma S2 is used directly in the proof of Theorem 1, the derivation as written does not establish the equivalence. A corrected derivation should start from the correctly centered EIF, e.g., φ = (r/π)(A−μ) + (1−r/π)(η−μ) with η=E[A|V], and then verify the calibration constraint properly.","section":"S3.1, Lemma S2"},{"comment":"The simulation results in Scenario 4 provide a clean empirical falsification of the theorem's partial-misspecification efficiency claim. In Scenario 4 the outcome regression and the optimal raking variable are misspecified, while the propensity score and missing-data model are correctly specified. Under Theorem 1, MGR should be asymptotically linear with the efficient influence function and the ASE-based coverage should approach 0.95. Instead, coverage is 0.891, 0.844, 0.804, and 0.780 at n=1000, 2000, 4000, and 8000 in Table 5, with similar patterns in Tables S10–S12. The systematic under-coverage shows that the variance estimator based on φ1,obs,P is not valid when either Q or η is misspecified, consistent with the O_P(n^{-1/2}) remainder term. The simulation section should be revised to acknowledge this limitation explicitly, rather than presenting MGR as fully efficient in this setti","section":"Section 3.4, Tables 5 and S10–S12"}],"minor_comments":[{"comment":"The abbreviations in the table footnotes contain repeated entries (e.g., 'RR: relative risk' appears several times in the same footnote). Please clean up the footnote formatting for consistency.","section":"Tables 2–5 and S2–S15"},{"comment":"The variance formula in Step 7 is written as a double sum over i and j of a matrix product, but the notation is unclear and likely should be a single sum of outer products of the estimated influence function. Please clarify the expression.","section":"Algorithm 2, Step 7"},{"comment":"In the proof, the text refers to '[(A2) or (B2)]' and '[(A1) or (B1)]' when the relevant assumptions for MGR are B1–B4. This inconsistent use of the CGR assumptions (A1–A3) is confusing and should be corrected.","section":"S3.1, Proof of Theorem 1"},{"comment":"The column labeled 'Y' uses '—' for some rows without explanation, and the table formatting for the continuous-outcome rows is not fully aligned. Please add a footnote clarifying the '—' entries.","section":"Table S1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses a practically important problem and includes a substantial simulation study with reproducible code. The main theoretical claim is, however, overstated: the proof of Theorem 1 does not establish asymptotic linearity or efficiency under partial misspecification, and the equivalence lemma contains an algebraic error. The core idea of MGR is likely salvageable — the weaker consistency claim and the full-efficiency claim when all working models are correct may be recoverable — but the theorem, its proof, and the simulation interpretation need to be rewritten accordingly. I recommend major revision rather than rejection, provided the authors can correct the algebra and appropriately qualify the efficiency claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the practical contribution is real. Marginal generalized raking gives you a bound-respecting, multiply robust estimator of treatment-specific means, ATE, and RR that can be fit with the survey package, and the simulations show it performs essentially identically to AIPCW-AIPTW in finite samples. The VCCC analysis is a nice illustration of where it beats a naive marginalization. That part of the paper is solid.\n\nThe soft spot is Theorem 1. As stated, it claims asymptotic linearity with the efficient influence function under partial misspecification — one of B3/B4 and one of B1/B2. The proof doesn't deliver that. In Lemma S3 the remainder is the sum of two products of differences. If exactly one nuisance in each pair is misspecified, one factor in each product is O_P(n^{-1/2}) while the other converges to a non-zero limit, so the remainder is O_P(n^{-1/2}), not o_P. The influence function representation therefore does not hold; you get consistency but not efficiency, and the variance estimator based on the efficient influence function will undercover. The Scenario 4 results show exactly that pattern. There is also an algebraic slip in Lemma S2 — the definition η^u = η + μ is then used as η^u + μ, which double-counts μ and makes the cancellation to μ = A invalid. That needs fixing too.\n\nI'm not saying the MGR idea is wrong. The equivalence to AIPCW-AIPTW under correct specification is plausible and the numerics back it up. But the multiply robust efficiency claim is the headline, and it is not supported by the supplied proof. The paper needs a major revision: either prove a correct influence function under partial misspecification (with the extra terms) or restate the theorem to claim consistency under those conditions and relegate efficiency to the fully correct case. The variance estimation section should then be adjusted accordingly.\n\nWho is this for? Applied researchers in two-phase and EHR settings will find the method useful regardless of the theory fix. Methodologists will want the proof cleaned up before citing the efficiency result. It deserves a serious referee — the idea is worth publishing, but not as is.\n\nSend it to review with a 'major revision' recommendation.","headline":"MGR is a useful practical twist on raking for marginal estimands, but Theorem 1's efficiency claim under partial misspecification does not follow from the proof — the remainder is only O_P(n^{-1/2}).","tokens_in":38519,"tokens_out":2998,"would_cite":false,"duration_ms":34987,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62F12"],"pacs":[],"model":"deepseek-v4-flash","headline":"Generalized raking can be aimed directly at marginal estimands such as average treatment effects and relative risks, and when calibrated on the efficient influence function it matches the asymptotic performance of the best augmented inverse","keywords":["marginal generalized raking","generalized raking","efficient influence function","missing data","two-phase studies","average treatment effect","multiple robustness","AIPCW"],"falsifier":"Estimate the empirical influence function of MGR in a simulated two-phase sample where the treatment-propensity model is misspecified but the outcome regression is correct, and compare its variance to the theoretical efficient-variance formula; Theorem 1 predicts exact agreement (up to Monte Carlo error) regardless of which single model is correct at each level. A statistically significant mismatch would falsify the asymptotic-linearity claim under partial misspecification.","tokens_in":37574,"feed_emoji":"📊","tokens_out":5981,"duration_ms":70873,"temperature":0.7,"pith_summary":"Generalized raking, a survey-sampling device that recalibrates inverse-probability weights to match known or estimated totals, has until now been used mainly to estimate regression coefficients. This paper claims that the same machinery can be aimed directly at marginal targets—mean outcomes under a treatment, average treatment effects, relative risks—by calibrating the weights on the projection of the marginal estimand's efficient influence function onto the always-observed variables. The resulting estimator, marginal generalized raking (MGR), is claimed to be asymptotically equivalent to the best augmented inverse-probability-weighted estimator and to be efficient when the parametric working models are right, while staying multiply robust: consistent if one model at the missing-data level and one model at the outcome/exposure level is correct. This matters because marginal quantities are the natural targets in causal analyses with missing or error-prone data, and raking is already implemented in standard software, so optimal efficiency becomes available without a bespoke estimator. Simulations and a cohort of people living with HIV are used to show that MGR delivers the promised efficiency and beats the naive strategy of fitting a conditional raked regression and then marginalizing.","feed_headline":"Treatment-effect raking matches the optimal missing-data estimator","feed_subtitle":"Survey-style weight calibration on the efficient influence function gives multiply robust, efficient effect estimates from incomplete data.","key_machinery":"The optimal raking variable η(v) = E{φ1,P(Y, X, L) | V = v}, the projection of the parametric efficient influence function for the marginal estimand onto the always-observed variables V. This is the auxiliary variable used in the calibration constraint; it plays the same role in the missing-data raking as the optimal augmentation term in AIPCW, and it is the object whose estimation (via regression, multiple imputation, or error-prone proxies) determines whether the calibration removes the influence of the missing-data mechanism. The efficient influence function itself, built from the treatment-specific-mean influence function and its projection, carries the argument: calibrating weights to i","core_discovery":"The central claim is Theorem 1: when the missing-data probability and the outcome/treatment models are estimated parametrically, the MGR estimator is asymptotically linear with influence function equal to the efficient influence function φ1,obs,P, provided at least one of the missing-data model or the EIF projection is consistent and at least one of the treatment-propensity model or outcome-regression model is consistent. Consequently, if all four nuisance models are correct, MGR is semiparametrically efficient. The proof proceeds by showing MGR is asymptotically equivalent to the AIPCW-AIPTW estimator and that the remainder is a product of errors from the two levels, so it vanishes whenever","pith_inferences":["A natural next step, which the paper flags but does not take, is a nonparametric version where the outcome regression, propensity score, and missing-data model are estimated by flexible machine-learning tools; the parametric result here sets the efficiency baseline for such an extension.","The calibration constraint equates the weighted sum of η over the observed subsample to the full-cohort sum of η, so comparing calibrated and uncalibrated inverse-probability weights in a given dataset could serve as a practical diagnostic for how informative—or how misspecified—the EIF projection is.","Because MGR is asymptotically equivalent to AIPCW-AIPTW, it likely inherits known finite-sample sensitivities of augmented weighting; raking's nonnegative, bounded weights may soften extreme-weight problems, but that is a testable conjecture rather than a claim of this paper.","The multiple-imputation route to estimating η suggests that in error-prone electronic health record data, standard MI software plus a raking step could replace bespoke measurement-error estimators, provided the imputation model is rich enough to capture the EIF projection."],"forward_implications":["MGR gives efficient estimation of average treatment effects and relative risks in two-phase or error-prone studies whenever one model at each of two levels is correct, with the same asymptotic variance as the optimal AIPCW estimator.","Because the weights are calibrated, the estimates stay inside the parameter space (e.g., risk differences respect bounds), unlike some AIPCW implementations that can fall outside the range of the estimand.","Directly raking on the marginal EIF removes the need to fit regression parameters first and marginalize, so MGR is asymptotically at least as efficient as conditional GR followed by the delta method, and more efficient when the propensity model is misspecified.","The multiply robust consistency property—correct specification of either the missing-data model or the EIF projection, plus either the outcome regression or the propensity score—generalizes double robustness to two levels of nuisance models.","The approach can be implemented with standard raking software, making semiparametric efficient missing-data estimation accessible in routine practice."],"fun_headline_variants":["Double-robust raking for efficient effect estimates","Missing-data raking is multiply robust and efficient","Marginal raking matches optimal missing-data estimator","Raking for margins: efficient and multiply robust","Efficient influence via generalized raking"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The argument leans on the assumption that at least one model in each of the two pairs—missingness (π or η) and outcome/exposure (g or Q)—is correctly specified, because the error term that drives the theory is a product of one error from each level; if both models in either pair are wrong, the asymptotic-linearity and efficiency statement collapses.","fun_headline_variants_meta":{"raw":{"variants":["Double-robust raking for efficient effect estimates","Missing-data raking is multiply robust and efficient","Marginal raking matches optimal missing-data estimator","Raking for margins: efficient and multiply robust","Efficient influence via generalized raking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000334,"raw_usage":{"total_tokens":1643,"prompt_tokens":647,"completion_tokens":996,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":391,"completion_tokens_details":{"reasoning_tokens":926}},"tokens_in":391,"tokens_out":996,"duration_ms":9785,"temperature":1.0,"reasoning_tokens":926,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T03:08:53.649944+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Estimate the empirical influence function of MGR in a simulated two-phase sample where the treatment-propensity model is misspecified but the outcome regression is correct, and compare its variance to the theoretical efficient-variance formula; Theorem 1 predicts exact agreement (up to Monte Carlo error) regardless of which single model is correct at each level. A statistically significant mismatch would falsify the asymptotic-linearity claim under partial misspecification.","supporting_citations":[],"review_version":1}