{"id":"c5e72518-da56-4f4c-9c9f-ab2fc78f314a","arxiv_id":"2504.13467","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A sequential covariate-balancing estimator attains the semiparametric efficiency bound for moment-equation parameters under non-monotone MNAR models encoded by regular pattern graphs.","lead":"This paper presents an estimation method for regression parameters when data contain several non-monotone missing-not-at-random patterns. The method builds balancing weights sequentially over a pattern graph and is proven asymptotically efficient, with simulations and a survey analysis illustrating its behavior.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proof of Lemma H.8 is incomplete: the balancing-error bound that underpins Theorem 4.3 is asserted, not derived.","rationale":"The paper proposes a sequential balancing estimator for moment-equation parameters under non-monotone missingness and claims semiparametric efficiency. The reader's conditional verdict focused on the untestable overlap and unique-minimizer assumptions (D.1.A and D.1.B). I agree those are limitations, but they are explicit assumptions of the theorem and are standard in the missing-data literature. The most load-bearing concern I find is a proof gap in Lemma H.8, which is essential to the asymptotic linear expansion in Theorem D.4 and hence to the efficiency claim in Theorem 4.3. The proof of Lemma H.8 states an inequality without derivation and concludes the rate without controlling the relevant norm of β^r_θ uniformly. This is an internal-completeness issue rather than a disagreement with the consensus. The concrete test I propose is to fill in the derivation from the KKT conditions and check the implied rate; this would settle whether the gap is merely expository or hides a missing condition. Since this concern reinforces the need for careful verification, it does not change the reader's conditional verdict, but it shifts the reason toward proof completeness rather than only untestable assumptions.","tokens_in":43291,"tokens_out":14679,"duration_ms":126904,"concrete_test":"Derive the inequality in Lemma H.8 from the first-order (KKT) conditions of the penalized sequential loss (16). Then verify whether Assumptions D.1.D, D.1.G, and D.3.A imply sup_θ ‖β^r_θ‖_2 = O(√K_r) (or another bound sufficient for λ√PEN2(Φ_r^T β^r_θ) = o(N^{-1/2})) uniformly over θ ∈ Θ. A complete derivation that yields the stated rate would close the gap; failure would show Theorem 4.3 is unproven as written.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central efficiency claim (Theorem 4.3) rests on Theorem D.4, whose proof uses Lemma H.8 to show that the balancing-error term S^r_{θ,3} is asymptotically negligible. However, the proof of Lemma H.8 in Appendix H does not actually derive the key inequality. It states: \"Notice that S^r_{θ,3} is related to the balancing error: sup_θ |√N S^r_{θ,3}| ≤ λ { γ√K_r + 2(1-γ)√PEN2(Φ_r^T \\hat α[r]) } √PEN2(Φ_r^T β^r_θ)\" and then immediately concludes the rate op(1) without proving this inequality or showing that √PEN2(Φ_r^T β^r_θ) is uniformly bounded in a way that makes the product o(N^{-1/2}). Because S^r_{θ,3} is exactly the balancing condition that the sequential loss is designed to enforce, an omitted derivation here leaves a genuine gap in the asymptotic expansion. If this inequality or the claimed rate fails, the influence-function representation in Theorem D.4 does not follow, and Theorem 4.3's efficiency conclusion is unsupported. This concern is distinct from the overlap assumption (D.1.A), which is a standard, untestable distributional condition rather than a mathematical lacuna in the proof.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies estimation of parameters defined by moment equations when data have multiple non-monotone missing patterns and missingness is not at random. Building on pattern graphs and the authors' prior work on balancing weights, it proposes a sequential balancing loss that estimates the propensity odds O_r sequentially from complete cases, constructs weights ŵ = Σ_r Q̂_r, and solves a weighted estimating equation. The paper derives a semiparametric efficiency bound (Theorem 4.2), claims consistency and semiparametric efficiency of the sequential estimator (Theorem 4.3), and presents simulations and a real-data application. The theoretical development relies on a long appendix with empirical-process and bracketing-number arguments, including sieve approximations and several technical lemmas.","tokens_in":43602,"tokens_out":12542,"duration_ms":116385,"significance":"The intended contribution is significant: if Theorem 4.3 were rigorously established, the paper would provide a computationally simple, stable, and efficient estimator for a flexible class of MNAR/pattern-mixture models, extending the pattern-graph framework from density/imputation settings to general estimating equations. The efficiency-bound calculation in Theorem 4.2 is nontrivial, and the bracketing-number lemmas (H.12-H.15) are potentially reusable. The simulations show improvements over entropy and local balancing. However, the central asymptotic proof currently has a load-bearing gap and an apparent mismatch with the algorithm actually proposed, so the efficiency claim is not yet supported.","major_comments":[{"comment":"The decomposition at the start of the proof of Theorem D.4 is not the decomposition for θ_seq as defined in Algorithm 1. The proof writes P̂_N ψ_θ − E{ψ_θ(L)} = Σ_{r∈R} [N^{-1}Σ_i 1_{R_i=1_d} O_r(L[r]_i; α̂[r]) ψ_θ(L_i) − E{1_{R_i=r} ψ_θ(L)}]. But Algorithm 1 defines weights as ŵ = Σ_r Q̂_r, with Q̂_r = Ô_r Σ_{s∈Pa(r)} Q̂_s, so the complete-case weighted sum contains products of odds along paths, not the single factors Ô_r. Unless O_r is redefined to mean Q_r, which the paper does not do, the proof establishes asymptotic normality for a different estimator, and the influence function displayed at the end of Appendix F is not the V_θ of Theorem 4.2 for a nontrivial pattern graph. This mismatch directly undermines Theorem 4.3 as stated.","section":"Appendix F (proof of Theorem D.4)"},{"comment":"The proof of Lemma H.8 asserts the key balancing-error bound sup_θ |√N S^r_{θ,3}| ≤ λ{γ√K_r + 2(1−γ)√PEN2(Φ_r^T α̂[r])}√PEN2(Φ_r^T β^r_θ) and then concludes op(1). The inequality is never derived, and the notation PEN2 and γ is not defined anywhere in the paper. No argument shows that √PEN2(Φ_r^T β^r_θ) is uniformly bounded in a way that makes the product negligible, nor is the dependence on the recursive estimator Q̂_Pa(r) controlled. Since S^r_{θ,3} is precisely the balancing error that the sequential loss is designed to enforce, this is a load-bearing gap in the proof of Theorem D.4 and, consequently, in the efficiency claim of Theorem 4.3.","section":"Appendix H, Lemma H.8"},{"comment":"The uniform-overlap condition c_0 ≤ O_r(l[r]) ≤ C_0 in Assumption D.1.A is used at several load-bearing points, including Theorem D.2, Lemmas H.4-H.5, and the proof of Theorem D.4, but it involves distributions of missing variables and cannot be verified from the observed data alone. The manuscript should state explicitly that the efficiency guarantee is conditional on this assumption and should provide a sensitivity analysis or a diagnostic for the overlap condition in the simulations; otherwise the practical scope of Theorem 4.3 is unclear.","section":"Assumptions D.1.A and Theorem 4.3"}],"minor_comments":[{"comment":"The text says that the local estimations 'fail in around 5% dataset' under the simulation setting; please define what constitutes failure, for example non-convergence, extreme estimated weights, or undefined estimates.","section":"Section 5 (Simulation)"},{"comment":"Figure 5 labels patterns with six digits (e.g., 111111, 111101) while the analysis reports five predictors; please reconcile the dimensions of the patterns with the variables used.","section":"Section 6 and Figure 5"},{"comment":"The 'Input' line describes {1_{R_i=1_d} Q̂_Pa(r)(L_i)} as input, but these quantities are produced by earlier recursive steps; please rephrase to indicate that they are available from previous stages of the algorithm.","section":"Algorithm 1"},{"comment":"The reference to Dong, Wong, and Chan (2024) is listed as 'Balancing method for non-monotone missing data' without a publication venue; if this is a preprint or dissertation chapter, please indicate its status.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious contribution with a meaningful efficiency claim, but the proof of that claim currently contains a nontrivial gap in Lemma H.8 and an apparent mismatch between the estimator analyzed in Appendix F and the estimator defined by Algorithm 1. These issues are likely fixable, so I recommend major revision rather than rejection. The editor may wish to ask for a complete derivation of the balancing-error bound and a careful rewrite of the decomposition in the proof of Theorem D.4 before further review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this one carefully before you put money on the efficiency claim. The paper is a serious extension of Chen (2022) and the local balancing work: it handles general moment equations, derives the semiparametric efficiency bound under pattern-graph MNAR assumptions, and proposes a sequential balancing estimator designed to avoid the extrapolation and error-accumulation problems of local weighting. The efficiency bound derivation in Appendix C is standard semiparametric theory done carefully for all three mixture-coefficient types, and I think it is correct. The sequential loss is a sensible construction, and the simulation does show the method beating entropy and local balancing in the graphs considered.\n\nNow the soft spots. The proof of Theorem 4.3 depends on Lemma H.8, and the lemma's proof is not a proof. It states an inequality relating the balancing-error term S^r_{θ,3} to the penalized objective, says 'Notice that', and then jumps to the conclusion. That inequality is doing real work: it is exactly how the balancing property gives the op(1) rate. Without a derivation, or a reference to a known result that implies it, the influence-function decomposition in Theorem D.4 is not complete. This is not a minor typo; it is a missing argument at the hinge of the main theorem. I would not call the claim false—the result is plausible and likely true given the structure—but the manuscript as written does not establish it.\n\nTwo smaller issues. The overlap assumption (D.1.A) is bounded propensity odds uniformly; the authors call it standard, and it is, but it is not testable and can easily fail with small complete-case probabilities. That is a limitation to state clearly. And the simulation appears to be nearly correctly specified: the true odds are polynomial degree four and the basis functions are splines up to degree four, so the simulation is closer to an oracle check than to an honest evaluation under misspecification. Also, no code or data is provided, which makes it harder to check the finite-sample behavior. These are fixable in revision.\n\nVerdict: worth refereeing. The efficiency bound alone is a contribution, and the estimator design is interesting enough that the missing lemma deserves a careful look, not a desk reject. Give it to a referee who knows the balancing literature and ask specifically about Lemma H.8.","headline":"Solid extension of pattern-graph missing-data estimation; the efficiency bound looks right, but a real proof gap in Lemma H.8 means the efficiency claim is not yet fully supported.","tokens_in":44076,"tokens_out":2530,"would_cite":true,"duration_ms":22388,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D10","62F12","62G05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that a sequentially balanced weighting estimator is semiparametrically efficient for parameters defined by moment equations under non-monotone, missing-not-at-random patterns.","keywords":["missing not at random","non-monotone missing data","pattern graphs","balancing weights","semiparametric efficiency","weighted estimating equations","propensity odds","covariate balancing"],"falsifier":"Construct a data-generating process satisfying the pattern-graph identification assumptions but with one propensity odds $O_r(l[r])$ unbounded above, e.g., $P(R\\in\\mathrm{Pa}(r)|l[r])$ decaying exponentially in an observed covariate; then run the sequential balancing estimator on many replications and check whether the fitted weights develop heavy tails and whether the coverage of 95% confidence intervals degrades, which would refute the uniformity claim on which efficiency rests.","tokens_in":43090,"feed_emoji":"🎯","tokens_out":7082,"duration_ms":63428,"temperature":0.7,"pith_summary":"The paper addresses estimation of parameters defined by moment conditions, such as regression coefficients, when data have multiple non-monotone missingness patterns and missingness may depend on unobserved values. Its proposal is to reweight complete cases with weights built by sequentially balancing each missing pattern against the complete cases, then solve a weighted estimating equation. The authors derive the semiparametric efficiency bound for the parameter under the identifying assumptions encoded by a regular pattern graph and prove that the sequential balancing estimator attains this bound. As a result, the simple weighted estimator is asymptotically as precise as any regular estimator, including imputation-based or augmented procedures. Simulations illustrate that the sequential weights are more stable and give smaller bias and mean squared error than entropy-based or local balancing weights.","feed_headline":"Balancing weights hit the efficiency bound for missing data","feed_subtitle":"A sequential weighting scheme matches the best possible asymptotic variance under non-monotone, not-at-random missingness.","key_machinery":"The load-bearing device is a recursive factorization of propensity odds along paths in a pattern graph, a directed acyclic graph whose vertices are missingness patterns and whose edges encode the identifying assumption that a pattern's unobserved variables follow the same conditional distribution as a mixture over its parents. For a missingness pattern $r$, the odds of seeing pattern $r$ rather than the complete case is $Q_r(l)=O_r(l[r])\\sum_{s\\in \\mathrm{Pa}(r)} Q_s(l)$, with $Q_{1^d}=1$, where $O_r(l[r])$ is the local odds of pattern $r$ versus its parents given observed variables. The sequential balancing loss $L_r\\{O_r(l[r];\\alpha_r),R\\}=1_{R=1^d}O_r(l[r];\\alpha_r)\\hat Q_{\\mathrm{Pa}(r)}(l)-1_{R=r}\\log O_r(l[r];\\alpha_r)$ is minimized at each pattern in increasing missingness order, which enforces empirical balance between pattern $r$ and reweighted complete cases and controls error accumulation along the factorization. The final weights $\\hat w(L)=\\sum_{r\\in\\mathcal R}\\hat Q_r(L)$ are inserted into the weighted estimating equation.","core_discovery":"Under the identifying assumptions of a regular pattern graph, the paper establishes that the estimator $\\hat\\theta_N$ obtained by solving $\\frac{1}{N}\\sum_{i: R_i=1^d} \\hat w(L_i)\\psi_{\\theta}(L_i)=0$ with $\\hat w(L)=\\sum_{r\\in\\mathcal R}\\hat Q_r(L)$ is consistent for $\\theta_0$ and satisfies $N^{1/2}(\\hat\\theta_N-\\theta_0)\\to_d N(0,D_{\\theta_0}^{-1}V_{\\theta_0}D_{\\theta_0}^{-\\top})$, where $D_{\\theta_0}$ is the derivative of the moment and $V_{\\theta_0}$ is the variance of the efficient influence function. The paper shows this variance coincides with the asymptotic variance bound for all regular estimators derived in Theorem 4.2, so the estimator attains semiparametric efficiency.","pith_inferences":["A natural extension beyond the paper is to causal effect estimation with multiple versions of treatment or noncompliance patterns, since target causal contrasts can be written as moment equations.","The efficiency result implies that when the overlap assumption holds, the choice between imputation and weighted estimation is second-order; practical gains would come from stability and computation, where the sequential weights may dominate.","A diagnostic extension: evaluate the empirical balance condition for each pattern's basis functions; large imbalance under the fitted weights would signal that either the pattern graph or the overlap assumption is untenable.","One could test the overlap assumption indirectly by monitoring the distribution of fitted weights on complete cases; weights diverging toward infinity for a non-negligible fraction of observations would flag a violation."],"forward_implications":["The estimator reaches the same asymptotic variance lower bound as any regular estimator, so under the stated assumptions no imputation or augmented IPW strategy can improve first-order efficiency.","Because the sequential loss uses complete cases directly in every pattern's optimization, the weight estimators avoid extrapolating odds models to the complete-case region, which the paper argues removes a source of instability.","The method extends balancing-weight estimation from the single-parent CCMV assumption to general regular pattern graphs with multiple parents and to generalized mixture coefficients.","Simulations show reductions in bias and mean squared error relative to entropy and local balancing weights, including settings where the identifying pattern graph is misspecified.","A consistent sandwich variance estimator accompanies the estimator, so confidence intervals and hypothesis tests can be constructed in practice."],"supporting_citations":[{"why":"Introduces regular pattern graphs and the recursive odds representation that the paper's identification argument builds on.","marker":"Chen (2022)"},{"why":"Develops balancing weights for the CCMV single-parent case, the starting point the sequential procedure generalizes.","marker":"Dong et al. (2024)"},{"why":"Shows covariate balancing can be achieved by minimizing tailored loss functions, the basis for both local and sequential losses.","marker":"Zhao (2019)"},{"why":"Supplies kernel-based covariate balancing conditions that motivate the balancing equations used here.","marker":"Wong and Chan (2018)"},{"why":"Establishes the stable balancing-weights approach that motivates preference for balanced weights over extreme inverse probabilities.","marker":"Zubizarreta (2015)"},{"why":"Provides the semiparametric efficiency-bound machinery used to derive the asymptotic variance lower bound.","marker":"Newey (1990)"},{"why":"Defines regular estimators and asymptotic efficiency in the sense the paper targets.","marker":"Begun et al. (1983)"},{"why":"Supplies empirical-process and bracketing-entropy tools used in the proof of uniform asymptotic normality.","marker":"Van der Vaart (2000)"}],"fun_headline_variants":["Balancing weights hit the efficiency bound for missing data","Sequential balancing achieves optimal variance under missingness","Pattern graphs plus sequential weighting: efficient estimation","Efficient missing-data estimation via balancing weights"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire consistency and efficiency argument rests on uniform overlap: for each missing pattern the true propensity odds $O_r(l[r])$ must be bounded away from zero and infinity on the support of $l[r]$, and the complete-case probability must also be bounded away from zero, so no pattern pairing becomes arbitrarily rare.","fun_headline_variants_meta":{"raw":{"variants":["Balancing weights hit the efficiency bound for missing data","Sequential balancing achieves optimal variance under missingness","Pattern graphs plus sequential weighting: efficient estimation","Efficient missing-data estimation via balancing weights"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00059,"raw_usage":{"total_tokens":2734,"prompt_tokens":878,"completion_tokens":1856,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":1798}},"tokens_in":494,"tokens_out":1856,"duration_ms":11699,"temperature":1.0,"reasoning_tokens":1798,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:07:17.312032+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a data-generating process satisfying the pattern-graph identification assumptions but with one propensity odds $O_r(l[r])$ unbounded above, e.g., $P(R\\in\\mathrm{Pa}(r)|l[r])$ decaying exponentially in an observed covariate; then run the sequential balancing estimator on many replications and check whether the fitted weights develop heavy tails and whether the coverage of 95% confidence intervals degrades, which would refute the uniformity claim on which efficiency rests.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces regular pattern graphs and the recursive odds representation that the paper's identification argument builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Develops balancing weights for the CCMV single-parent case, the starting point the sequential procedure generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows covariate balancing can be achieved by minimizing tailored loss functions, the basis for both local and sequential losses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies kernel-based covariate balancing conditions that motivate the balancing equations used here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the semiparametric efficiency-bound machinery used to derive the asymptotic variance lower bound."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines regular estimators and asymptotic efficiency in the sense the paper targets."}],"review_version":1}