REVIEW 4 major objections 4 minor 1 cited by
A restricted ANOVA decomposition, called RAY, turns the intractable semiparametric efficient estimating function for nonmonotone missing data into a tractable, multiply robust Z-estimator, with an adaptive variant that attains the best asym
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 14:37 UTC pith:ZIHHK7ZC
load-bearing objection A clever, original decomposition that deserves referee time, but Theorem 3.11 looks mis-specified and the advertised efficiency-gap bound never appears. the 4 major comments →
Towards Efficient Inference under Nonmonotone Missingness with General Imputation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that the RAY approximation in equation (9) is a valid estimating function for the target parameter and yields a computationally tractable estimator that is consistent and asymptotically normal, and that the adaptive estimator IBM(Adaptive) minimizes asymptotic variance over the class (17). Under MCAR, the operator M in the efficient influence function decomposes so that each RAY component is an almost-eigen-function with eigenvalue equal to the proportion of patterns containing that component. Ignoring the remainder terms, which are compositions of conditional expectations, gives a closed-form equation that is multiply robust: it remains unbiased regardless of the qualit
What carries the argument
The RAY decomposition (5) splits any full-data function into signed conditional-expectation terms over observed patterns; this makes the operator M approximately diagonal, with each RAY component scaled by λ_s, the proportion of patterns containing subset s. The resulting closed-form estimating function (9) assigns random weights ω_r (10) to imputed terms from each pattern. The adaptive estimator (17) parameterizes a larger family of such weights and optimizes them under zero-sum constraints by quadratic programming.
Load-bearing premise
The paper discards the remainder terms in the RAY decomposition, arguing they are negligible, but the general bound on the efficiency loss from doing so is not supplied; the argument rests on numerical evidence for a Gaussian example and the conditional-independence special case.
What would settle it
Run the Remark 3.3 Gaussian simulation with exchangeable correlation ρ = -0.4 and compare the variance of the remainder terms Rem_s(f) with the main terms λ_s P_s(f). If, as the paper's Figure 2 hints, the remainder variance is not negligible, the RAY approximation's efficiency loss relative to the semiparametric bound is substantial, and since the paper does not provide the promised general bound, the near-efficiency claim would be unsupported in that regime.
If this is right
- General Z-estimation problems (means, regression coefficients, etc.) under nonmonotone MCAR missingness become computationally tractable.
- Estimators remain consistent even when pre-trained or misspecified imputation models are used, because of multiple robustness.
- IBM(Adaptive) is guaranteed to be asymptotically no worse than the naive complete-case estimator, and beats both IBM(RAY) and IBM(PS) in the class (17).
- RAY weights are numerically more stable than pattern-stratification weights when complete-case proportion is small.
- The framework suggests a route to MAR extensions, though the paper notes this remains challenging.
Where Pith is reading between the lines
- The same almost-eigen decomposition could be applied to other semiparametric operator-inversion problems outside missing data, e.g., causal inference with coarsened data.
- The absent general efficiency-gap bound is the key open step; without it, near-efficiency claims are only numerically supported, not theorem-backed.
- Combining RAY with cross-fitted flexible nuisance models could make the method nonparametric and data-adaptive in high dimensions.
- The quadratic-programming tuning could be adapted to other loss criteria, trading off variance components differently.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies estimation and inference for parameters defined by Z-estimating equations when data exhibit blockwise, non-monotone missingness under a missing-completely-at-random (MCAR) assumption. It introduces the Restricted ANOVA HierarchY (RAY) decomposition, which approximates the semiparametrically efficient observed-data estimating function by separating a tractable leading term from intractable compositions of conditional expectations. The authors define IBM(RAY), IBM(PS), and IBM(Adaptive) estimators, where pre-trained AI models are used to approximate conditional expectations and tuning parameters are chosen to minimize a scalarization of asymptotic variance. They prove unbiasedness of the estimating functions, consistency and asymptotic normality under fixed tunings, and claim that IBM(Adaptive) achieves the lowest asymptotic variance in a class containing RAY and PS estimators. Simulations and a single-cell multi-omics application illustrate efficiency gains.
Significance. The RAY decomposition is an original and potentially valuable tool: it gives a tractable estimating equation in a setting where the efficient influence function is generally intractable, and the unbiasedness argument in Proposition 3.5/3.6 is clean and does not require the imputation functions to be correct. The multiply robust property, in the sense of consistency for arbitrary pre-trained imputation functions, is a genuine strength. If the asymptotic variance formulas and the adaptive optimality claim are corrected, the paper would be a strong contribution to the prediction-powered inference and missing-data literatures. However, as written, several load-bearing variance calculations are incomplete or internally inconsistent, and one advertised contribution (a general efficiency-gap bound) is absent; these issues must be resolved before the central claims can be accepted.
major comments (4)
- [Theorem 3.8] The variance formula for the imputation terms is incomplete when F_r has nonzero mean, which is allowed by Proposition 3.6 and is the intended imputation-mode case. Since E[omega_r]=0 and R is independent of Z, V(sum_r alpha_r omega_r F_r) = sum_{r,r'} alpha_r alpha_{r'} eta_{r,r'} E[F_r F_{r'}^T], not eta_{r,r'} Cov(F_r,F_{r'}). The displayed formula drops the E[F_r]E[F_{r'}]^T term. Consequently the G and L in Corollary 3.9 do not define the minimizer of the asymptotic variance for general AI models, and the efficiency-gain expression l(Sigma_{alpha*}) = l(A^{-1}V(psi^F)A^{-1})/pi_{[M]} - L^T G^{-1}L is not established. The proof in Appendix A.3 simply says covariance terms 'simplify' and does not address the nonzero-mean case. Please either restrict Theorem 3.8 to mean-zero F_r (which excludes biased imputation-mode models) or derive and use the correct formulas.
- [Theorem 3.11] The displayed Sigma_alpha for IBM(Adaptive) is internally inconsistent: the first and third terms are A^{-1}(.)A^{-1}-sandwiched, but the middle cross term sum_r (1/pi_{[M]}) alpha_{r,[M]}[Cov(psi^F,F_r)+Cov(psi^F,F_r)^T] is not. Expanding V(psi_alpha) from (17) shows the cross term must be sandwiched identically. As printed, the quadratic program minimizes a quantity that is not l(Sigma_alpha) for general d, so the claimed optimality of IBM(Adaptive) within class (17)-(18) does not follow. In addition, no proof of Theorem 3.11 is provided in Appendix A (A.5 gives the representation of PS/RAY, A.6 gives an efficiency comparison). This is a load-bearing gap.
- [Abstract; Section 6] The paper's first abstract promises 'a general bound for the efficiency gap otherwise' for RAY and an 'extension of RAY under missing at random.' No such general bound appears in the body; Remark 3.3 and Figure 2 give only a Gaussian numerical check, and Remark 3.4 treats the conditional-independence case where remainders vanish. Section 6 explicitly states that MAR extension is an open future direction. The near-efficiency claim of RAY is therefore supported only by examples, and the advertised bound is absent. Either derive the bound or revise the claims.
- [Theorem 3.8/3.11] The covariance displays use A^{-1} on both sides of the covariance, but Appendix A.3's own Taylor expansion gives asymptotic covariance A^{-1}Sigma(A^T)^{-1}. For non-symmetric A (general Z-estimation), the displayed Sigma_alpha and the definitions of G and L in Corollary 3.9 are incorrect. Replace the right factor by A^{-T} throughout, or state and verify a symmetry assumption on A.
minor comments (4)
- [Front matter] The manuscript contains two different abstracts and titles. The first abstract promises results (efficiency-gap bound, MAR extension) that the second abstract and Section 6 do not claim. Please reconcile them.
- [Appendix A.4] In the proof of Corollary 3.9, the objective is first written with M, then G is defined, and the gradient is written with 2M alpha + 2L. The notation should be consistent; the gradient equation should use G throughout.
- [Remark 3.3] The Gaussian example in Figure 2 lacks simulation details (sample size, number of replications, and how the variance is computed). Please add these details to the caption or remark.
- [Theorem 3.1] The inverse M^{-1} is asserted to be uniquely defined as the Neumann series sum (I-M)^k. This requires conditions for convergence in the relevant operator norm; if only a formal inverse is meant, state that explicitly.
Circularity Check
No circularity found: RAY and IBM estimators are derived from explicit semiparametric algebra and genuine optimization; the main gaps are missing proofs/bound, not circular reasoning.
full rationale
The paper's derivation chain is self-contained. The RAY approximation (9) is obtained by an explicit algebraic decomposition (Lemma 3.2, Eq. 6) and by dropping remainder terms; its unbiasedness (Propositions 3.5 and 3.6) follows from the zero-mean property of the RAY weights ω_r, which is proven directly from the inclusion–exclusion structure (Appendix A.2). The IBM(RAY) and IBM(PS) estimators are special cases of the class (17), and IBM(Adaptive) is defined as the minimizer of a quadratic scalarization of the asymptotic variance over that class (Theorem 3.11, Eq. 18). This is a genuine optimization over an explicitly parameterized class, not a fitted parameter renamed as a prediction. Self-citations (Testa et al. 2025a,b) appear only in remarks and simulation context and are not load-bearing for the central claims. Important weaknesses exist, but they are not circularity: the abstract promises 'a general bound for the efficiency gap otherwise' that never appears; the multiply-robust statement is explicitly said to have its formal statement omitted; and Theorem 3.11's displayed covariance formula appears to drop the A^{-1} sandwich on the cross term, which would undermine the claimed adaptive optimality. These are correctness/completeness concerns, not instances of the derivation reducing to its inputs. No circular step can be exhibited from the paper's equations.
Axiom & Free-Parameter Ledger
free parameters (2)
- alpha_r (RAY tuning weights) =
estimated by minimizing asymptotic variance (Corollary 3.9)
- alpha_{r,s} (Adaptive tuning weights) =
estimated by constrained quadratic programming (Theorem 3.11)
axioms (4)
- domain assumption MCAR: distribution of R independent of the full data
- domain assumption Strict positivity pi_[M] > c > 0
- standard math Standard Z-estimator regularity conditions
- domain assumption Conditional independence X2 ⊥ Y | X1 for efficiency-bound attainment
read the original abstract
Missing data are ubiquitous in classical survey and longitudinal studies as well as modern multi-modality data analysis. A longstanding challenge arises under nonmonotone missingness, where different units may observe arbitrary subsets of all variables. We study parameter estimation and inference problem under this setting. Semiparametric efficiency theory characterizes the efficient estimator through inversion of an operator constructed from pattern-specific conditional expectations. However, this estimator is generally not tractable due to compositions of conditional expectations across patterns. We introduce the Restricted ANOVA hierarchY (RAY), a functional decomposition that reveals an almost-eigen structure of the operator under missing completely at random. This structure yields a closed-form, computable approximation to the efficient estimator. RAY estimator is applicable to general Z-estimation problems, and it remains unbiased for arbitrary independent imputation functions. In theory, we establish verifiable sufficient conditions where RAY attains the efficiency lower bound, and offer a general bound for the efficiency gap otherwise. We further develop adaptive RAY estimator, which attains the minimal asymptotic variance within a broader class containing RAY and other existing estimators. Finally, we investigate the extension of RAY under missing at random mechanism. Simulations and a single-cell multi-omics application demonstrate the efficiency gains of the proposed estimators.
Figures
Forward citations
Cited by 1 Pith paper
-
Semiparametric Mediation Analysis with Separately Observed Mediator and Outcome under Unmeasured Confounding
A semiparametric data fusion framework restores identification of mediation effects from separately observed mediator and outcome data using shared IVs under unmeasured confounding and no-interaction plus latent align...
Reference graph
Works this paper leans on
-
[1]
L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774
Pith/arXiv arXiv 2023
-
[2]
N., Bates, S., Fannjiang, C., Jordan, M
Angelopoulos, A. N., Bates, S., Fannjiang, C., Jordan, M. I., and Zrnic, T. (2023a). Prediction-powered inference. Science , 382(6671):669--674
-
[3]
Angelopoulos, A. N., Duchi, J. C., and Zrnic, T. (2023b). Ppi++: Efficient prediction-powered inference. arXiv preprint arXiv:2311.01453
-
[4]
and SZEGEDY, B
BACKHAUSZ, \'A . and SZEGEDY, B. (2019). On the almost eigenvectors of random regular graphs. The Annals of Probability , 47(3):1677--1725
2019
-
[5]
Chen, H. Y. (2004). Nonparametric and semiparametric models for missing covariates in parametric regression. Journal of the American Statistical Association , 99(468):1176--1189
2004
-
[6]
Chen, X., McCormick, T., Mukherjee, B., and Wu, Z. (2025). A unified framework for inference with general missingness patterns and machine learning imputation. arXiv preprint arXiv:2508.15162
arXiv 2025
-
[7]
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters
2018
-
[8]
H., and Jewell, N
Das, M., Kennedy, E. H., and Jewell, N. P. (2024). Doubly robust capture-recapture methods for estimating population size. Journal of the American Statistical Association , 119(546):1309--1321
2024
-
[9]
Diakit \'e , A. O., Moreau, C., Bezgin, G., Bhagwat, N., Rosa-Neto, P., Poline, J.-B., Girard, S., Barry, A., Initiative, A. D. N., et al. (2025). Adapdiscom: An adaptive sparse regression method for high-dimensional multimodal data with block-wise missingness and measurement errors. arXiv preprint arXiv:2508.00120
arXiv 2025
-
[10]
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Muller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., Podell, D., Dockhorn, T., English, Z., Lacey, K., Goodwin, A., Marek, Y., and Rombach, R. (2024). Scaling rectified flow transformers for high-resolution image synthesis. ArXiv , abs/2403.03206
Pith/arXiv arXiv 2024
-
[11]
Gronsbell, J., Gao, J., Shi, Y., McCaw, Z. R., and Cheng, D. (2024). Another look at inference after prediction. arXiv preprint arXiv:2411.19908
arXiv 2024
-
[12]
Hoeffding, W. (1948). A Class of Statistics with Asymptotically Normal Distribution . The Annals of Mathematical Statistics , 19(3):293 -- 325
1948
-
[13]
Hooker, G. (2004). Discovering additive structure in black box functions. In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining , pages 575--580
2004
-
[14]
Horn, R. A. and Johnson, C. R. (2012). Matrix analysis . Cambridge university press
2012
-
[15]
Huang, J., Wang, H., Lei, Y., and Chen, Y. (2025). Efficient semiparametric inference for distributed data with blockwise missingness. arXiv preprint arXiv:2508.16902
Pith/arXiv arXiv 2025
-
[16]
Ji, W., Lei, L., and Zrnic, T. (2025). Predictions as surrogates: Revisiting surrogate outcomes in the age of ai. arXiv preprint arXiv:2501.09731
Pith/arXiv arXiv 2025
-
[17]
and Rothenh \"a usler, D
Jin, Y. and Rothenh \"a usler, D. (2023). Modular regression: improving linear models by incorporating auxiliary data. Journal of Machine Learning Research , 24(351):1--52
2023
-
[18]
Kundu, P., Tang, R., and Chatterjee, N. (2019). Generalized meta-analysis for multiple regression models across studies with disparate covariate information. Biometrika , 106(3):567--585
2019
-
[19]
Li, Y., Yang, X., Wei, Y., and Liu, M. (2024). Adaptive and efficient learning with blockwise missing and semi-supervised data. arXiv preprint arXiv:2405.18722
Pith/arXiv arXiv 2024
-
[20]
Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al. (2024). Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437
Pith/arXiv arXiv 2024
-
[21]
and Ulam, S
Metropolis, N. and Ulam, S. (1949). The monte carlo method. Journal of the American statistical association , 44(247):335--341
1949
-
[22]
Miao, J., Miao, X., Wu, Y., Zhao, J., and Lu, Q. (2023). Assumption-lean and data-adaptive post-prediction inference. arXiv preprint arXiv:2311.14220
Pith/arXiv arXiv 2023
-
[23]
P., Lareau, C
Mimitou, E. P., Lareau, C. A., Chen, K. Y., Zorzetto-Fernandes, A. L., Hao, Y., Takeshima, Y., Luo, W., Huang, T.-S., Yeung, B. Z., Papalexi, E., et al. (2021). Scalable, multimodal profiling of chromatin accessibility, gene expression and protein levels in single cells. Nature biotechnology , 39(10):1246--1258
2021
-
[24]
F., Chakraborti, T., Holmes, C., Copping, R., Hagenbuch, N., Biedermann, S., Noonan, J., Lehmann, B., Shenvi, A., et al
Mitra, R., McGough, S. F., Chakraborti, T., Holmes, C., Copping, R., Hagenbuch, N., Biedermann, S., Noonan, J., Lehmann, B., Shenvi, A., et al. (2023). Learning from data with structured missingness. Nature Machine Intelligence , 5(1):13--23
2023
-
[25]
and Witten, D
Motwani, K. and Witten, D. (2023). Revisiting inference after prediction. J. Mach. Learn. Res. , 24:394:1--394:18
2023
-
[26]
Owen, A. B. (2013). Monte carlo theory, methods and examples
2013
-
[27]
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. ArXiv , abs/2204.06125
Pith/arXiv arXiv 2022
-
[28]
Robins, J. M. (1997). Non-response models for the analysis of non-monotone non-ignorable missing data. Statistics in medicine , 16(1):21--37
1997
-
[29]
Robins, J. M. and Gill, R. D. (1997). Non-response models for the analysis of non-monotone ignorable missing data. Statistics in medicine , 16(1):39--56
1997
-
[30]
M., Rotnitzky, A., and and, L
Robins, J. M., Rotnitzky, A., and and, L. P. Z. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American Statistical Association , 89(427):846--866
1994
-
[31]
Sohn, K., Lee, H., and Yan, X. (2015). Learning structured output representation using deep conditional generative models. Advances in neural information processing systems , 28
2015
-
[32]
Song, S., Lin, Y., and Zhou, Y. (2024). Semi-supervised inference for block-wise missing data without imputation. Journal of Machine Learning Research , 25(99):1--36
2024
-
[33]
and Tchetgen Tchetgen, E
Sun, B. and Tchetgen Tchetgen, E. J. (2018). On inverse probability weighting for nonmonotone missing at random data. Journal of the American Statistical Association , 113(521):369--379
2018
-
[34]
M., Hauth, A., Millican, K., et al
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al. (2023). Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805
Pith/arXiv arXiv 2023
-
[35]
Testa, L., Boschi, T., Chiaromonte, F., Kennedy, E. H., and Reimherr, M. (2025a). Doubly-robust functional average treatment effect estimation. arXiv preprint arXiv:2501.06024
-
[36]
Testa, L., Xu, Q., Lei, J., and Roeder, K. (2025b). Semiparametric semi-supervised learning for general targets under distribution shift and decaying overlap. arXiv preprint arXiv:2505.06452
-
[37]
Tsiatis, A. A. (2006). Semiparametric theory and missing data , volume 4. Springer
2006
-
[38]
Van der Vaart, A. W. (2000). Asymptotic statistics , volume 3. Cambridge university press
2000
-
[39]
H., and Leek, J
Wang, S., McCormick, T. H., and Leek, J. T. (2020). Methods for correcting inference based on outcomes predicted by machine learning. Proceedings of the National Academy of Sciences , 117(48):30266--30275
2020
-
[40]
Xu, Z., Witten, D., and Shojaie, A. (2025). A unified framework for semiparametrically efficient semi-supervised learning. arXiv preprint arXiv:2502.17741
Pith/arXiv arXiv 2025
-
[41]
Xue, F., Ma, R., and Li, H. (2021). Statistical inference for high-dimensional linear regression with blockwise missing data. arXiv preprint arXiv:2106.03344
Pith/arXiv arXiv 2021
-
[42]
and Qu, A
Xue, F. and Qu, A. (2021). Integrating multisource block-wise missing data in model selection. Journal of the American Statistical Association , 116(536):1914--1927
2021
-
[43]
Yang, S. (2022). A double robust approach for non-monotone missingness in multi-stage data. arXiv preprint arXiv:2201.01010
Pith/arXiv arXiv 2022
-
[44]
Ying, C., Jin, J., Guo, Y., Li, X., Liang, M., and Zhao, J. (2025). Towards the efficient inference by incorporating automated computational phenotypes under covariate shift. arXiv preprint arXiv:2505.22632
Pith/arXiv arXiv 2025
-
[45]
Yu, G., Li, Q., Shen, D., and Liu, Y. (2020). Optimal sparse linear prediction for block-missing multi-modality data without imputation. Journal of the American Statistical Association , 115(531):1406--1419
2020
-
[46]
Zhang, Y., Chakrabortty, A., and Bradic, J. (2023). Double robust semi-supervised inference for the mean: selection bias under mar labeling with decaying overlap. Information and Inference: A Journal of the IMA , 12(3):2066--2159
2023
-
[47]
Zhao, S. and Candès, E. (2025). Imputation-powered inference. arXiv preprint arXiv: 2509.13778
arXiv 2025
-
[48]
and Cand \`e s, E
Zrnic, T. and Cand \`e s, E. J. (2024). Cross-prediction-powered inference. Proceedings of the National Academy of Sciences , 121(15):e2322083121
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.