Pith. sign in

REVIEW 1 major objections 4 minor 1 cited by

Asymptotics of Nonparametric Estimation under general non-monotone MAR missingness: A Bayesian Approach

T0 review · 1 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read This paper proves that under general non-monotone missing-at-random missingness, a Bayesian Dirichlet-mixture-of-normals posterior can estimate the complete-data density at the same minimax rate as if the data were fully observed, up to log

desk verdict First nonparametric Bayesian contraction under general non-monotone MAR; rate is plausible, but the support condition in Thm 5.1 needs to be made explicit. read the letter →

arxiv 2603.23449 v2 pith:IMNVX7K6 submitted 2026-03-24 math.ST stat.TH

classification math.STstat.TH MSC 62G2062C1062G07
keywords missingatrandomnon-monotonemissingnessposteriorcontractionBayesiannonparametricsdensityestimationDirichletmixtureofnormalspropensityscoreminimaxrate
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper targets a long-standing gap: under general, non-monotone missing-at-random (MAR) missingness, no nonparametric consistency result for the complete-data distribution existed, and common M-estimators can be inconsistent. The authors establish that Bayesian posterior inference built on the ignorable likelihood remains valid: the posterior contracts around the true complete-data density at the same rate as if the data were fully observed, up to logarithmic factors. The key adaptation is a missingness-aware Kullback-Leibler divergence that replaces the usual KL divergence in the prior-mass condition, and a constructive test showing that Hellinger testing still works automatically under MAR when the probability of a fully observed row is bounded away from zero. In the density-estimation application with a Dirichlet mixture of normals, the contraction rate matches the minimax rate up to logs. A practical consequence is an algorithm that takes incomplete rows and returns samples from a consistent estimate of the unmasked distribution.

What carries the argument

Two objects carry the argument. First, the relative KL divergence fKL(Pθ*∥Pθ) = E_{(X,M)~Pθ*} log[pθ*^{(M)}(X^{(M)})/pθ^{(M)}(X^{(M)})], which is nonnegative under MAR and has a unique zero exactly when the complete-data densities coincide, provided the propensity score is positive. It replaces the ordinary KL in the prior-mass part of the contraction theorem. Second, the pattern-weighted Hellinger distance eH², which sums squared differences of observed marginal densities weighted by P(M=m|x^{(m)}); MAR lets it be sandwiched as δ H² ≤ eH² ≤ H², where δ is the propensity lower bound. This sandwich yields uniform tests of power with type-I error e^{-n δ ε²/8}, recovering the classical automat

What would settle it

Take any MAR mechanism with P(M=0|X=x)=0 on a set A of positive probability, and two densities that agree on every observed marginal but differ on A. Simulate data from such a mechanism and check whether the posterior of the proposed Dirichlet-mixture sampler fails to contract to the true density in Hellinger distance — or, equivalently, whether two different complete-data densities produce the same observable likelihood. If the posterior still concentrates near the true density, the identifiability claim would be contradicted; if it drifts, the positivity condition is confirmed as load-bearin

Watch

Extended reading notes

Core claim

The central claim is that, contrary to what the fragmented non-monotone MAR literature suggested, the missingness mechanism does not change the minimax exponent for density estimation: with a positive propensity score, the posterior based only on the observed fragments concentrates in Hellinger distance at n^{-β/(2β+d)} up to log factors. This is achieved by showing that the two classical ingredients of Bayesian nonparametric contraction — prior mass in a KL-type neighbourhood and existence of uniform tests — have missing-data analogues that are no harder to satisfy than their complete-data versions. The prior-mass condition is expressed through a new relative KL divergence that measures how

Load-bearing premise

The complete-data density is recoverable from the observed fragments only because the probability of seeing a fully observed row is assumed to be bounded away from zero at every point; if that probability can be zero on some region, two distributions that differ only there are indistinguishable from the observed data.

Editorial extensions

If this is right

  • The complete-data density can be estimated consistently at the minimax rate up to log factors under general non-monotone MAR, so any downstream estimate built from the density inherits this rate.
  • The ignorable Bayesian likelihood — which ignores the missingness mechanism — is frequentist-valid in a nonparametric sense whenever the propensity score is bounded away from zero.
  • The proposed algorithm returns posterior samples from a consistent estimate of the unmasked distribution, giving a principled alternative to ad hoc imputation for distributional questions.
  • The general contraction theorem does not depend on density estimation specifically, so the same framework can be reused for other nonparametric models under MAR.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A sharp experiment suggested by the paper: run the proposed estimator under a MAR mechanism whose propensity score is zero on a region of positive mass; the posterior should drift to an indistinguishable alternative, confirming that the positivity condition is not merely technical.
  • The paper itself notes it is unclear whether a prior that satisfies the prior-mass condition for complete data automatically satisfies the adapted prior-mass condition for missing data; verifying this on a case-by-case basis is a natural next step.
  • Because the estimator generates fresh samples from the learned mask-free density rather than imputing missing entries, it reframes imputation as a distributional task — whole-distribution distances, not cell-wise accuracy, become the right benchmark.
  • A semiparametric Bernstein–von Mises result under MAR, which the paper flags for future work, would add confidence intervals to everything estimated from the recovered density.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. The paper develops a general Bayesian nonparametric posterior contraction theory for data with general non-monotone missing-at-random (MAR) mechanisms. It introduces a missing-data-adapted Kullback-Leibler divergence fKL and an adapted Hellinger distance eH, proves that the Hellinger testing condition holds under MAR plus a positivity condition on the propensity score, states a general rate theorem, and specializes it to density estimation with a Dirichlet-mixture-of-normals prior. The advertised main result is that the complete-data density can be estimated at the minimax rate up to logarithmic factors under Hölder smoothness. The paper also provides an MCMC algorithm and simulation evidence.

Significance. If the main theorem were correct as stated, this would be the first nonparametric posterior-contraction result for general non-monotone MAR, and the testing construction is a genuinely interesting contribution. The paper is mostly rigorous, with detailed proofs and an implemented algorithm; the simulation study is honest about settings that violate the assumptions. However, as explained in the major comment, Theorem 5.1 is not proved as stated because the proof relies on a support condition that is absent from the theorem statement and is not satisfied by the Dirichlet-mixture prior for a natural class of Hölder densities. The central idea is defensible, but the main density-estimation claim needs to be corrected or qualified.

major comments (1)
  1. [§5.2, Theorem 5.1 and its proof] The proof of Theorem 5.1 rests on Theorem C.2, Corollary C.3, and Lemma C.4, all of which require the additional condition that there exists θ⋆ such that p_θ⋆ is continuous and supp(p_θ) ⊂ supp(p_θ⋆) for every θ in the model. In particular, the construction q_θj(x,m) = p^(m)_θj(x^(m)) p_θ⋆(x^(-m)|x^(m)) P(M=m|x) in Eq. (31) is only a probability density when supp(p_θj) ⊆ supp(p_θ⋆). The Dirichlet-mixture-of-normals prior used in Theorem 5.1 assigns positive mass only to densities with full support R^d. Thus if p_θ⋆ is compactly supported — which Assumption 5.1 explicitly permits — the support condition fails for every θ in the prior support. Consequently, Theorem 4.4 cannot supply the testing condition and Lemma C.4 cannot be used to verify the prior-mass condition Assumption 4.1(i). The theorem as stated is therefore unproved. This is load-bearing: it invalidates the claimed minimax-rat
minor comments (4)
  1. [Appendix B, Proposition B.1] In the first inequality of the proof, 'dx^(m)' should be 'dx' (or 'dx^(0)'); the displayed lower bound integrates over the full vector x, not the observed subvector.
  2. [Appendix C, Corollary C.3] Near the end of the proof, the expression '-δε²/6 + ε²δ/24' is written as '= δε²/8'; it should be '-δε²/8'. The intended inequality is clear, but the sign typo is confusing.
  3. [Theorem 5.1 statement] The exponent t is given as 't > βd+βd/τ+d+β / 2β+d 2', which is hard to parse. It should be written as t > 2(βd+βd/τ+d+β)/(2β+d); also 'desribed' is a typo.
  4. [Theorem 4.3 proof] In the last display of the proof, the factor e^{-n\barε_n^2(2+C1)} appears; the preceding steps have e^{+n\barε_n^2(2+C1)}. The convergence conclusion is unaffected because ε_n ≥ \barε_n, but the sign should be corrected for clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity found; the contraction proof rests on external Ghosal–van der Vaart theory and no fitted quantity is relabeled as a prediction.

full rationale

The central rate theorem (Theorem 5.1) is derived by verifying the three components of the prior-mass-and-testing framework (Assumption 4.1 and the new Hellinger testing result, Theorem 4.4), and the verification uses external, parameter-free bounds from Ghosal and van der Vaart (2017) for Dirichlet-mixture priors and inverse-Wishart priors. The MAR-specific objects fKL, eV, and eH are internally defined distances used to reformulate the prior-mass and testing conditions; they are not fitted outputs that are then fed back into the derivation. Proposition 4.2 only establishes identifiability of the true density from observed fragments under the positivity assumption, not a rate result. Self-citations to Näf et al. (2026) and Chérief-Abdellatif and Näf (2025) are motivational or contextual and do not carry the proof of Theorem 5.1. The reviewer-flagged support-condition gap is real and non-circular: Theorem 5.1 omits the supp(pθ)⊂supp(pθ⋆) continuity condition that Theorem 4.4, Corollary C.3, Lemma C.4, and the construction qθj(x,m)=p^{(m)}_{θj}(x^{(m)})p_{θ⋆}(x^{(-m)}|x^{(m)})P(M=m|x) require, so the stated Hölder-class result is not fully proved for densities with zeros or compact support. This is a correctness/assumption gap, not a circular reduction: the conclusion is not equivalent to an input by construction, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

No fitted free parameters appear in the theory; δ is an assumed lower bound, not a fitted constant. The fKL divergence and adapted Hellinger eH are new mathematical definitions but serve internal roles and are not postulated empirical entities. The central derivation uses standard Bayesian nonparametrics machinery from Ghosal-van der Vaart as an external benchmark.

assumptions (7)
  • domain assumption Rubin MAR: P(M=m|X=x)=P(M=m|X^{(m)}=x^{(m)}) for all m and almost all x (Assumption 2.1).
    Defines the missingness regime under which all results hold; without it the ignorable likelihood factorization and fKL properties break down.
  • domain assumption Empty pattern has zero probability: P(M=1)=0 (Assumption 2.2).
    Excludes completely unobserved records; used in the fKL proofs and pattern sums.
  • domain assumption Positivity: P(M=0|X=x)>δ>0 for almost all x (Assumption 2.3).
    Load-bearing: gives identifiability (Prop 4.2), the lower bound δH²≤eH², and the test exponent δ/8. Without it the complete-data density may be non-identifiable.
  • domain assumption Well-specified model: P_θ* belongs to the model {P_θ}, and the missingness mechanism is fixed but unknown.
    Posterior contraction is stated under a true θ*; misspecification of the complete-data model is not covered.
  • domain assumption Support and regularity: there exists θ⋆ such that p_θ⋆ is continuous and supp(p_θ)⊂supp(p_θ⋆) for all θ (Theorem 4.4).
    Required for constructing the modified densities ¯q in Theorem C.2 and for Lemma C.4. In the density application it is implicit in Assumption 5.1(9.10), but not stated in Theorem 5.1.
  • domain assumption Hölder class Assumption 5.1 with mixed derivatives, integrability of L/p, and exponential tail.
    Defines the target class and drives the minimax rate n^{-β/(2β+d)}; inherited from Ghosal-van der Vaart Theorem 9.9.
  • standard math Classical prior-mass and testing framework of Ghosal and van der Vaart (2017): sieve entropy bounds, Dirichlet-process/inverse-Wishart prior-mass bounds, and i.i.d. testing lemmas.
    Background results imported without full re-derivation, especially in the proof of Theorem 5.1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Asymptotics of Nonparametric Estimation under general non-monotone MAR missingness: A Bayesian Approach." pith.science (2026). https://pith.science/paper/IMNVX7K6

@misc{pith2026260323449,
  author       = {Pith},
  title        = {Pith review of: Asymptotics of Nonparametric Estimation under general non-monotone MAR missingness: A Bayesian Approach},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IMNVX7K6}},
  note         = {Machine review of arXiv:2603.23449}
}
read the original abstract

Missing values are ubiquitous in statistical practice, with potentially detrimental consequences for any statistical analysis. As such, a wealth of methods and theoretical results have been developed in the last decades. However, many questions remain open, in particular in the case of general non-monotone missing at random (MAR), where nonparametric results are still lacking. In this paper, we extend nonparametric Bayesian theory to this MAR setting. We introduce a general theorem of posterior contraction under MAR and an additional positivity condition and apply this result to density estimation as well as regression problems. In particular, we show that, despite the missing values, the complete-data density can be estimated with the minimax posterior contraction rate up to logarithmic factors. To the best of our knowledge, this is the first nonparametric result showing that the complete-data distribution can be consistently estimated under Rubin's MAR definition. As a consequence, we obtain an algorithm that takes incomplete data and returns a sample from a consistent estimate of the complete-data distribution.

Figures

Figures reproduced from arXiv: 2603.23449 by the authors.

Figure 1
Figure 1. Three Data matrices with missing values, each with three different patterns. Each [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Mean of X1 (Left) and correlation of X1, X2 (Right) estimation under the MAR mecha￾nism in (3). reconstructing the original distribution, which is not visible in case of the mean, but gets apparent when trying to assess the correlation parameter. Since both mice rf and missForest are rather ad hoc methods based on the multiple imputa￾tion by chained equation (mice, Van Buuren and Groothuis-Oudshoorn (2011)) idea, ne… view at source ↗
Figure 3
Figure 3. Example with Pθ ∗ chosen to be N (0, Σ). Top: Quantile Estimate of X1 for n = 500 (left) and n = 1000 (right), Bottom: negative energy distance between the newly generated sam￾ple/imputation and the full data for n = 500 (left) and n = 1000 (right). 16 [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Example with Pθ ∗ chosen to be a mixture of Gaussians with different means and a correlation of 0.7 between X1 and X2. Top: Quantile Estimate of X1 for n = 500 (left) and n = 1000 (right), Bottom: negative energy distance between the newly generated sample/imputation a…
Figure 5
Figure 5. Figure 5: Example with Pθ ∗ chosen to be uniform on [0, 1]3 with correlation between X1 and X2 induced by a copula. Left: Quantile Estimate of X1, Right: negative energy distance between the newly generated sample/imputation and the full data. We used n = 1000. 18 [PITH_FULL_IM…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach

    stat.ME 2026-08 conditional novelty 8.0 of 10

    A sieve maximum likelihood estimator attains near-minimax Hellinger rates for density estimation under general non-monotone missing at random, with missingness affecting only the constant.

Reference graph

Works this paper leans on

28 extracted references · 3 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Castillo, I. (2024). Bayesian nonparametric statistics, st-flour lecture notes.arXiv preprint arXiv:2402.16422

  2. [2]

    and Yu, C

    Chen, S. and Yu, C. L. (2016). Parameter estimation through semiparametric quantile regression imputation. Electronic Journal of Statistics, 10:3621–3647

  3. [3]

    and Sadinle, M

    Chen, Y.-C. and Sadinle, M. (2019). Nonparametric pattern-mixture models for inference with missing data. arXiv preprint arXiv:1904.11085. Ch´ erief-Abdellatif, B.-E. and N¨ af, J. (2025). Parametric MMD estimation with missing values: Robustness to missingness and data model misspecification.arXiv preprint arXiv:2503.00448. Daniel Malinsky, I. S. and Tch...

  4. [4]

    Frahm, G., Nordhausen, K., and Oja, H. (2020). M-estimation with incomplete and dependent multivariate data.Journal of Multivariate Analysis, 176:104569

  5. [5]

    and van der Vaart, A

    Ghosal, S. and van der Vaart, A. (2017).Fundamentals of Nonparametric Bayesian Inference. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press

  6. [6]

    Grzesiak, K., Muller, C., Josse, J., and N¨ af, J. (2025). Do we need dozens of methods for real world missing value imputation?arXiv preprint arXiv:2511.04833

  7. [7]

    and Yang, S

    Guan, Q. and Yang, S. (2024). A unified inference framework for multiple imputation using martingales. Statistica Sinica, 34:1649–1673

  8. [8]

    Little, R. J. A. and Rubin, D. B. (2019).Statistical Analysis with Missing Data., 3rd Edition. Wiley

Show all 28 references
  1. [9]

    and Fan, Y

    Liu, Y. and Fan, Y. (2023). Biased-sample empirical likelihood weighting for missing data problems: an alternative to inverse probability weighting.Journal of the Royal Statistical Society Series B: Statistical Methodology, 85(1):67–83

  2. [10]

    A., Berrett, T

    Ma, T., Verchand, K. A., Berrett, T. B., Wang, T., and Samworth, R. J. (2024). Estimation beyond missing (completely) at random.Preprint arXiv:2410.10704

  3. [11]

    and Frellsen, J

    Mattei, P.-A. and Frellsen, J. (2019). MIW AE: Deep generative modelling and imputation of incomplete data sets. InProceedings of the 36th International Conference on Machine Learning, volume 97, pages 4413–4423

  4. [12]

    and Rubin, D

    Mealli, F. and Rubin, D. B. (2015). Clarifying missing at random and related definitions, and implications when coupled with exchangeability.Biometrika, 102(4):995–1000

  5. [13]

    Murphy, K. P. (2012).Machine Learning: A Probabilistic Perspective. The MIT Press

  6. [14]

    Neal, R. M. (2000). Markov chain sampling methods for dirichlet process mixture models.Journal of Computational and Graphical Statistics, 9(2):249–265. N¨ af, J. (2026). A practical guide to modern imputation.arXiv preprint arXiv:2601.14796. N¨ af, J., Scornet, E., and Josse, ...

  7. [15]

    Pinheiro, J. C. and Bates, D. M. (1996). Unconstrained parametrizations for variance-covariance matrices. Statistics and Computing, 6(3):289–296

  8. [16]

    Qin, J., Zhang, B., and Leung, D. H. Y. (2009). Empirical likelihood in missing data problems.Journal of the American Statistical Association, 104(488):1492–1503. 42

  9. [17]

    Rubin, D. B. (1976). Inference and missing data.Biometrika, 63(3):581–592

  10. [18]

    Missing at Random

    Seaman, S., Galati, J., Jackson, D., and Carlin, J. (2013). What is meant by “Missing at Random”? Statistical Science, 28(2):257–268

  11. [19]

    Seaman, S. R. and Vansteelandt, S. (2018). Introduction to double robust methods for incomplete data.Stat Sci, 33(2):184–197

  12. [20]

    Smith, R. L. (1992). Some interlacing properties of the schur complement of a hermitian matrix.Linear Algebra and its Applications, 177:137–144

  13. [21]

    Stekhoven, D. J. and B¨ uhlmann, P. (2011). MissForest—non-parametric missing value imputation for mixed- type data.Bioinformatics, 28(1):112–118

  14. [22]

    and Tchetgen, E

    Sun, B. and Tchetgen, E. J. T. (2018). On inverse probability weighting for nonmonotone missing at random data.Journal of the American Statistical Association, 113(521):369–379. Sz´ ekely, G. J. (2003). E-statistics: the energy of statistical samples. Technical Report 05, Bowl...

  15. [23]

    and Kano, Y

    Takai, K. and Kano, Y. (2013). Asymptotic inference with incomplete data.Communications in Statistics - Theory and Methods, 42(17):3174–3190. Van Buuren, S. (2018).Flexible Imputation of Missing Data. Second Edition. Chapman & Hall/CRC Press. Van Buuren, S. and Groothuis-Oudsh...

  16. [24]

    and Robins, J

    Wang, N. and Robins, J. M. (1998). Large-sample theory for parametric multiple imputation procedures. Biometrika, 85(4):935–948

  17. [25]

    and Rao, J

    Wang, Q. and Rao, J. N. K. (2002). Empirical likelihood-based inference under imputation for missing response data.The Annals of Statistics, 30(3):896–924

  18. [26]

    Yu, J., Ying, Q., Wang, L., Jiang, Z., and Liu, S. (2025). Missing data imputation by reducing mutual information with rectified flows.arXiv preprint arXiv:2505.11749

  19. [27]

    and Dong, X

    Yuan, X. and Dong, X. (2019). Weighted empirical likelihood for quantile regression with non ignorable missing covariates.Communications in Statistics - Theory and Methods, 48(12):3068–3084

  20. [28]

    and Cand` es, E

    Zhao, S. and Cand` es, E. (2025). Imputation-powered inference.arXiv preprint arXiv:2509.13778. 43

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.