{"id":"45251857-fe53-4987-bb93-b32d161c813c","arxiv_id":"2608.10113","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A sieve maximum likelihood estimator attains near-minimax Hellinger rates for density estimation under general non-monotone missing at random, with missingness affecting only the constant.","lead":"This paper shows that a likelihood-based density estimator can recover a full data distribution at nearly optimal statistical rates even when values are missing in arbitrary patterns, as long as missingness depends only on observed values and every data point has a chance of being fully seen. It provides the first frequentist guarantees for nonparametric estimation under general non-monotone missing at random, backed by a practical EM algorithm.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Assumption 2.3 (uniform positivity of the propensity score) is the load-bearing condition; the rate and even identification rely on it, and MAR alone does not imply it.","rationale":"The reader's weakest_assumption correctly identifies Assumption 2.3. My independent check of the proof confirms that the uniform lower bound c0>0 is used in Proposition 4.3 to convert the adapted-Hellinger rate from Theorem 4.4 into a full-data Hellinger rate with factor 1/sqrt(c0), and that the pointwise positivity version is used in Proposition 4.1 for identifiability. Both uses are explicit and the associated claims are correct; no hidden dependence or circular step was found. The main mathematical structure of Theorems 4.4 and 5.1 appears internally consistent: the entropy transfer in Lemma B.3, the Bernstein-norm argument in Proposition B.2, and the Gaussian-mixture sieve verification in the proof of Theorem 5.1 all check out at the level of the stated assumptions. The simulations support the practical behavior, though they do not validate the logarithmic factor or the 1/sqrt(c0) constant. The reliance on the companion paper and on Kaji (2026) is a legitimate concern for verification, but the parts cited from the companion are short and independently derivable; the lower bound c0 H^2 <= eH^2 follows by retaining only the m=0 term. Therefore I do not see a flaw that would change the conditional verdict: the paper is correct conditional on Assumption 2.3, and that assumption is indeed the weakest load-bearing point, exactly as the reader stated. My recommendation remains UNCHANGED, meaning the reader's CONDITIONAL verdict should stand.","tokens_in":48458,"tokens_out":28102,"duration_ms":271802,"concrete_test":"Construct a bivariate design satisfying Assumptions 2.1 and 2.2 but violating Assumption 2.3: let the fully observed pattern have propensity w00(x1,x2) = 1/(1 + x1^2 + x2^2), which is positive everywhere but has infimum 0. Choose two densities p0 and p1 that differ only on a set where w00 is extremely small, while all observed-block marginals p^(m) coincide. Compute fKL(p0||p1) numerically and verify it is 0 while H(p0,p1)>0, then run the proposed EM sieve estimator on data from this mechanism and check that the Hellinger error to p0 does not vanish as n grows. If the error remains bounded away from zero, this confirms that positivity, not merely MAR, is the boundary of the stated result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central rate result, Theorem 5.1, is conditional on Assumption 2.3, and this assumption does real work in two distinct places. First, Proposition 4.1's identifiability statement fKL(P0||P)=0 iff P=P0 uses the pointwise version P(M=0|X=x)>0; without it, one can construct P1 with KL(P0||P1)>0 but fKL(P0||P1)=0, so the observed-data likelihood does not identify the complete-data density. Second, the rate factor 1/sqrt(c0) enters through Proposition 4.3's lower bound c0 H^2(p1,p2) <= eH^2(p1,p2), which is obtained by keeping only the fully observed pattern m=0 in Definition 4.2. If c0=0, eH degenerates on the set where the fully observed pattern has zero probability, the Hellinger rate cannot be recovered from the eH rate, and the theorem's bound diverges. The paper is transparent about this: it flags in Section 4 and Remark 1 that positivity is needed beyond MAR itself. The concern is therefore not an internal inconsistency but a scope limitation: the title and abstract emphasize general non-monotone MAR, yet the guarantees require a strong positivity condition that is not implied by MAR and can be arbitrarily small in applications. The companion-paper reliance, while a verification risk, does not appear to hide a mathematical error: the cited bounds have short direct derivations, and the lower bound above is immediate. As long as Assumption 2.3 is stated as a premise, the proof chain appears coherent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a frequentist sieve maximum likelihood theory for estimation under non-monotone missing at random (MAR) missingness. It introduces an adapted Kullback-Leibler divergence and an adapted Hellinger functional on observed-data marginals, and proves (Theorem 4.4) that an approximate maximizer of the observed-data likelihood over any sieve satisfying bracketing entropy, approximation, and moment conditions converges at rate eH(p0, p_hat n)=O_P(delta_n or n^{-1/2}), with H-rate degraded by a 1/sqrt(c0) factor under a uniform positivity assumption on the propensity score. Applying this to a Gaussian mixture sieve under a Holder smoothness class with exponential tails (Theorem 5.1), it obtains H(p0, p_hat n)=O_P(C4/sqrt(c0) n^{-beta/(2beta+d)} (log n)^t), the minimax rate up to logarithmic factors, with missingness affecting only the constant. The estimator is implemented by a MAR-specific EM algorithm with BIC/AIC selection; simulations in d=20 compare it with complete-data Gaussian KDE. The proofs are detailed in an appendix, although two key identifiability and rate-bridge propositions are cited from an unpublished companion paper.","tokens_in":48760,"tokens_out":15108,"duration_ms":141098,"significance":"If correct, this is a substantial contribution: it provides the first general frequentist rate guarantees for nonparametric likelihood-based estimation under non-monotone MAR, and it removes the bounded-density-ratio condition via Kaji's Hellinger bounds. The adapted-divergence machinery is clean, and the adapted Hellinger functional is well suited to pattern-marginal data. The result that missingness changes only the constant in the minimax rate is striking and is stated with transparent assumptions. The authors are also candid about the role of the positivity assumption, and the new entropy localization arguments in Lemmas B.6-B.9 are worked out in detail. Reproducible code is provided. The significance is moderated, however, by the reliance on a strong uniform positivity condition, by the restriction to Holder classes with strong tail and moment conditions and known smoothness, and by the fact that two load-bearing results (Propositions 4.1 and 4.3) are borrowed from an unpublished companion preprint rather than proved in the manuscript.","major_comments":[{"comment":"The identifiability of the adapted KL divergence (Proposition 4.1) and the two-sided bound c0 H^2(p1,p2) <= eH^2(p1,p2) <= H^2(p1,p2) (Proposition 4.3) are both cited from the companion preprint Naf and Cherief-Abdellatif (2026), which is not peer-reviewed. These propositions are load-bearing: Proposition 4.1 justifies the target of the sieve MLE, and Proposition 4.3 is used in Lemma B.3 and in the final transition from eH to H in Theorems 4.4 and 5.1. For a self-contained journal submission, the full proofs of these propositions should be included in the appendix or derived in the paper, rather than referenced to an unpublished source.","section":"Section 4, Propositions 4.1 and 4.3"},{"comment":"The quantity ell_k is computed as sum_i sum_v (1-m_iv) phi(...) + log pi_k, and then ell = log sum_k exp(ell_k). As written, this is not the observed-data log-likelihood: the component contribution should use log phi rather than phi, and the sum over observations should be outside the log-sum-exp, i.e. ell = sum_i log sum_k exp(log pi_k + sum_v (1-m_iv) log phi(...)). The convergence check and the BIC/AIC criterion therefore use an incorrect objective, which can affect the simulation results and model selection; the pseudocode should be corrected and the simulations should be verified with the corrected likelihood.","section":"Section A.1, Algorithm 1, lines 19-22"},{"comment":"Theorem 5.1 is stated for the sieve P_n with general, possibly non-diagonal covariance matrices, but the implemented estimator restricts to diagonal covariances, sieve fP_n. The paper asserts that the same asymptotic result could be shown for fP_n, but no proof is given. Since fP_n subset P_n does not imply that an approximate maximizer over fP_n is an approximate maximizer over P_n in the sense of condition (7'), the theorem does not formally cover the estimator used in the simulations. Please extend Theorem 5.1 to the diagonal sieve or provide the missing argument.","section":"Section 5 and Section 6"}],"minor_comments":[{"comment":"The claim that the estimator attains the minimax rate 'for any prescribed smoothness level' should be qualified: Theorem 5.1 requires Assumption 5.1 with known parameters beta and tau, and the sieve tuning in (11) depends on beta. The result is a fixed-smoothness guarantee, not an adaptive one.","section":"Abstract and Section 1"},{"comment":"Assumption 2.3 is not implied by MAR, and the paper itself notes in Section 4 that without pointwise positivity the adapted KL functional no longer identifies P0. The phrase 'natural positivity condition' in the abstract understates this: the rate constant 1/sqrt(c0) can be arbitrarily large. A short discussion of when c0 can be expected to be bounded away from zero, and of the consequences of small c0, would help the reader calibrate the scope of the result.","section":"Section 2, Assumption 2.3 and Remark 1"},{"comment":"The responsibility update in Algorithm 2 has a typo: the denominator uses pi_j p_{i ell} where it should be pi_ell p_{i ell}. The same typo appears in the E-step exposition following equation (15).","section":"Algorithm 2, line 8"},{"comment":"The captions refer to colors and dashed lines (e.g., 'dashed blue line'), but the figures appear without legends or labeled curves. Please add direct labels or a legend so the results are interpretable.","section":"Figures 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The central proof chain appears internally coherent conditional on Assumptions 2.1-2.3 and 5.1, and the paper is transparent about the positivity condition. The main issues are fixable: the companion-paper citations for Propositions 4.1 and 4.3, the incorrect log-likelihood in Algorithm 1, and the missing proof for the diagonal sieve used in practice. The self-citation pattern is heavy, but the borrowed results are clearly flagged and appear to be separate mathematical facts. I would like to see the revised manuscript before making a final recommendation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jeff, this is the real thing: the first frequentist rate guarantees for sieve MLE under general non-monotone MAR, with no modeling of the missingness mechanism. The main theorem (4.4) is a proper extension of Kaji's empirical-process bounds to observed-data marginals, and the adapted Hellinger machinery is clean. The density application (Thm 5.1) attains near-minimax H\\\"older rates for any smoothness, with MAR entering only through the constant 1/sqrt(c0). That is a substantial advance over the previous state of the art, which had only Bayesian contraction (their companion paper) or restrictive pattern-graph assumptions.\n\nThe paper is also honest about its main condition. Assumption 2.3 (uniform positivity of the propensity score) is load-bearing twice: it makes the adapted KL identify the complete-data density, and it converts eH rates into H rates. MAR alone does not imply it, and the rate degrades as c0 shrinks. The paper flags this in Remark 1 and the abstract says 'natural positivity condition.' Good. It is a scope limitation, not a hidden flaw.\n\nThe weak point that deserves referee attention is the reliance on the companion paper (N\\\"af and Ch\\'erief-Abdellatif 2026) for Propositions 4.1 and 4.3. The appendix proves the adapted inequalities and the entropy transfer, but the identification statement (Prop 4.1) is taken on faith. It may well be correct—the stress-test found no mathematical error, and the lower bound in Prop 4.3 is immediate—but an editor should ask for the companion proof or a self-contained appendix. This is a verification issue, not a correctness issue.\n\nThe simulations are adequate: they show the estimator works on a hard non-monotone MAR example and in d=20, beating complete-data KDE in some settings. They won't convince a skeptic about finite-sample behavior, but the theory is the point.\n\nBottom line: this deserves a serious referee. It is a major result within missing-data theory. I would recommend sending to a strong theory journal, with the companion-paper dependency as the main request for revision.","headline":"A genuine first: frequentist sieve MLE rates under general non-monotone MAR, with a load-bearing positivity assumption that is honestly disclosed.","tokens_in":49299,"tokens_out":2418,"would_cite":true,"duration_ms":23186,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G07","62G20","62D10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that under general non-monotone missing at random with a positivity condition, a sieve maximum-likelihood estimator recovers the complete-data density at the minimax rate up to a logarithmic factor, with missingness…","keywords":["missing at random","non-monotone missingness","sieve maximum likelihood","nonparametric density estimation","minimax rate","Gaussian mixture sieve","adapted Hellinger functional","positivity condition"],"falsifier":"Take a true density $p_0$ and a missingness mechanism with $P(M=0|X=x)=0$ on a cube of positive Lebesgue measure, and choose $p_1$ equal to $p_0$ outside that cube but different inside. The observed-data likelihood is identical under $p_0$ and $p_1$, so no sieve MLE can converge to the true density in Hellinger distance on that region; a simulation with such a mechanism would show the estimated density failing to recover $p_0$ exactly where the chance of full observation is zero.","tokens_in":1860,"feed_emoji":"📊","tokens_out":2556,"duration_ms":86363,"temperature":0.7,"pith_summary":"Missing at random (MAR) promises that the missingness mechanism can be ignored when maximizing likelihood, but for non-monotone missingness—where each variable can be missing in arbitrary combinations—frequentist guarantees for nonparametric estimation have been scarce. This paper establishes that a sieve maximum-likelihood estimator, maximizing only the observed-data marginal likelihood over a growing space of Gaussian mixtures, recovers the complete-data density consistently under general MAR, with no model for the missingness mechanism and no restriction on the configuration of missingness patterns. The rate is the minimax rate for a $\\beta$-smooth density up to a logarithmic factor, and missingness enters only through a constant $1/\\sqrt{c_0}$ arising from the positivity condition. This matters because it turns the long-held intuition that likelihood estimation ignores MAR into a theorem, while also providing a practical EM algorithm that performs like a complete-data kernel density estimator.","feed_headline":"Non-monotone missingness costs only a constant in density estimation","feed_subtitle":"Under missing at random, a sieve likelihood estimator matches the complete-data rate for any smoothness level.","key_machinery":"The load-bearing object is the adapted Hellinger functional, defined by summing pattern-wise squared Hellinger differences between observed-block marginals, weighted by $P(M=m|x^{(m)})$: $$\\widetilde{H}^2(f,g)=\\sum_m \\int \\big(\\sqrt{$f^{{(m)}}$($x^{{(m)}}$)}-\\sqrt{$g^{{(m)}}$($x^{{(m)}}$)}\\big)^2\\, P(M=m|$x^{{(m)}}$)\\,$dx^{{(m)}}$.$$ Under positivity it behaves like a distance and is squeezed between $c_0 H^2$ and $H^2$, so convergence in it implies convergence of the complete density and bracketing-entropy bounds transfer from complete data to observed data. The proof also uses a Gaussian-mixture sieve whose covariance eigenvalues are allowed to diverge with $n$, a localization argument that caps the eigenvalues for densities near the truth, and recent Hellinger bounds on the Kullback-Leibler divergence and the Bernstein norm that remove the need for bounded density ratios.","core_discovery":"The central discovery is Theorem 5.1: under MAR, a positivity condition $P(M=0|X=x)\\ge c_0>0$, and a H\\\"older smoothness condition with exponent $\\beta$, any approximate maximizer $\\hat p_n$ of the observed-data likelihood over the Gaussian-mixture sieve satisfies $H(p_0,\\hat p_n)=O_P\\big((C_4/\\sqrt{c_0})\\, n^{-\\beta/(2\\beta+d)}(\\log n)^t\\big)$. Thus the complete-data density is estimated at the minimax rate up to a logarithmic factor, for any prescribed smoothness level, even when each coordinate can be missing in arbitrary non-monotone patterns. The missingness mechanism is never modeled or estimated; it affects only the multiplicative constant through $c_0$. Theorem 4.4 supplies the general sieve-MLE rate under MAR that powers this density result, and the paper shows the estimator is approximated in practice by a simple EM algorithm operating directly on the incomplete data.","pith_inferences":["If positivity fails on a region of positive measure, the adapted Hellinger functional ceases to identify the complete density there; pattern-specific identification assumptions would be needed, so the positivity condition marks the boundary of what MAR alone identifies.","The same proof architecture should transfer to other sieves, such as nonparametric maximum likelihood for Gaussian location mixtures or sieve estimators in nonparametric regression, making Theorem 4.4 a template for ignoring estimators beyond density estimation.","A practical diagnostic extension would estimate the chance of full observation and flag samples where its minimum is near zero, since the rate constant $1/\\sqrt{c_0}$ warns that estimation will be poor in low-observation regions even when consistency holds.","The logarithmic factor originates in the entropy of the Gaussian-mixture sieve rather than in the missingness, so removing it may be possible with sharper entropy or localization arguments."],"forward_implications":["Ignoring maximum likelihood is now a valid nonparametric estimator under general non-monotone MAR: consistency and rates hold without specifying a model for why values are missing.","The rate is minimax up to logarithmic factors: additional missingness patterns, or complex dependencies between missingness and observed values, do not slow the density estimator beyond the constant factor $1/\\sqrt{c_0}$.","Any statistic that is a continuous functional of the density—such as quantiles or smooth M-estimators—inherits consistency when applied to the estimated density.","The estimator is computable: a closed-form EM update on the incomplete data approximates the sieve maximizer, so the theoretical guarantee has a practical algorithm attached.","Complete-data kernel density estimators with Gaussian kernels are limited to smoothness $\\beta\\le 2$, while the Gaussian-mixture sieve reaches near-minimax rates for every $\\beta>0$."],"supporting_citations":[{"why":"Introduces MAR and the ignorability property that justifies maximizing the observed-data likelihood while ignoring the missingness mechanism; this is the paper's starting point.","marker":"Rubin (1976)"},{"why":"Shows consistency and asymptotic normality of the ignoring MLE in parametric models, the frequentist baseline the paper extends to nonparametric sieves.","marker":"Takai and Kano (2013)"},{"why":"Gives Bayesian posterior contraction rates under MAR along with the adapted KL and Hellinger tools and the positivity condition that the paper adapts to the frequentist sieve MLE.","marker":"Näf and Chérief-Abdellatif (2026)"},{"why":"Supplies Hellinger bounds on the KL divergence and the Bernstein norm that let the paper avoid the bounded-density-ratio assumption in the sieve-MLE proof.","marker":"Kaji (2026)"},{"why":"Provides the empirical-process theorem and bracketing-entropy conditions that form the skeleton of Theorem 4.4.","marker":"van der Vaart and Wellner (2023)"},{"why":"Contributes the Gaussian-mixture sieve construction, approximation results, and entropy bounds that Theorem 5.1 adapts to missing data.","marker":"Ghosal and van der Vaart (2017a)"},{"why":"Establishes the minimax lower bound for $\\beta$-smooth density estimation that the paper's rate matches up to a logarithmic factor.","marker":"Tsybakov (2009)"}],"fun_headline_variants":["Sieve MLE matches complete-data rate despite arbitrary missingness","MAR missingness: rate unaffected, only constant","Non-monotone MAR? No rate penalty for density estimation","Density estimation at minimax rate under general missingness","Arbitrary missingness costs only a constant in density risk"],"cache_read_input_tokens":51328,"weakest_assumption_plain":"Every data point must have at least a fixed positive chance of being observed completely; if some region of the covariate space never produces a fully observed record, the observed data cannot tell apart densities that differ only there, and the claimed rate fails.","fun_headline_variants_meta":{"raw":{"variants":["Sieve MLE matches complete-data rate despite arbitrary missingness","MAR missingness: rate unaffected, only constant","Non-monotone MAR? No rate penalty for density estimation","Density estimation at minimax rate under general missingness","Arbitrary missingness costs only a constant in density risk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000786,"raw_usage":{"total_tokens":3487,"prompt_tokens":985,"completion_tokens":2502,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":2420}},"tokens_in":601,"tokens_out":2502,"duration_ms":15874,"temperature":1.0,"reasoning_tokens":2420,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:14.768357+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a true density $p_0$ and a missingness mechanism with $P(M=0|X=x)=0$ on a cube of positive Lebesgue measure, and choose $p_1$ equal to $p_0$ outside that cube but different inside. The observed-data likelihood is identical under $p_0$ and $p_1$, so no sieve MLE can converge to the true density in Hellinger distance on that region; a simulation with such a mechanism would show the estimated density failing to recover $p_0$ exactly where the chance of full observation is zero.","supporting_citations":[],"review_version":1}