Pith. sign in

REVIEW 3 major objections 4 minor 300 references

Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper proves that under general non-monotone missing at random with a positivity condition, a sieve maximum-likelihood estimator recovers the complete-data density at the minimax rate up to a logarithmic factor, with missingness…

desk verdict A genuine first: frequentist sieve MLE rates under general non-monotone MAR, with a load-bearing positivity assumption that is honestly disclosed. read the letter →

arxiv 2608.10113 v1 pith:FBC5RM43 submitted 2026-08-10 stat.ME

classification stat.ME MSC 62G0762G2062D10
keywords missingatrandomnon-monotonemissingnesssievemaximumlikelihoodnonparametricdensityestimationminimaxrateGaussianmixtureadaptedHellingerfunctionalpositivitycondition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Missing at random (MAR) promises that the missingness mechanism can be ignored when maximizing likelihood, but for non-monotone missingness—where each variable can be missing in arbitrary combinations—frequentist guarantees for nonparametric estimation have been scarce. This paper establishes that a sieve maximum-likelihood estimator, maximizing only the observed-data marginal likelihood over a growing space of Gaussian mixtures, recovers the complete-data density consistently under general MAR, with no model for the missingness mechanism and no restriction on the configuration of missingness patterns. The rate is the minimax rate for a $\beta$-smooth density up to a logarithmic factor, and missingness enters only through a constant $1/\sqrt{c_0}$ arising from the positivity condition. This matters because it turns the long-held intuition that likelihood estimation ignores MAR into a theorem, while also providing a practical EM algorithm that performs like a complete-data kernel density estimator.

What carries the argument

The load-bearing object is the adapted Hellinger functional, defined by summing pattern-wise squared Hellinger differences between observed-block marginals, weighted by $P(M=m|x^{(m)})$: $$\widetilde{H}^2(f,g)=\sum_m \int \big(\sqrt{$f^{{(m)}}$($x^{{(m)}}$)}-\sqrt{$g^{{(m)}}$($x^{{(m)}}$)}\big)^2\, P(M=m|$x^{{(m)}}$)\,$dx^{{(m)}}$.$$ Under positivity it behaves like a distance and is squeezed between $c_0 H^2$ and $H^2$, so convergence in it implies convergence of the complete density and bracketing-entropy bounds transfer from complete data to observed data. The proof also uses a Gaussian-mixture sieve whose covariance eigenvalues are allowed to diverge with $n$, a localization argument that caps the eigenvalues for densities near the truth, and recent Hellinger bounds on the Kullback-Leibler divergence and the Bernstein norm that remove the need for bounded density ratios.

What would settle it

Take a true density $p_0$ and a missingness mechanism with $P(M=0|X=x)=0$ on a cube of positive Lebesgue measure, and choose $p_1$ equal to $p_0$ outside that cube but different inside. The observed-data likelihood is identical under $p_0$ and $p_1$, so no sieve MLE can converge to the true density in Hellinger distance on that region; a simulation with such a mechanism would show the estimated density failing to recover $p_0$ exactly where the chance of full observation is zero.

Watch

Extended reading notes

Core claim

The central discovery is Theorem 5.1: under MAR, a positivity condition $P(M=0|X=x)\ge c_0>0$, and a H\"older smoothness condition with exponent $\beta$, any approximate maximizer $\hat p_n$ of the observed-data likelihood over the Gaussian-mixture sieve satisfies $H(p_0,\hat p_n)=O_P\big((C_4/\sqrt{c_0})\, n^{-\beta/(2\beta+d)}(\log n)^t\big)$. Thus the complete-data density is estimated at the minimax rate up to a logarithmic factor, for any prescribed smoothness level, even when each coordinate can be missing in arbitrary non-monotone patterns. The missingness mechanism is never modeled or estimated; it affects only the multiplicative constant through $c_0$. Theorem 4.4 supplies the general sieve-MLE rate under MAR that powers this density result, and the paper shows the estimator is approximated in practice by a simple EM algorithm operating directly on the incomplete data.

Load-bearing premise

Every data point must have at least a fixed positive chance of being observed completely; if some region of the covariate space never produces a fully observed record, the observed data cannot tell apart densities that differ only there, and the claimed rate fails.

Editorial extensions

If this is right

  • Ignoring maximum likelihood is now a valid nonparametric estimator under general non-monotone MAR: consistency and rates hold without specifying a model for why values are missing.
  • The rate is minimax up to logarithmic factors: additional missingness patterns, or complex dependencies between missingness and observed values, do not slow the density estimator beyond the constant factor $1/\sqrt{c_0}$.
  • Any statistic that is a continuous functional of the density—such as quantiles or smooth M-estimators—inherits consistency when applied to the estimated density.
  • The estimator is computable: a closed-form EM update on the incomplete data approximates the sieve maximizer, so the theoretical guarantee has a practical algorithm attached.
  • Complete-data kernel density estimators with Gaussian kernels are limited to smoothness $\beta\le 2$, while the Gaussian-mixture sieve reaches near-minimax rates for every $\beta>0$.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If positivity fails on a region of positive measure, the adapted Hellinger functional ceases to identify the complete density there; pattern-specific identification assumptions would be needed, so the positivity condition marks the boundary of what MAR alone identifies.
  • The same proof architecture should transfer to other sieves, such as nonparametric maximum likelihood for Gaussian location mixtures or sieve estimators in nonparametric regression, making Theorem 4.4 a template for ignoring estimators beyond density estimation.
  • A practical diagnostic extension would estimate the chance of full observation and flag samples where its minimum is near zero, since the rate constant $1/\sqrt{c_0}$ warns that estimation will be poor in low-observation regions even when consistency holds.
  • The logarithmic factor originates in the entropy of the Gaussian-mixture sieve rather than in the missingness, so removing it may be possible with sharper entropy or localization arguments.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper develops a frequentist sieve maximum likelihood theory for estimation under non-monotone missing at random (MAR) missingness. It introduces an adapted Kullback-Leibler divergence and an adapted Hellinger functional on observed-data marginals, and proves (Theorem 4.4) that an approximate maximizer of the observed-data likelihood over any sieve satisfying bracketing entropy, approximation, and moment conditions converges at rate eH(p0, p_hat n)=O_P(delta_n or n^{-1/2}), with H-rate degraded by a 1/sqrt(c0) factor under a uniform positivity assumption on the propensity score. Applying this to a Gaussian mixture sieve under a Holder smoothness class with exponential tails (Theorem 5.1), it obtains H(p0, p_hat n)=O_P(C4/sqrt(c0) n^{-beta/(2beta+d)} (log n)^t), the minimax rate up to logarithmic factors, with missingness affecting only the constant. The estimator is implemented by a MAR-specific EM algorithm with BIC/AIC selection; simulations in d=20 compare it with complete-data Gaussian KDE. The proofs are detailed in an appendix, although two key identifiability and rate-bridge propositions are cited from an unpublished companion paper.

Significance. If correct, this is a substantial contribution: it provides the first general frequentist rate guarantees for nonparametric likelihood-based estimation under non-monotone MAR, and it removes the bounded-density-ratio condition via Kaji's Hellinger bounds. The adapted-divergence machinery is clean, and the adapted Hellinger functional is well suited to pattern-marginal data. The result that missingness changes only the constant in the minimax rate is striking and is stated with transparent assumptions. The authors are also candid about the role of the positivity assumption, and the new entropy localization arguments in Lemmas B.6-B.9 are worked out in detail. Reproducible code is provided. The significance is moderated, however, by the reliance on a strong uniform positivity condition, by the restriction to Holder classes with strong tail and moment conditions and known smoothness, and by the fact that two load-bearing results (Propositions 4.1 and 4.3) are borrowed from an unpublished companion preprint rather than proved in the manuscript.

major comments (3)
  1. [Section 4, Propositions 4.1 and 4.3] The identifiability of the adapted KL divergence (Proposition 4.1) and the two-sided bound c0 H^2(p1,p2) <= eH^2(p1,p2) <= H^2(p1,p2) (Proposition 4.3) are both cited from the companion preprint Naf and Cherief-Abdellatif (2026), which is not peer-reviewed. These propositions are load-bearing: Proposition 4.1 justifies the target of the sieve MLE, and Proposition 4.3 is used in Lemma B.3 and in the final transition from eH to H in Theorems 4.4 and 5.1. For a self-contained journal submission, the full proofs of these propositions should be included in the appendix or derived in the paper, rather than referenced to an unpublished source.
  2. [Section A.1, Algorithm 1, lines 19-22] The quantity ell_k is computed as sum_i sum_v (1-m_iv) phi(...) + log pi_k, and then ell = log sum_k exp(ell_k). As written, this is not the observed-data log-likelihood: the component contribution should use log phi rather than phi, and the sum over observations should be outside the log-sum-exp, i.e. ell = sum_i log sum_k exp(log pi_k + sum_v (1-m_iv) log phi(...)). The convergence check and the BIC/AIC criterion therefore use an incorrect objective, which can affect the simulation results and model selection; the pseudocode should be corrected and the simulations should be verified with the corrected likelihood.
  3. [Section 5 and Section 6] Theorem 5.1 is stated for the sieve P_n with general, possibly non-diagonal covariance matrices, but the implemented estimator restricts to diagonal covariances, sieve fP_n. The paper asserts that the same asymptotic result could be shown for fP_n, but no proof is given. Since fP_n subset P_n does not imply that an approximate maximizer over fP_n is an approximate maximizer over P_n in the sense of condition (7'), the theorem does not formally cover the estimator used in the simulations. Please extend Theorem 5.1 to the diagonal sieve or provide the missing argument.
minor comments (4)
  1. [Abstract and Section 1] The claim that the estimator attains the minimax rate 'for any prescribed smoothness level' should be qualified: Theorem 5.1 requires Assumption 5.1 with known parameters beta and tau, and the sieve tuning in (11) depends on beta. The result is a fixed-smoothness guarantee, not an adaptive one.
  2. [Section 2, Assumption 2.3 and Remark 1] Assumption 2.3 is not implied by MAR, and the paper itself notes in Section 4 that without pointwise positivity the adapted KL functional no longer identifies P0. The phrase 'natural positivity condition' in the abstract understates this: the rate constant 1/sqrt(c0) can be arbitrarily large. A short discussion of when c0 can be expected to be bounded away from zero, and of the consequences of small c0, would help the reader calibrate the scope of the result.
  3. [Algorithm 2, line 8] The responsibility update in Algorithm 2 has a typo: the denominator uses pi_j p_{i ell} where it should be pi_ell p_{i ell}. The same typo appears in the E-step exposition following equation (15).
  4. [Figures 2 and 3] The captions refer to colors and dashed lines (e.g., 'dashed blue line'), but the figures appear without legends or labeled curves. Please add direct labels or a legend so the results are interpretable.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the rate theorem is derived by verifying sieve entropy and bias conditions, and the self-citations transmit independent lemmas rather than the target conclusion.

full rationale

The paper's central claim, Theorem 5.1, states that an approximate maximizer of the observed-data likelihood over a Gaussian-mixture sieve satisfies H(p0, p_hat_n) = O_P((C4/sqrt(c0)) n^{-beta/(2beta+d)} (log n)^t). This conclusion is obtained by checking the three conditions (4'), (5'), and (6') of the general sieve-MLE rate theorem, Theorem 4.4, on a concrete sieve; the proof in Appendix B supplies the entropy and approximation arguments rather than assuming the rate. The quantity delta_n is a candidate rate verified through bracketing entropy bounds and a bias-correction construction, not a parameter fitted to the conclusion. Missingness enters only through the constant c0, via Proposition 4.3, whose bound c0 H^2(p1,p2) <= eH^2(p1,p2) <= H^2(p1,p2) is an immediate consequence of the definition of the adapted Hellinger functional and Assumption 2.3: the m=0 pattern contributes H^2(p1,p2) with weight P(M=0|x) >= c0, and the upper bound follows from Cauchy-Schwarz and summing propensity weights to one. Although Proposition 4.3 and Proposition 4.1 are cited from the companion paper by Naf and Cherief-Abdellatif (2026), they are parameter-free mathematical lemmas whose assumptions do not include the frequentist rate theorem being proved, so they are independent support rather than a self-citation chain. The approximate-maximizer condition (7') is the standard condition from the external sieve-MLE theory of Kaji and van der Vaart and Wellner, and the paper verifies it in the proof. There is no fitted input renamed as a prediction, no uniqueness theorem imported to forbid alternatives, and no ansatz smuggled in via citation. Assumption 2.3 is indeed load-bearing, but the abstract and Remark 1 explicitly present it as a stated positivity condition; a scope limitation is not circularity. The simulation study is empirical and does not enter the proof chain. Therefore the derivation does not reduce to its own inputs and no circular step is present.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim rests on standard missing data assumptions (MAR, positivity, no empty pattern) plus a Holder class condition with tail restrictions. The sieve parameters C4 and t are theoretical constants, not fitted to data. No new physical or structural entities are postulated.

free parameters (3)
  • C4
    Large enough constant in the rate delta_n = C4 n^{-beta/(2beta+d)} (log n)^t; chosen to satisfy entropy and bias conditions in the proof, not fitted to data.
  • t
    Exponent in the logarithmic factor, required to satisfy inequality (10); a theoretical constant chosen in the proof, not fitted.
  • Covariance regularization lambda
    Regularization added to covariance updates in Algorithm 1 to ensure positivity; chosen by the user in practice, not part of the main theorem.
assumptions (5)
  • domain assumption Assumption 2.1 (MAR): P(M=m|X=x) = P(M=m|X^(m)=x^(m)) for all m, x.
    Core missingness assumption; used throughout to rewrite observed-data likelihood as a propensity-weighted marginal expectation.
  • domain assumption Assumption 2.2: P(M=1)=0 (the completely missing pattern has probability zero).
    Avoids observing no coordinates; needed for the observed-data likelihood to be well-defined.
  • domain assumption Assumption 2.3: P(M=0|X=x) >= c0 > 0.
    Positivity of the propensity score; turns the adapted Hellinger pseudometric into a metric and sets the constant in the rate.
  • domain assumption Assumption 5.1: Holder class with moment conditions and exponential tail.
    Smoothness and tail assumptions on the true density p0 needed for Gaussian mixture approximation and entropy bounds.
  • standard math Empirical process theory: van der Vaart and Wellner (2023) Theorem 3.4.1 and Kaji (2026) inequalities.
    Framework for sieve MLE rates; the paper adapts these tools to the MAR setting.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach." pith.science (2026). https://pith.science/paper/FBC5RM43

@misc{pith2026260810113,
  author       = {Pith},
  title        = {Pith review of: Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FBC5RM43}},
  note         = {Machine review of arXiv:2608.10113}
}
read the original abstract

Missing data constitute a pervasive challenge in empirical research. Consequently, there is an ever-growing number of methods designed to address this challenge, with multiple imputation and inverse probability weighting the dominant strategies. Despite this, theoretical guarantees remain limited, particularly in the challenging case of non-monotone missing at random (MAR). When guarantees exist, they are often confined to simplified settings such as missing completely at random, monotone or block-wise missingness, or rest on restrictive assumptions about the missingness mechanism. In this paper, we utilize the theory of sieve maximum likelihood to establish a general rate of convergence under MAR that requires no modeling of the missingness mechanism and no restriction on the configuration of missing patterns, beyond MAR itself and a natural positivity condition. Applying this result to density estimation, we show that the complete-data density can be estimated at the minimax rate over a H\"older class, up to a logarithmic factor, for any prescribed smoothness level. The missingness does not affect the rate and enters only through a constant. The estimator is approximated in practice by a simple expectation-maximization (EM) algorithm operating on the incomplete data directly. In simulations, it performs comparably to the kernel density estimator supplied with the complete data across a wide range of missingness levels.

Figures

Figures reproduced from arXiv: 2608.10113 by the authors.

Figure 1
Figure 1. Three Data matrices with missing values, each with three different patterns. Each [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Estimation of the 0.1 quantile of X1 (left) and correlation between X1, X2 (right), for B = 50, with n = 5000 observations each. The dashed blue line corresponds to the correct value. “missForest” refers to the widely-used method of Stekhoven and B¨uhlmann (2011), “mean” to mean imputation, where missing values are replaced by the sample mean, “MLE” to the ignoring MLE in (1) and “Mest” to the ignoring M-estimator, … view at source ↗
Figure 3
Figure 3. Estimation of 0.1 quantile of X1 for data in the motivating example, with B = 50 and n = 5000 observations for each batch. G(x) = max(Φ(x), c0) to ensure P(M = m1 | x) ≥ c0. Capability in recovering complete-data density for d = 20. Here, we compare our method fitted on data with varying levels of missingness, against the Gaussian Kernel Density Estimator (KDE) with complete data. We directly apply scipy.stats.gauss… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Data generated from a 20-component Gaussian Mixture with randomly simulated loca [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: shows that our method outperforms Gaussian KDE at all missingness levels in terms of energy distance. This implies that the point estimate from our method is very close to the truth on a global scale since the energy distance measures the overall closeness between two …
Figure 6
Figure 6. Figure 6: Multivariate logistic data consists of variables with standard logistic marginal distribu [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: Data sampled from N20(0, Σ) with X1, X2, . . . , X19 correlated, and X20 independent from others. For our method, we use shared diagonal covariances for all components of base models, and select the model by the AIC criterion. The number of random initializations is 30…
Figure 8
Figure 8. Figure 8: Multivariate logistic data consists of variables with standard logistic marginal distribu [PITH_FULL_IMAGE:figures/full_fig_p027_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 59 canonical work pages

  1. [1]

    Dempster, A. P. and Laird, N. M. and Rubin, D. B. , title =. Journal of the Royal Statistical Society: Series B (Methodological) , volume =

  2. [2]

    and Krishnan, Thriyambakam , title =

    McLachlan, Geoffrey J. and Krishnan, Thriyambakam , title =

  3. [3]

    and Yu, Bin , title =

    Balakrishnan, Sivaraman and Wainwright, Martin J. and Yu, Bin , title =. The Annals of Statistics , volume =. 2017 , doi =

  4. [4]

    The Annals of Statistics , volume =

    Saha, Sujayam and Guntuboyina, Adityanand , title =. The Annals of Statistics , volume =. 2020 , month =

  5. [5]

    arXiv preprint arXiv:2602.19203 , year =

    A Calibration Framework for Inference with Partially Observed Data , author =. arXiv preprint arXiv:2602.19203 , year =

  6. [6]

    Annals of Statistics , year =

    Chen, Yen-Chi , title =. Annals of Statistics , year =

  7. [7]

    , title =

    Tsybakov, Alexandre B. , title =. 2009 , doi =

  8. [8]

    2019 , publisher =

    High-Dimensional Statistics: A Non-Asymptotic Viewpoint , author =. 2019 , publisher =

Show all 300 references
  1. [9]

    The Annals of Statistics , year =

    Shen, Xiaotong and Wong, Wing Hung , title =. The Annals of Statistics , year =

  2. [10]

    The Annals of Statistics , year =

    Wong, Wing Hung and Shen, Xiaotong , title =. The Annals of Statistics , year =

  3. [11]

    The Annals of Statistics , year =

    Shen, Xiaotong , title =. The Annals of Statistics , year =

  4. [12]

    and van der Vaart, Aad W

    Ghosal, Subhashis and Ghosh, Jayanta K. and van der Vaart, Aad W. , title =. The Annals of Statistics , year =

  5. [13]

    Journal of the American Statistical Association , year =

    Chen, Xiaohong and Hong, Han and Tarozzi, Alessandro , title =. Journal of the American Statistical Association , year =

  6. [14]

    The Japanese Economic Review , year =

    Kaji, Tetsuya , title =. The Japanese Economic Review , year =

  7. [15]

    , title =

    Ghosal, Subhashis and van der Vaart, Aad W. , title =. 2017 , doi =

  8. [16]

    and Wellner, Jon A

    van der Vaart, Aad W. and Wellner, Jon A. , title =. 2023 , doi =

  9. [17]

    Rates of Convergence for Minimum Contrast Estimators , journal =

    Birg. Rates of Convergence for Minimum Contrast Estimators , journal =. 1993 , volume =

  10. [18]

    arXiv preprint arXiv:2508.15162 , year =

    A Unified Framework for Inference with General Missingness Patterns and Machine Learning Imputation , author =. arXiv preprint arXiv:2508.15162 , year =

  11. [19]

    Nijman, S. W. J. and Leeuwenberg, A. M. and Beekers, I. and Verkouter, I. and Jacobs, J. J. L. and Bots, M. L. and Asselbergs, F. W. and Moons, K. G. M. and Debray, T. P. A. , title =. Journal of Clinical Epidemiology , year =

  12. [20]

    2008 , edition =

    Ambrosio, Luigi and Gigli, Nicola and Savaré, Giuseppe , title =. 2008 , edition =

  13. [21]

    arXiv preprint arXiv:2411.02549 , year =

    Distributionally Robust Optimization , author =. arXiv preprint arXiv:2411.02549 , year =

  14. [22]

    arXiv preprint arXiv:2401.14655 , year =

    Distributionally Robust Optimization and Robust Statistics , author =. arXiv preprint arXiv:2401.14655 , year =

  15. [23]

    arXiv preprint arXiv:2209.01754 , year =

    Sahoo, Roshni and Lei, Lihua and Wager, Stefan , title =. arXiv preprint arXiv:2209.01754 , year =

  16. [24]

    , title =

    Wei, Ying and Ma, Yanyuan and Carroll, Raymond J. , title =. Biometrika , volume =

  17. [25]

    and Yang, Y

    Wei, Y. and Yang, Y. , title =. Statistica Sinica , year =

  18. [26]

    Xuerong Chen and Alan T. K. Wan and Yong Zhou , title =. Journal of the American Statistical Association , volume =

  19. [27]

    Bayesian nonparametric statistics,

    Castillo, Isma. Bayesian nonparametric statistics,. arXiv preprint arXiv:2402.16422 , year=

  20. [28]

    Stochastic Gradient Methods for Distributionally Robust Optimization with f-divergences , url =

    Namkoong, Hongseok and Duchi, John C , booktitle =. Stochastic Gradient Methods for Distributionally Robust Optimization with f-divergences , url =

  21. [29]

    and Namkoong, Hongseok , title =

    Duchi, John C. and Namkoong, Hongseok , title =. Annals of Statistics , year =

  22. [30]

    Scientific Data , author =

    Introducing the. Scientific Data , author =. 2022 , pages =

  23. [31]

    , title =

    Huber, Peter J. , title =. The Annals of Mathematical Statistics , year =

  24. [32]

    2009 , publisher =

    Optimal Transport: Old and New , author =. 2009 , publisher =

  25. [33]

    Bayesian Analysis , volume =

    Variational Inference for Dirichlet Process Mixtures , author =. Bayesian Analysis , volume =

  26. [34]

    Tchetgen Tchetgen , title =

    BaoLuo Sun and Eric J. Tchetgen Tchetgen , title =. Journal of the American Statistical Association , volume =

  27. [35]

    2019 , publisher =

    Statistical Analysis with Missing Data., 3rd Edition , author =. 2019 , publisher =

  28. [36]

    Roderick J. A. Little and Rubin B. Donald , title =. 1986 , publisher =

  29. [37]

    Rubin , title =

    Donald B. Rubin , title =

  30. [38]

    , title =

    Rubin, Donald B. , title =. Biometrika , volume =

  31. [39]

    Rubin , title =

    Donald B. Rubin , title =. Journal of the American Statistical Association , volume =

  32. [40]

    2013 , publisher =

    Seaman, Shaun and Galati, John and Jackson, Dan and Carlin, John , journal =. 2013 , publisher =

  33. [41]

    Biometrika , volume =

    Farewell, D M and Daniel, R M and Seaman, S R , title =. Biometrika , volume =

  34. [42]

    , title =

    Mealli, Fabrizia and Rubin, Donald B. , title =. Biometrika , volume =

  35. [43]

    International Statistical Review , volume =

    Doretti, Marco and Geneletti, Sara and Stanghellini, Elena , title =. International Statistical Review , volume =

  36. [44]

    Copas , title =

    Guobing Lu and John B. Copas , title =. The Annals of Statistics , number =

  37. [45]

    Statistical Methods in Medical Research , volume =

    Can one assess whether missing data are missing at random in medical studies? , author =. Statistical Methods in Medical Research , volume =

  38. [46]

    Psychometrika , year =

    Rabe-Hesketh, Sophia and Skrondal, Anders , title =. Psychometrika , year =

  39. [47]

    Kenward , title =

    Geert Molenberghs and Caroline Beunckens and Cristina Sotto and Michael G. Kenward , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , number =

  40. [48]

    Communications in Statistics - Theory and Methods , volume =

    Keiji Takai and Yutaka Kano , title =. Communications in Statistics - Theory and Methods , volume =

  41. [49]

    and Henley, Steven S

    Golden, Richard M. and Henley, Steven S. and White, Halbert and Kashner, T. Michael , title =. Econometrics , volume =. 2019 , number =

  42. [50]

    2020 , author =

    M-estimation with incomplete and dependent multivariate data , journal =. 2020 , author =

  43. [51]

    Statistical Methods in Medical Research , volume =

    On weighting approaches for missing data , author =. Statistical Methods in Medical Research , volume =

  44. [52]

    2020 , author =

    Semiparametric inference with missing data: Robustness to outliers and model misspecification , journal =. 2020 , author =

  45. [53]

    Tchetgen Tchetgen , title =

    Daniel Malinsky and Ilya Shpitser and Eric J. Tchetgen Tchetgen , title =. Journal of the American Statistical Association , volume =

  46. [54]

    Ibrahim and Stuart R

    Joseph G. Ibrahim and Stuart R. Lipsitz and Ming-Hui Chen , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume =

  47. [55]

    2016 , author =

    Multiply imputing missing values in data sets with mixed measurement scales using a sequence of generalised linear models , journal =. 2016 , author =

  48. [56]

    and Winterstein, Almut G

    Xu, Dandan and Daniels, Michael J. and Winterstein, Almut G. , journal =. Sequential

  49. [57]

    Journal of the American Statistical Association , volume =

    Joseph G Ibrahim and Ming-Hui Chen and Stuart R Lipsitz and Amy H Herring , title =. Journal of the American Statistical Association , volume =

  50. [58]

    Murray , title =

    Jared S. Murray , title =. Statistical Science , number =

  51. [59]

    Statistical Methods in Medical Research , volume =

    Multiple imputation of discrete and continuous data by fully conditional specification , author =. Statistical Methods in Medical Research , volume =

  52. [60]

    Journal of Statistical Software , volume =

    Stef. Journal of Statistical Software , volume =. 2011 , pages =

  53. [61]

    Flexible Imputation of Missing Data

    Stef. Flexible Imputation of Missing Data. Second Edition , publisher =

  54. [62]

    On the stationary distribution of iterative imputations , year =

    Jingchen Liu and Andrew Gelman and Jennifer Hill and Yu-Sung Su and Jonathan Kropko , journal =. On the stationary distribution of iterative imputations , year =

  55. [63]

    Raghunathan , title =

    Jian Zhu and Trivellore E. Raghunathan , title =. Journal of the American Statistical Association , volume =

  56. [64]

    and Bartlett, Jonathan W

    Shah, Anoop D. and Bartlett, Jonathan W. and Carpenter, James and Nicholas, Owen and Hemingway, Harry , title =. American Journal of Epidemiology , volume =

  57. [65]

    and Bühlmann, Peter , title =

    Stekhoven, Daniel J. and Bühlmann, Peter , title =. Bioinformatics , volume =

  58. [66]

    2022 , note =

    missForest: Nonparametric Missing Value Imputation using Random Forest , author =. 2022 , note =

  59. [67]

    1997 , publisher =

    Analysis of Incomplete Multivariate Data , author =. 1997 , publisher =

  60. [68]

    Multiple Imputation for Multivariate Missing-Data Problems: A Data Analyst's Perspective , volume =

    Schafer, Joseph and Olsen, Maren , year =. Multiple Imputation for Multivariate Missing-Data Problems: A Data Analyst's Perspective , volume =

  61. [69]

    Biometrika , volume =

    Large-sample theory for parametric multiple imputation procedures , author =. Biometrika , volume =. 1998 , publisher =

  62. [70]

    Statistical Science , number =

    Xiao-Li Meng , title =. Statistical Science , number =

  63. [71]

    Statistica Sinica , year =

    Guan, Qian and Yang, Shu , title =. Statistica Sinica , year =

  64. [72]

    and Heumann, C

    Schomaker, M. and Heumann, C. , title =. Statistics in Medicine , year =

  65. [73]

    Statistical Methods in Medical Research , volume =

    Shaun R Seaman and Rachael A Hughes , title =. Statistical Methods in Medical Research , volume =

  66. [74]

    , title =

    Rubin, Donald B. , title =. Journal of Business & Economic Statistics , volume =

  67. [75]

    Roderick J. A. Little , title =. Biometrika , number =

  68. [76]

    Roderick J. A. Little , title =. Journal of the American Statistical Association , volume =

  69. [77]

    and Michiels, B

    Molenberghs, G. and Michiels, B. and Kenward, M. G. and Diggle, P. J. , title =. Statistica Neerlandica , volume =

  70. [78]

    arXiv preprint arXiv:1904.11085 , year =

    Nonparametric pattern-mixture models for inference with missing data , author =. arXiv preprint arXiv:1904.11085 , year =

  71. [79]

    Little, Roderick J. A. , title =. Journal of Business & Economic Statistics , number =

  72. [80]

    Statistical Methods in Medical Research , volume =

    Lauren J Beesley and Irina Bondarenko and Michael R Elliot and Allison W Kurian and Steven J Katz and Jeremy MG Taylor , title =. Statistical Methods in Medical Research , volume =

  73. [81]

    Statistical Methods in Medical Research , volume =

    Multiple imputation for non-monotone missing not at random data using the no self-censoring model , author =. Statistical Methods in Medical Research , volume =

  74. [82]

    Proceedings of the 37th International Conference on Machine Learning , pages =

    Missing Data Imputation using Optimal Transport , author =. Proceedings of the 37th International Conference on Machine Learning , pages =

  75. [83]

    Yoon, Jinsung and Jordon, James and van der Schaar, Mihaela , booktitle =

  76. [84]

    2019 , volume =

    Mattei, Pierre-Alexandre and Frellsen, Jes , booktitle =. 2019 , volume =

  77. [85]

    , booktitle =

    Li, Steven Cheng-Xian and Jiang, Bo and Marlin, Benjamin M. , booktitle =

  78. [86]

    Ouyang, Yidong and Xie, Liyan and Li, Chongxuan and Cheng, Guang , booktitle =

  79. [87]

    Rethinking the Diffusion Models for Missing Data Imputation: A Gradient Flow Perspective , volume =

    Chen, Zhichao and Li, Haoxuan and Wang, Fangyikang and Zhang, Odin and Xu, Hu and Jiang, Xiaoyu and Song, Zhihuan and Wang, Hao , booktitle =. Rethinking the Diffusion Models for Missing Data Imputation: A Gradient Flow Perspective , volume =

  80. [88]

    Advances in Neural Information Processing Systems , volume =

    Missing Data Imputation by Reducing Mutual Information with Rectified Flows , author =. Advances in Neural Information Processing Systems , volume =

  81. [89]

    arXiv preprint arXiv:2206.07769 , year =

    Jarrett, Daniel and Cebere, Bogdan and Liu, Tennison and Curth, Alicia and van der Schaar, Mihaela , title =. arXiv preprint arXiv:2206.07769 , year =

  82. [90]

    Are deep learning models superior for missing data imputation in surveys?

    Zhenhua Wang and Olanrewaju Akande and Jason Poulos and Fan Li , journal =. Are deep learning models superior for missing data imputation in surveys?. 2022 , volume =

  83. [91]

    Frontiers in Big Data , volume =

    Jäger, Sebastian and Allhorn, Arndt and Bießmann, Felix , title =. Frontiers in Big Data , volume =

  84. [92]

    2025 , journal =

    Do we Need Dozens of Methods for Real World Missing Value Imputation? , author =. 2025 , journal =

  85. [93]

    and Roberts, M

    Shadbahr, T. and Roberts, M. and Stanczuk, J. and others , title =. Communications Medicine , year =

  86. [94]

    The Annals of Applied Statistics , number =

    Jeffrey Näf and Meta-Lina Spohn and Loris Michel and Nicolai Meinshausen , title =. The Annals of Applied Statistics , number =

  87. [95]

    arXiv preprint arXiv:2507.11297 , primaryClass=

    How to rank imputation methods? , author=. arXiv preprint arXiv:2507.11297 , primaryClass=. 2025 , eprint=

  88. [96]

    What Is a Good Imputation Under

    Jeffrey Näf and Erwan Scornet and Julie Josse , journal =. What Is a Good Imputation Under

  89. [97]

    2026 , journal =

    A Practical Guide to Modern Imputation , author =. 2026 , journal =

  90. [98]

    Näf, Jeffrey and Spohn, Meta-Lina and Michel, Loris and Meinshausen, Nicolai , title =

  91. [99]

    Spohn, Meta-Lina and Näf, Jeffrey and Michel, Loris and Meinshausen, Nicolai , journal =

  92. [100]

    2025 , journal =

    Imputation-Powered Inference , author =. 2025 , journal =

  93. [101]

    arXiv preprint arXiv:2310.17434 , year =

    The `Why' behind including `Y' in your imputation model , author =. arXiv preprint arXiv:2310.17434 , year =

  94. [102]

    2021 , journal =

    R-miss-tastic: a unified platform for missing values methods and workflows , author =. 2021 , journal =

  95. [103]

    arXiv preprint arXiv:1902.06931 , year =

    On the consistency of supervised learning with missing values , author =. arXiv preprint arXiv:1902.06931 , year =

  96. [104]

    What's a good imputation to predict with missing values? , volume =

    Le Morvan, Marine and Josse, Julie and Scornet, Erwan and Varoquaux, Gael , booktitle =. What's a good imputation to predict with missing values? , volume =

  97. [105]

    Proceedings of the 39th International Conference on Machine Learning , pages =

    Near-optimal rate of consistency for linear models with missing values , author =. Proceedings of the 39th International Conference on Machine Learning , pages =. 2022 , volume =

  98. [106]

    and Wahl, Sven and Raffler, Julia and others , title =

    Do, Kyla T. and Wahl, Sven and Raffler, Julia and others , title =. Metabolomics , year =

  99. [107]

    BMC Bioinformatics , year =

    Kokla, Marietta and Virtanen, Jyrki and Kolehmainen, Marjukka and Paananen, Jussi and Hanhineva, Kati , title =. BMC Bioinformatics , year =

  100. [108]

    BMC Medical Research Methodology , year =

    Dong, Weinan and Fong, Daniel Yee Tak and Yoon, Jin-sun and Wan, Eric Yuk Fai and Bedford, Laura Elizabeth and Tang, Eric Ho Man and Lam, Cindy Lo Kuen , title =. BMC Medical Research Methodology , year =

  101. [109]

    , title =

    Deng, Grace and Han, Cuize and Matteson, David S. , title =. Data Mining and Knowledge Discovery , year =

  102. [110]

    Statistical Theory and Related Fields , volume =

    Fang Fang and Shenliao Bao , title =. Statistical Theory and Related Fields , volume =

  103. [111]

    , title =

    Troyanskaya, Olga and Cantor, Michael and Sherlock, Gavin and Brown, Pat and Hastie, Trevor and Tibshirani, Robert and Botstein, David and Altman, Russ B. , title =. Bioinformatics , volume =

  104. [112]

    Journal of Machine Learning Research , year =

    Dimitris Bertsimas and Colin Pawlowski and Ying Daisy Zhuo , title =. Journal of Machine Learning Research , year =

  105. [113]

    Applied Artificial Intelligence , volume =

    Anil Jadhav and Dhanya Pramod and Krishnan Ramanathan , title =. Applied Artificial Intelligence , volume =

  106. [114]

    Statistical Analysis and Data Mining , volume =

    Random Forest Missing Data Algorithms , author =. Statistical Analysis and Data Mining , volume =

  107. [115]

    Computational Statistics , year =

    Ramosaj, Burim and Pauly, Markus , title =. Computational Statistics , year =

  108. [116]

    and Mukherjee, Ashin and Singal, Amit G

    Waljee, Akbar K. and Mukherjee, Ashin and Singal, Amit G. and Zhang, Yiwei and Warren, Jeffrey and Balis, Ulysses and Marrero, Jorge and Zhu, Ji and Higgins, Peter , title =. BMJ Open , year =

  109. [117]

    , title =

    Hong, Shangzhi and Lynn, Henry S. , title =. BMC Medical Research Methodology , year =

  110. [118]

    Journal of Statistical Computation and Simulation , volume =

    Rianne Margaretha Schouten and Peter Lugtig and Gerko Vink , title =. Journal of Statistical Computation and Simulation , volume =

  111. [119]

    2014 , author =

    Recursive partitioning for missing data imputation in the presence of interaction effects , journal =. 2014 , author =

  112. [120]

    and Reiter, Jerome P

    Burgette, Lane F. and Reiter, Jerome P. , title =. American Journal of Epidemiology , volume =

  113. [121]

    Journal of Statistical Computation and Simulation , volume =

    Vincent Audigier and François Husson and Julie Josse , title =. Journal of Statistical Computation and Simulation , volume =

  114. [122]

    Multiple imputation in principal component analysis , volume =

    Josse, Julie and Pagès, Jérôme and Husson, Francois , year =. Multiple imputation in principal component analysis , volume =

  115. [123]

    Journal of Statistical Software , year =

    Julie Josse and Fran. Journal of Statistical Software , year =

  116. [124]

    Murray and Jerome P

    Jared S. Murray and Jerome P. Reiter , title =. Journal of the American Statistical Association , volume =

  117. [125]

    Scientific Reports , year =

    Deng, Yi and Chang, Changgee and Ido, Moges Seyoum and Long, Qi , title =. Scientific Reports , year =

  118. [126]

    2018 , booktitle =

    Biessmann, Felix and Salinas, David and Schelter, Sebastian and Schmidt, Philipp and Lange, Dustin , title =. 2018 , booktitle =

  119. [127]

    Pattern Recognition , volume =

    Handling incomplete heterogeneous data using. Pattern Recognition , volume =. 2020 , author =

  120. [128]

    GigaScience , volume =

    Qiu, Yeping Lina and Zheng, Hong and Gevaert, Olivier , title =. GigaScience , volume =

  121. [129]

    Proceedings of the Ninth Asian Conference on Machine Learning , pages =

    Recovering Probability Distributions from Missing Data , author =. Proceedings of the Ninth Asian Conference on Machine Learning , pages =

  122. [130]

    Journal of the Royal Statistical Society: Series B (Statistical Methodology) , year =

    Han, Peisong and Kong, Linglong and Zhao, Jiwei and Zhou, Xingcai , title =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , year =

  123. [131]

    The Annals of Statistics , volume =

    Empirical likelihood-based inference under imputation for missing response data , author =. The Annals of Statistics , volume =. 2002 , publisher =

  124. [132]

    Jing Qin and Biao Zhang and Denis H. Y. Leung , title =. Journal of the American Statistical Association , volume =

  125. [133]

    Communications in Statistics - Theory and Methods , volume =

    Xiaohui Yuan and Xiaogang Dong , title =. Communications in Statistics - Theory and Methods , volume =. 2019 , publisher =

  126. [134]

    Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =

    Biased-sample empirical likelihood weighting for missing data problems: an alternative to inverse probability weighting , author =. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =

  127. [135]

    Electronic Journal of Statistics , volume =

    Parameter estimation through semiparametric quantile regression imputation , author =. Electronic Journal of Statistics , volume =. 2016 , publisher =

  128. [136]

    , title =

    Lichman, M. , title =. 2013 , url =

  129. [137]

    2022 , note =

    energy: E-Statistics: Multivariate Inference via the Energy of Data , author =. 2022 , note =

  130. [138]

    Székely , title =

    Gabor J. Székely , title =. 2003 , volume =

  131. [139]

    2023 , journal =

    Prediction with Incomplete Data under Agnostic Mask Distribution Shift , author =. 2023 , journal =

  132. [140]

    , title =

    Zhu, Ziwei and Wang, Tengyao and Samworth, Richard J. , title =. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =

  133. [141]

    Verchand and Thomas B

    Tianyi Ma and Kabir A. Verchand and Thomas B. Berrett and Tengyao Wang and Richard J. Samworth , journal =. Estimation beyond

  134. [142]

    Advances in Neural Information Processing Systems , volume =

    Graphical Models for Inference with Missing Data , author =. Advances in Neural Information Processing Systems , volume =

  135. [143]

    Artificial Intelligence and Statistics , pages =

    Karthika Mohan and Judea Pearl , title =. Artificial Intelligence and Statistics , pages =. 2014 , note =

  136. [144]

    Journal of the American Statistical Association , pages =

    Karthika Mohan and Judea Pearl , title =. Journal of the American Statistical Association , pages =

  137. [145]

    Proceedings of the Twenty Sixth International Conference on Systems Inference , year =

    Karthika Mohan and Judea Pearl and Tian Jin , title =. Proceedings of the Twenty Sixth International Conference on Systems Inference , year =

  138. [146]

    Uncertainty in Artificial Intelligence , pages =

    Razieh Nabi and Rohit Bhattacharya , title =. Uncertainty in Artificial Intelligence , pages =. 2023 , note =

  139. [147]

    Proceedings of the Twenty Seventh International Conference on Machine Learning (ICML-20) , year =

    Razieh Nabi and Rohit Bhattacharya and Ilya Shpitser , title =. Proceedings of the Twenty Seventh International Conference on Machine Learning (ICML-20) , year =

  140. [148]

    Statistica Sinica , year =

    Razieh Nabi and Rohit Bhattacharya and Ilya Shpitser and James Robins , title =. Statistica Sinica , year =

  141. [149]

    2015 , note =

    Ilya Shpitser and Karthika Mohan and Judea Pearl , title =. 2015 , note =

  142. [150]

    Consistent Estimation of Functions of Data Missing Non-Monotonically and Not at Random , volume =

    Shpitser, Ilya , booktitle =. Consistent Estimation of Functions of Data Missing Non-Monotonically and Not at Random , volume =

  143. [151]

    Proceedings of The 35th Uncertainty in Artificial Intelligence Conference , pages =

    Identification In Missing Data Models Represented By Directed Acyclic Graphs , author =. Proceedings of The 35th Uncertainty in Artificial Intelligence Conference , pages =. 2020 , volume =

  144. [152]

    Daniel and Michael G

    Rhian M. Daniel and Michael G. Kenward and Simon N. Cousens and Bianca L. Using causal diagrams to guide analysis in missing data problems , journal =

  145. [153]

    Parametric

    Ch. Parametric. arXiv preprint arXiv:2503.00448 , year =

  146. [154]

    Asymptotics of Nonparametric Estimation under general non-monotone

    N. Asymptotics of Nonparametric Estimation under general non-monotone. arXiv preprint arXiv:2603.23449 , year =

  147. [155]

    Generative Modeling under Non-Monotone

    Gitte Kremling and N. Generative Modeling under Non-Monotone. arXiv preprint arXiv:2604.04567 , year =

  148. [156]

    arXiv preprint arXiv:2501.16174 , year =

    Measuring Heterogeneity in Machine Learning with Distributed Energy Distance , author =. arXiv preprint arXiv:2501.16174 , year =

  149. [157]

    Statistical Science , volume =

    Introduction to Double Robust Methods for Incomplete Data , author =. Statistical Science , volume =

  150. [158]

    2022 , volume =

    Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression , journal =. 2022 , volume =

  151. [159]

    Proceedings of The 27th International Conference on Artificial Intelligence and Statistics , pages =

    B\'. Proceedings of The 27th International Conference on Artificial Intelligence and Statistics , pages =. 2024 , volume =

  152. [160]

    Journal of Machine Learning Research , year =

    Jeffrey Näf and Corinne Emmenegger and Peter Bühlmann and Nicolai Meinshausen , title =. Journal of Machine Learning Research , year =

  153. [161]

    Loris Michel and Domagoj. drf:. 2021 , note =

  154. [162]

    Machine Learning , volume =

    Random forests , author =. Machine Learning , volume =. 2001 , publisher =

  155. [163]

    1984 , publisher =

    Classification and regression trees , author =. 1984 , publisher =

  156. [164]

    Manual---setting up, using and understanding random forests V4.0 , author =

  157. [165]

    The Annals of Statistics , number =

    Susan Athey and Julie Tibshirani and Stefan Wager , title =. The Annals of Statistics , number =

  158. [166]

    Julie Tibshirani and Susan Athey and Erik Sverdrup and Stefan Wager , year =. grf:

  159. [167]

    Journal of the American Statistical Association , volume =

    Estimation and inference of heterogeneous treatment effects using random forests , author =. Journal of the American Statistical Association , volume =. 2018 , publisher =

  160. [168]

    2017 , journal =

    Estimation and Inference of Heterogeneous Treatment Effects using Random Forests , author =. 2017 , journal =

  161. [169]

    Journal of Machine Learning Research , volume =

    Quantile regression forests , author =. Journal of Machine Learning Research , volume =

  162. [170]

    Biostatistics , volume =

    Survival ensembles , author =. Biostatistics , volume =. 2006 , publisher =

  163. [171]

    Journal of Computational and Graphical Statistics , volume =

    Torsten Hothorn and Achim Zeileis , title =. Journal of Computational and Graphical Statistics , volume =. 2021 , publisher =

  164. [172]

    Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery , volume =

    Multivariate random forests , author =. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery , volume =. 2011 , publisher =

  165. [173]

    Journal of Machine Learning Research , year =

    Lucas Mentch and Giles Hooker , title =. Journal of Machine Learning Research , year =

  166. [174]

    Journal of Machine Learning Research , volume =

    Analysis of a random forests model , author =. Journal of Machine Learning Research , volume =

  167. [175]

    Journal of the American Statistical Association , volume =

    Random forests and adaptive nearest neighbors , author =. Journal of the American Statistical Association , volume =. 2006 , publisher =

  168. [176]

    1055) , author =

    Random forests and adaptive nearest neighbors (Technical Report No. 1055) , author =. University of Wisconsin , year =

  169. [177]

    arXiv preprint arXiv:1405.0352 , year =

    Asymptotic theory for random forests , author =. arXiv preprint arXiv:1405.0352 , year =

  170. [178]

    Journal of Computational and Graphical Statistics , volume =

    Model-based recursive partitioning , author =. Journal of Computational and Graphical Statistics , volume =. 2008 , publisher =

  171. [179]

    and Ziegler, Andreas , journal =

    Wright, Marvin N. and Ziegler, Andreas , journal =

  172. [180]

    R package version , volume =

    RandomForestSRC: Random forests for survival, regression and classification (RF-SRC) , author =. R package version , volume =

  173. [181]

    Journal of Computational and Graphical Statistics , volume =

    Tao Shi and Steve Horvath , title =. Journal of Computational and Graphical Statistics , volume =

  174. [182]

    The Annals of Applied Statistics , year =

    Ishwaran, Hemant and Kogalur, Udaya and Blackstone, Eugene and Lauer, Michael , title =. The Annals of Applied Statistics , year =

  175. [183]

    Node harvest , volume =

    Meinshausen, Nicolai , year =. Node harvest , volume =. The Annals of Applied Statistics , publisher =

  176. [184]

    European Conference on Machine Learning , pages =

    Ensembles of multi-objective decision trees , author =. European Conference on Machine Learning , pages =. 2007 , organization =

  177. [185]

    International Conference on Machine Learning , pages =

    Narrowing the gap: Random forests in theory and in practice , author =. International Conference on Machine Learning , pages =

  178. [186]

    Probability Machines Consistent Probability Estimation Using Nonparametric Learning Machines , volume =

    Malley, James and Kruppa, J and Dasgupta, Abhijit and Malley, Karen and Ziegler, Andreas , year =. Probability Machines Consistent Probability Estimation Using Nonparametric Learning Machines , volume =

  179. [187]

    arXiv preprint arXiv:1804.05753 , year =

    Rfcde: Random forests for conditional density estimation , author =. arXiv preprint arXiv:1804.05753 , year =

  180. [188]

    Rosenthal , title =

    Cédric Beaulac and Jeffrey S. Rosenthal , title =. Computational Statistics , year =

  181. [189]

    Journal of Machine Learning Research , pages =

    Saar-Tsechansky, Maytal and Provost, Foster , title =. Journal of Machine Learning Research , pages =. 2007 , volume =

  182. [190]

    Ross , publisher =

    Quinlan, J. Ross , publisher =. C4.5: programs for machine learning , year =

  183. [191]

    Geaur and Islam, Md , year =

    Rahman, Md. Geaur and Islam, Md , year =. Missing Value Imputation Using Decision Trees and Decision Forests by Splitting and Merging Records: Two Novel Techniques , volume =

  184. [192]

    Applied Artificial Intelligence , volume =

    Bhekisipho Twala , title =. Applied Artificial Intelligence , volume =

  185. [193]

    Principles of Data Mining and Knowledge Discovery , year =

    Ad Feelders , title =. Principles of Data Mining and Knowledge Discovery , year =

  186. [194]

    The Annals of Statistics , number =

    Peter Bühlmann and Bin Yu , title =. The Annals of Statistics , number =

  187. [195]

    2009 , author =

    Standard errors for bagged and random forest estimators , journal =. 2009 , author =

  188. [196]

    Holland , title =

    Paul W. Holland , title =. Journal of the American Statistical Association , volume =

  189. [197]

    Journal of the American Statistical Association , volume =

    Donald B Rubin , title =. Journal of the American Statistical Association , volume =. 2005 , publisher =

  190. [198]

    and Rubin, Donald B

    Rosenbaum, Paul R. and Rubin, Donald B. , title =. Biometrika , volume =

  191. [199]

    2009 , publisher =

    Causality , author =. 2009 , publisher =

  192. [200]

    Biometrika , volume =

    Causal diagrams for empirical research , author =. Biometrika , volume =. 1995 , publisher =

  193. [201]

    2023 , publisher =

    Stefan Wager and Jonathan Taylor , title =. 2023 , publisher =

  194. [202]

    2018 , publisher =

    Double/debiased machine learning for treatment and structural parameters , author =. 2018 , publisher =

  195. [203]

    Econometrica , volume =

    Inference on counterfactual distributions , author =. Econometrica , volume =

  196. [204]

    Journal of Econometrics , year =

    Network and panel quantile effects via distribution regression , author =. Journal of Econometrics , year =

  197. [205]

    Journal of the Royal Statistical Society: Series B, Statistical Methodology , volume =

    Nonparametric methods for doubly robust estimation of continuous treatment effects , author =. Journal of the Royal Statistical Society: Series B, Statistical Methodology , volume =. 2017 , publisher =

  198. [206]

    Review of Economics and Statistics , volume =

    Nonparametric estimation of average treatment effects under exogeneity: A review , author =. Review of Economics and Statistics , volume =. 2004 , publisher =

  199. [207]

    Econometrica , volume =

    Large sample properties of matching estimators for average treatment effects , author =. Econometrica , volume =. 2006 , publisher =

  200. [208]

    Proceedings of the National Academy of Sciences , volume =

    Metalearners for estimating heterogeneous treatment effects using machine learning , author =. Proceedings of the National Academy of Sciences , volume =. 2019 , publisher =

  201. [209]

    Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

    Robust and Agnostic Learning of Conditional Distributional Treatment Effects , author =. Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =. 2023 , volume =

  202. [210]

    Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =

    Cui, Yifan and Kosorok, Michael R and Sverdrup, Erik and Wager, Stefan and Zhu, Ruoqing , title =. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =

  203. [211]

    Journal of Machine Learning Research , year =

    Krikamol Muandet and Motonobu Kanagawa and Sorawit Saengkyongam and Sanparith Marukatat , title =. Journal of Machine Learning Research , year =

  204. [212]

    An Efficient Doubly-Robust Test for the Kernel Treatment Effect , volume =

    Martinez Taboada, Diego and Ramdas, Aaditya and Kennedy, Edward , booktitle =. An Efficient Doubly-Robust Test for the Kernel Treatment Effect , volume =

  205. [213]

    2022 , journal =

    Doubly Robust Kernel Statistics for Testing Distributional Treatment Effects Even Under One Sided Overlap , author =. 2022 , journal =

  206. [214]

    Conditional Distributional Treatment Effect with Kernel Conditional Mean Embeddings and

    Park, Junhyung and Shalit, Uri and Sch. Conditional Distributional Treatment Effect with Kernel Conditional Mean Embeddings and. Proceedings of the 38th International Conference on Machine Learning (ICML) , volume =

  207. [215]

    Proceedings of the Thirty-Eighth Conference on Uncertainty in Artificial Intelligence , pages =

    Feature selection for discovering distributional treatment effect modifiers , author =. Proceedings of the Thirty-Eighth Conference on Uncertainty in Artificial Intelligence , pages =. 2022 , volume =

  208. [216]

    2023 , eprint =

    Variable importance for causal forests: breaking down the heterogeneity of treatment effects , author =. 2023 , eprint =

  209. [217]

    arXiv preprint arXiv:2204.06030 , year =

    Variable importance measures for heterogeneous causal effects , author =. arXiv preprint arXiv:2204.06030 , year =

  210. [218]

    Yuchen Ma and Valentyn Melnychuk and Jonas Schweisthal and Stefan Feuerriegel , journal =

  211. [219]

    arXiv preprint arXiv:2110.01664 , year =

    Estimating Potential Outcome Distributions with Collaborating Causal Networks , author =. arXiv preprint arXiv:2110.01664 , year =

  212. [220]

    arXiv preprint arXiv:2010.04855 , year =

    Reproducing Kernel Methods for Nonparametric and Semiparametric Treatment Effects , author =. arXiv preprint arXiv:2010.04855 , year =

  213. [221]

    Treatment effects beyond the mean using distributional regression: Methods and guidance , year =

    Hohberg, Maike and Pütz, Peter and Kneib, Thomas , journal =. Treatment effects beyond the mean using distributional regression: Methods and guidance , year =

  214. [222]

    A Short Survey on Forest Based Heterogeneous Treatment Effect Estimation Methods: Meta-learners and Specific Models , year =

    Hao Jiang and Peng Qi and Jingying Zhou and Jack Zhou and Sharath Rao , journal =. A Short Survey on Forest Based Heterogeneous Treatment Effect Estimation Methods: Meta-learners and Specific Models , year =

  215. [223]

    Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , pages =

    Nonparametric Estimation of Heterogeneous Treatment Effects: From Theory to Learning Algorithms , author =. Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , pages =. 2021 , volume =

  216. [224]

    Klebanov and Ingmar Schuster and T

    I. Klebanov and Ingmar Schuster and T. Sullivan , journal =. A Rigorous Theory of Conditional Mean Embeddings , year =

  217. [225]

    Groll and T

    Guillermo Briseño Sanchez and Maike Hohberg and A. Groll and T. Kneib , journal =. Flexible instrumental variable distributional regression , year =

  218. [226]

    Luo and C

    Guojun Wu and Ge Song and Xiaoxiang Lv and S. Luo and C. Shi and Hongtu Zhu , journal =. DNet: Distributional Network for Distributional Individualized Treatment Effects , year =

  219. [227]

    Frontiers in Genetics , volume =

    Conditional Generative Adversarial Networks for Individualized Treatment Effect Estimation and Treatment Selection , author =. Frontiers in Genetics , volume =

  220. [228]

    and Cole, Stephen R

    Westreich, Daniel and Edwards, Jessie K. and Cole, Stephen R. and Platt, Robert W. and Mumford, Sunni L. and Schisterman, Enrique F. , title =. International Journal of Epidemiology , year =

  221. [229]

    The American Economic Review , volume =

    Evaluating the econometric evaluations of training programs with experimental data , author =. The American Economic Review , volume =

  222. [230]

    Dehejia and Sadek Wahba , title =

    Rajeev H. Dehejia and Sadek Wahba , title =. Journal of the American Statistical Association , volume =. 1999 , publisher =

  223. [231]

    2003 , author =

    Does 401(k) eligibility increase saving?: Evidence from propensity score subclassification , journal =. 2003 , author =

  224. [232]

    The Review of Economics and Statistics , volume =

    Chernozhukov, Victor and Hansen, Christian , title =. The Review of Economics and Statistics , volume =

  225. [233]

    Electronic Journal of Statistics , volume =

    Marginal integration for nonparametric causal inference , author =. Electronic Journal of Statistics , volume =. 2015 , publisher =

  226. [234]

    Fair Data Adaptation with Quantile Preservation , journal =

    Drago Ple. Fair Data Adaptation with Quantile Preservation , journal =. 2020 , volume =

  227. [235]

    Advances in Neural Information Processing Systems , pages =

    Counterfactual fairness , author =. Advances in Neural Information Processing Systems , pages =

  228. [236]

    Advances in Neural Information Processing Systems , pages =

    Avoiding discrimination through causal reasoning , author =. Advances in Neural Information Processing Systems , pages =

  229. [237]

    arXiv preprint arXiv:2502.11820 , year =

    A Diagnostic to Find and Help Combat Stochastic Positivity Issues -- with a Focus on Continuous Treatments , author =. arXiv preprint arXiv:2502.11820 , year =

  230. [238]

    Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics , pages =

    Characterization of Overlap in Observational Studies , author =. Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics , pages =. 2020 , volume =

  231. [239]

    Journal of Machine Learning Research , volume =

    A kernel two-sample test , author =. Journal of Machine Learning Research , volume =

  232. [240]

    Advances in Neural Information Processing Systems , pages =

    A kernel method for the two-sample-problem , author =. Advances in Neural Information Processing Systems , pages =

  233. [241]

    , title =

    Gretton, Arthur and Sejdinovic, Dino and Strathmann, Heiko and Balakrishnan, Sivaraman and Pontil, Massimiliano and Fukumizu, Kenji and Sriperumbudur, Bharath K. , title =. Advances in Neural Information Processing Systems , publisher =. 2012 , volume =

  234. [242]

    Advances in Neural Information Processing Systems , pages =

    Optimal kernel choice for large-scale two-sample tests , author =. Advances in Neural Information Processing Systems , pages =

  235. [243]

    Advances in Neural Information Processing Systems , publisher =

    Chwialkowski, Kacper P and Ramdas, Aaditya and Sejdinovic, Dino and Gretton, Arthur , title =. Advances in Neural Information Processing Systems , publisher =. 2015 , volume =

  236. [244]

    Advances in Neural Information Processing Systems , publisher =

    Jitkrittum, Wittawat and Szab\'. Advances in Neural Information Processing Systems , publisher =. 2016 , volume =

  237. [245]

    Advances in Neural Information Processing Systems , pages =

    B-test: A non-parametric, low variance kernel two-sample test , author =. Advances in Neural Information Processing Systems , pages =

  238. [246]

    Journal of Machine Learning Research , volume =

    Hilbert Space Embeddings and Metrics on Probability Measures , author =. Journal of Machine Learning Research , volume =

  239. [247]

    Bernoulli , number =

    Bharath Sriperumbudur , title =. Bernoulli , number =

  240. [248]

    Metrizing Weak Convergence with Maximum Mean Discrepancies , journal =

    Carl-Johann Simon-Gabriel and Alessandro Barp and Bernhard Sch. Metrizing Weak Convergence with Maximum Mean Discrepancies , journal =. 2023 , volume =

  241. [249]

    Journal of Machine Learning Research , volume =

    Kernel distribution embeddings: Universal kernels, characteristic kernels and kernel metrics on distributions , author =. Journal of Machine Learning Research , volume =

  242. [250]

    Foundations and Trends in Machine Learning , title =

    Krikamol Muandet and Kenji Fukumizu and Bharath Sriperumbudur and Bernhard Sch\". Foundations and Trends in Machine Learning , title =. 2017 , volume =

  243. [251]

    The Annals of Statistics , number =

    Dino Sejdinovic and Bharath Sriperumbudur and Arthur Gretton and Kenji Fukumizu , title =. The Annals of Statistics , number =

  244. [252]

    Ji Zhao and Deyu Meng , journal =. Fast. 2015 , publisher =

  245. [253]

    Journal of Multivariate Analysis , volume =

    On a new multivariate two-sample test , author =. Journal of Multivariate Analysis , volume =. 2004 , publisher =

  246. [254]

    and Rafsky, Lawrence C , journal =

    Friedman, Jerome H. and Rafsky, Lawrence C , journal =. Multivariate generalizations of the. 1979 , publisher =

  247. [255]

    InterStat , volume =

    Testing for equal distributions in high dimension , author =. InterStat , volume =

  248. [256]

    Statistica Sinica , pages =

    Effect of high dimension: by an example of a two sample problem , author =. Statistica Sinica , pages =. 1996 , publisher =

  249. [257]

    Epps and Kenneth J

    T.W. Epps and Kenneth J. Singleton , title =. Journal of Statistical Computation and Simulation , volume =. 1986 , publisher =

  250. [258]

    arXiv preprint arXiv:1406.2083 , year =

    On the decreasing power of kernel and distance based nonparametric hypothesis tests in high dimensions , author =. arXiv preprint arXiv:1406.2083 , year =

  251. [259]

    , title =

    Rustamov, Raif M. , title =. Stat , volume =

  252. [260]

    arXiv preprint arXiv:2110.13452 , year =

    Alon Itai and Amir Globerson and Ami Wiesel , title =. arXiv preprint arXiv:2110.13452 , year =

  253. [261]

    Sriperumbudur and Kenji Fukumizu and Arthur Gretton and Bernhard Schölkopf and Gert R

    Bharath K. Sriperumbudur and Kenji Fukumizu and Arthur Gretton and Bernhard Schölkopf and Gert R. G. Lanckriet , title =. Electronic Journal of Statistics , publisher =

  254. [262]

    A Kernel Statistical Test of Independence , year =

    Gretton, Arthur and Fukumizu, Kenji and Teo, Choon Hui and Song, Le and Sch\". A Kernel Statistical Test of Independence , year =. Advances in Neural Information Processing Systems , volume =

  255. [263]

    arXiv preprint arXiv:1202.3775 , year =

    Kernel-based conditional independence test and application in causal discovery , author =. arXiv preprint arXiv:1202.3775 , year =

  256. [264]

    Entropy , year =

    Statistical Analysis of Distance Estimators with Density Differences and Density Ratios , author =. Entropy , year =

  257. [265]

    , title =

    Csiszar, I. , title =. Studia Scientiarum Mathematicarum Hungarica , year =

  258. [266]

    A Measure-Theoretic Approach to Kernel Conditional Mean Embeddings , volume =

    Park, Junhyung and Muandet, Krikamol , booktitle =. A Measure-Theoretic Approach to Kernel Conditional Mean Embeddings , volume =

  259. [267]

    Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics , pages =

    Kernel Conditional Density Operators , author =. Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics , pages =. 2020 , volume =

  260. [268]

    Kernel Embeddings of Conditional Distributions: A Unified Kernel Framework for Nonparametric Inference in Graphical Models , year =

    Song, Le and Fukumizu, Kenji and Gretton, Arthur , journal =. Kernel Embeddings of Conditional Distributions: A Unified Kernel Framework for Nonparametric Inference in Graphical Models , year =

  261. [269]

    2009 , booktitle =

    Song, Le and Huang, Jonathan and Smola, Alex and Fukumizu, Kenji , title =. 2009 , booktitle =

  262. [270]

    ICML , title =

    Gr\". ICML , title =

  263. [271]

    Journal of Machine Learning Research , year =

    Kenji Fukumizu and Le Song and Arthur Gretton , title =. Journal of Machine Learning Research , year =

  264. [272]

    Kernel Instrumental Variable Regression , volume =

    Singh, Rahul and Sahani, Maneesh and Gretton, Arthur , booktitle =. Kernel Instrumental Variable Regression , volume =

  265. [273]

    Lachlan McCalman and Simon Timothy O'Callaghan and Fabio Ramos , title =. 2013

  266. [274]

    Constructive Approximation , year =

    Minh, Ha Quang , title =. Constructive Approximation , year =

  267. [275]

    2014 , author =

    Characterization of Gaussian distribution on a Hilbert space from samples of random size , journal =. 2014 , author =

  268. [276]

    2019 , volume =

    Foundations and Trends in Machine Learning , title =. 2019 , volume =

  269. [277]

    Ramdas, Aaditya and Trillos, Nicol\'. On. Entropy , volume =. 2017 , number =

  270. [278]

    Neural Computation , volume =

    Imaizumi, Masaaki and Ota, Hirofumi and Hamaguchi, Takuo , title =. Neural Computation , volume =

  271. [279]

    Sliced and Radon

    Bonneel, Nicolas and Rabin, Julien and Peyr. Sliced and Radon. Journal of Mathematical Imaging and Vision , year =

  272. [280]

    Proceedings of the 34th International Conference on Machine Learning , pages =

    Martin Arjovsky and Soumith Chintala and L. Proceedings of the 34th International Conference on Machine Learning , pages =. 2017 , volume =

  273. [281]

    Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics , pages =

    Optimal Transport for Multi-source Domain Adaptation under Target Shift , author =. Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics , pages =. 2019 , volume =

  274. [282]

    Wasserstein variational inference , year =

    Ambrogioni, Luca and G\". Wasserstein variational inference , year =. Advances in Neural Information Processing Systems , volume =

  275. [283]

    International Conference on Machine Learning , pages =

    Minimizing f -Divergences by Interpolating Velocity Fields , author =. International Conference on Machine Learning , pages =. 2024 , organization =

  276. [284]

    An invitation to optimal transport,

    Figalli, Alessio and Glaudo, Federico , volume=. An invitation to optimal transport,. 2021 , publisher=

  277. [285]

    Minimax estimation of smooth densities in

    Niles-Weed, Jonathan and Berthet, Quentin , journal=. Minimax estimation of smooth densities in. 2022 , publisher=

  278. [286]

    arXiv preprint arXiv:2510.17608 , year=

    Non-asymptotic error bounds for probability flow ODEs under weak log-concavity , author=. arXiv preprint arXiv:2510.17608 , year=

  279. [287]

    Fundamentals of Nonparametric Bayesian Inference , publisher =

    Ghosal, Subhashis and van der Vaart, Aad , year =. Fundamentals of Nonparametric Bayesian Inference , publisher =

  280. [288]

    Castillo, Isma. On the. The Annals of Statistics , volume =. 2014 , publisher =

  281. [289]

    The Annals of Statistics , volume =

    Convergence rates of variational posterior distributions , author =. The Annals of Statistics , volume =. 2020 , publisher =

  282. [290]

    On the minimax optimality of

    Kunkel, Lea and Trabs, Mathias , year =. On the minimax optimality of

  283. [291]

    Asymptotic Statistics , publisher =

    Aad van der Vaart , year =. Asymptotic Statistics , publisher =

  284. [292]

    2008 , publisher =

    Introduction to Empirical Processes and Semiparametric Inference , author =. 2008 , publisher =

  285. [293]

    1993 , publisher =

    Efficient and adaptive estimation for semiparametric models , author =. 1993 , publisher =

  286. [294]

    Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume =

    Local maximum likelihood estimation and inference , author =. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume =. 1998 , publisher =

  287. [295]

    and Ritov, Ya'acov , title =

    Robins, James M. and Ritov, Ya'acov , title =. Statistics in Medicine , volume =

  288. [296]

    Nadaraya, E. A. , title =. Theory of Probability & Its Applications , volume =

  289. [297]

    Watson , journal =

    Geoffrey S. Watson , journal =. Smooth Regression Analysis , volume =

  290. [298]

    , publisher =

    Wahba, G. , publisher =. Spline Models for Observational Data , year =

  291. [299]

    Journal of the American Statistical Association , volume =

    Robust locally weighted regression and smoothing scatterplots , author =. Journal of the American Statistical Association , volume =. 1979 , publisher =

  292. [300]

    1986 , publisher =

    Density estimation for statistics and data analysis , author =. 1986 , publisher =

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.