A unified U-statistic framework tests equality of smooth parameters across k populations via Wald and ANOVA statistics, with fixed-d asymptotics, weighted bootstrap, and normal limits when d grows slower than n.
Two-Sample Testing with Missing Data via Energy Distance: Weighting and Imputation Approaches
2 Pith papers cite this work. Polarity classification is still indexing.
abstract
In this paper, we address the problem of two-sample testing in the presence of missing data under a variety of missingness mechanisms. Our focus is on the well-known energy distance-based two-sample test. In addition to the standard complete-case approach, we propose a modification of the test statistic that incorporates all available data, utilizing appropriate weights. The asymptotic null distribution of the test statistic is derived and two resampling procedures for approximating the corresponding p-values are proposed. We also propose a new bootstrap method specifically designed for a test statistic based on samples completed via common imputation methods. Through an extensive simulation study, we compare all methods in terms of type I error control and statistical power across a set of sample sizes, dimensions, distributions, missingness mechanisms, and missingness rates. Based on these results, we provide general recommendations for each considered scenario.
fields
stat.ME 2years
2026 2representative citing papers
Two new modifications to a Kendall tau-based test are proposed and analyzed for independence testing in high-dimensional data with missing observations, backed by theory and simulations.
citing papers explorer
-
Testing the equality of estimable parameters
A unified U-statistic framework tests equality of smooth parameters across k populations via Wald and ANOVA statistics, with fixed-d asymptotics, weighted bootstrap, and normal limits when d grows slower than n.
-
Testing independence in the presence of missing data: high-dimensional case
Two new modifications to a Kendall tau-based test are proposed and analyzed for independence testing in high-dimensional data with missing observations, backed by theory and simulations.