Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

A BHEP test for multivariate normality on incomplete data

T0 review · 2 major / 4 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read A characteristic-function test detects multivariate normality even when some data entries are missing, without needing a non-singular covariance estimate.

desk verdict Solid, usable incomplete-data BHEP with correct asymptotics under MCAR; practical advance, not a conceptual leap. read the letter →

arxiv 2607.03335 v1 pith:V36QJHQO submitted 2026-07-03 math.ST stat.TH

classification math.STstat.TH MSC 62G1062G09
keywords BHEPtestmultivariatenormalityincompletedatacharacteristicfunctionbootstrapmissingnessmechanism
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Standard tests for multivariate normality break down when observations have missing entries: discarding incomplete vectors wastes information and can produce singular covariance estimates, while naive imputation distorts the distribution. This paper builds a BHEP-style test that works directly with the observed incomplete vectors. It estimates means, covariances and missingness probabilities from the incomplete sample, then measures the L2 distance between the empirical characteristic function of the (projected, centered) incomplete data and the characteristic function that would hold under normality. The procedure remains valid even when the estimated covariance is singular. As sample size grows, the test statistic converges almost surely to zero under normality and to a positive constant under alternatives; under the null it converges in distribution to a known functional of a Gaussian process. Critical values are obtained by a parametric bootstrap that resamples both the data and the missingness pattern. Simulations and an air-quality data example show that the test maintains its size and has power against common alternatives even at high missing rates.

What carries the argument

The projected incomplete-data BHEP statistic Tn,π = n ∫ |φ̂n,π − ϕ̂n,π|^{2} dΦdπ, whose almost-sure limit and null limiting distribution are controlled by the influence functions of the incomplete-data estimators of mean, covariance and missingness probabilities.

What would settle it

Generate data from a non-normal multivariate distribution with a known missingness mechanism that is independent of the values, apply the proposed bootstrap test at nominal level 0.05 for large n, and check whether the empirical rejection rate stays near 1; if it stays near 0.05 the consistency claim is false.

Watch

Extended reading notes

Core claim

Under the stated regularity and independence conditions, the weighted sum of projected BHEP statistics based on incomplete data converges almost surely to a non-negative constant that vanishes if and only if the underlying distribution is multivariate normal, and converges in distribution under the null to the squared L2-norm of an explicitly characterized centered Gaussian process; a parametric bootstrap that redraws both data and missingness indicators therefore yields asymptotically valid critical values.

Load-bearing premise

The missingness indicators must be independent of the data values themselves and complete observations must occur with positive probability; if missingness depends on the unobserved values, the identification of the law of X fails.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper constructs a BHEP-type test for multivariate normality when data vectors may be incomplete. Under MCAR (X independent of the missingness indicator I) and P(I=1_d)>0, it proposes consistent estimators of μ, Σ and the missingness pmf p (Example 2, Theorem 1), forms a weighted sum of L^{2} distances between empirical and parametric characteristic functions of all coordinate projections (Section 4), and proves that (1/n)T_n o κ a.s. with κ=0 under H0 and κ>0 under H1 (Theorem 2) and that T_n converges in distribution under H0 to the squared L^{2}-norm of a centered Gaussian process whose covariance is given by the influence functions of the estimators (Theorem 3). Critical values are obtained by parametric bootstrap; size and power are examined by Monte-Carlo simulation and the procedure is illustrated on the airquality data.

Significance. The contribution fills a genuine gap: standard BHEP procedures become inapplicable or biased under missingness (singular covariance estimates, complete-case loss of power, imputation bias), while existing incomplete-data normality tests target only skewness or kurtosis. The paper supplies a fully specified, always-applicable characteristic-function test, explicit influence-function asymptotics under singular covariance estimates, and a practical bootstrap. The proofs (Appendix) rest on standard Hilbert-space CLTs, Taylor expansions and continuous mapping and appear carefully executed; the simulations (2000 replications, 199 bootstrap draws) and real-data example make the method immediately usable. The free parameters (projection weights, missingness family L_I) are declared and do not undermine the asymptotic claims.

major comments (2)
  1. Assumption 2 (X ⊥ I and p(1_d)>0) is load-bearing for identification of L(X) from L(I,I⊙X) and therefore for both consistency (Theorem 2) and the bootstrap. The manuscript states the assumption clearly and conditions all theorems on it, but never discusses the practical consequences of its violation (MNAR). A short paragraph in Section 2 or 6 acknowledging that the procedure is invalid under MNAR, and perhaps a brief simulation under a simple MNAR mechanism, would make the scope of the claims transparent without altering the theorems.
  2. Section 7, Table 1 (c=1 rows) and the accompanying text: the authors note that their test is more conservative and less powerful than the imputation-based BHEP of Aleksić & Milošević (2024) under low missingness, yet claim superiority under high missingness. The comparison is only qualitative and the competing procedure is not re-implemented under the same high-λ settings. Either a quantitative head-to-head under identical designs or a clearer statement that the two procedures target different regimes would strengthen the empirical claims.
minor comments (4)
  1. Abstract and title: “hypotesis” → “hypothesis”; several other typos appear throughout (e.g., “incompled”, “meachanism”, “demostrate”).
  2. Section 4: the closed-form expression for T_{n,π} involves expectations of exp(-|W|^{2}/2) that are moment-generating functions of generalized chi-squared laws; a one-line pointer to the explicit formula (or a stable numerical method) would aid reproducibility.
  3. Notation: the same symbol Φ_{d_π} is used both for the standard normal measure and, later, for the product measure; a brief clarification would avoid confusion.
  4. References: the arXiv identifier of the present paper appears as 2607.03335 (future date); this is presumably a placeholder and should be corrected on publication.

Circularity Check

1 steps flagged · score 1.0 of 10

No significant circularity: main a.s. limit and null weak-convergence results are derived from first principles under explicit assumptions; only a minor non-load-bearing self-citation appears for a secondary power argument.

  1. self citation load bearing [Section 6 (bootstrap consistency) and proof of Theorem 2 (final paragraph)]
    "using the same arguments as in (the proof of) Theorem 3 in Gaigall & Wübbolding (2025), we can combine Theorem 2 and Theorem 3 to obtain P(Tn>cn,1-α) o1 as n o∞ under the alternative H1. … see (the proof of) Theorem 3 in Gaigall & Wübbolding (2025) for details."

    A secondary claim (consistency of the level-α test under alternatives) is justified solely by reference to a proof in the authors’ own prior paper rather than being re-derived. The citation is not load-bearing for the paper’s main asymptotic theorems, which stand independently.

full rationale

The core claims (Theorem 2: (1/n)Tn o \kappa a.s. with \kappa=0 under H0 and \kappa>0 under H1; Theorem 3: Tn converges in distribution under H0 to the L2-norm of a centered Gaussian process whose covariance is the second-moment of the influence functions \Psi) are obtained by standard arguments that are fully written out: SLLN + continuity + dominated convergence for the a.s. limit (including the elementary identification that L(I,I⊙X) determines L(X) under Assumption 2), and second-order Taylor linearization of the empirical/parametric characteristic functions + Hilbert-space CLT + continuous mapping for the null limit. The influence functions are taken from the asymptotic expansions of the estimators (Assumption 4), which themselves are verified for the concrete estimators of Example 2 by SLLN/CLT/Slutsky/delta-method (Theorem 1 and Lemma 1). No free parameter is fitted to the same data that is later called a “prediction”; the bootstrap is ordinary parametric resampling under the fitted null Nd(etân,etân). The single self-citation (to the authors’ own 2025 Hilbert-space paper) is used only to import a continuity argument that ho>0 when the characteristic functions differ and to claim that the same arguments yield power o1; it is not required for the statements or proofs of Theorems 2–3 themselves. Hence the derivation chain does not reduce to its own inputs by construction.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The central asymptotic claims rest on four explicitly numbered assumptions (finite second moments and positive-definite covariance; independence of data and missingness with positive complete-observation probability; strong consistency of the estimators; asymptotic linear expansions of the estimators) plus the classical Hilbert-space CLT. No free parameters are fitted to obtain the theoretical limit; the only free choices are the projection weights wπ (set equal in the simulations) and the missingness family LI, both of which are user-specified rather than estimated from the target null.

free parameters (2)
  • projection weights wπ
    User-chosen non-negative weights summing to 1 with w_id > 0; equal weights used in all simulations and the real-data example. They affect finite-sample power but not the asymptotic validity statements.
  • missingness family LI
    The analyst must specify which missingness mechanisms are allowed (known, unrestricted, i.i.d. Bernoulli, monotone). The choice is not estimated from data and enters the definition of the estimators bp_n.
assumptions (5)
  • domain assumption E(|X|^2) < ∞ and Cov(X) positive definite (Assumption 1)
    Standard second-moment condition guaranteeing existence of mean and covariance; invoked throughout Sections 2–5.
  • domain assumption X ⊥ I and P(I = 1_d) > 0 (Assumption 2)
    Independence of data and missingness (MCAR) plus positive probability of complete observation; used to identify L(X) from the observed pairs (I, I⊙X).
  • domain assumption Strong consistency of bμ_n, bΣ_n, bp_n (Assumption 3)
    Required for the almost-sure limit of Tn/n; verified for the concrete estimators of Example 2 in Theorem 1.
  • domain assumption Asymptotic linear expansions of the estimators with square-integrable influence functions (Assumption 4)
    Needed for the Hilbert-space CLT that yields the limiting null distribution; again verified for the concrete estimators.
  • standard math Central limit theorem in separable Hilbert spaces (Laha & Rohatgi, Thm 7.5.1)
    Invoked in the proof of Theorem 3 to obtain weak convergence of the empirical process V_n.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A BHEP test for multivariate normality on incomplete data." pith.science (2026). https://pith.science/paper/V36QJHQO

@misc{pith2026260703335,
  author       = {Pith},
  title        = {Pith review of: A BHEP test for multivariate normality on incomplete data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V36QJHQO}},
  note         = {Machine review of arXiv:2607.03335}
}
read the original abstract

A BHEP test for the null hypotesis of multivariate normality on the basis of incomplete data is introduced. Estimators for the underlying unknown parameters in this situation are suggested. The test uses characteristic functions and circumvents the problem of singular covariance matrix estimates. As the sample size tends to infinity, an almost sure limit of the test statistic is obtained under the null hypothesis and under alternatives. The convergence in distribution under the null hypothesis is also proved. Critical values can be obtained using a bootstrap procedure. Simulation studies investigate size and power of the test and confirm the adequacy of the approach. A real data example demonstrates the application of the test.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A unified approach for testing in Hilbert spaces on incomplete data

    math.ST 2026-07 conditional novelty 4.0 of 10

    Incomplete observations in high-dimensional or functional spaces can be tested by averaging standard tests over finite-dimensional projections, provided missingness is independent of the data and every coordinate subs...

Reference graph

Works this paper leans on

16 extracted references · 9 canonical work pages · cited by 1 Pith paper

  1. [1]

    G., & Miloˇ sevi´ c, B

    Aleksi´ c, D. G., & Miloˇ sevi´ c, B. (2024). To impute or not? testing multivariate normality on incomplete dataset: revisiting the bhep test.Journal of Applied Statistics, 1–18. doi: 10.1080/02664763.2024.2438798

  2. [2]

    Baringhaus, L., & Henze, N. (1988). A consistent test for multivariate normality based on the empirical characteristic function.Metrika: International Journal for Theoretical and Applied Statistics,35, 339-348. doi: 10.1007/BF02613322

  3. [3]

    M., Cleveland, W

    Chambers, J. M., Cleveland, W. S., Kleiner, B., & Tukey, P. A. (1983).Graphical methods for data analysis. Pacific Grove, CA: Wadsworth & Brooks/Cole. (R dataset ’airquality’)

  4. [4]

    Das, A., & Geisler, W. S. (2021). A method to integrate and classify normal distributions. Journal of Vision,21, 1-17. doi: 10.1167/jov.21.10.1 24

  5. [5]

    Devroye, L. (1990). A note on linnik’s distribution.Statistics & Probability Letters,9(4), 305-306. doi: 10.1016/0167-7152(90)90136-U

  6. [6]

    Ebner, B., & Henze, N. (2020). Tests for multivariate normality—a critical review with emphasis on weightedL 2-statistics.TEST: An Official Journal of the Spanish Society of Statistics and Operations Research,29, 845-892. doi: 10.1007/s11749-020-00740-0

  7. [7]

    Gaigall, D. (2020). Testing marginal homogeneity of a continuous bivariate distribution with possibly incomplete paired data.Metrika: International Journal for Theoretical and Applied Statistics,83, 437-465. doi: 10.1007/s00184-019-00742-5

  8. [8]

    Gaigall, D., Wu, S., & Liang, H. (2025). A general approach for testing independence in hilbert spaces.Journal of Multivariate Analysis,206, 105384. doi: 10.1016/j.jmva.2024 .105384

Show all 16 references
  1. [9]

    Gaigall, D., & W¨ ubbolding, P. (2025). A goodness-of-fit test for geometric Brownian motion.Computational Statistics & Data Analysis,210, 108196. doi: 10.1016/j.csda .2025.108196

  2. [10]

    Henze, N., & Jim´ enez-Gamero, M. D. (2020, 05). A test for gaussianity in hilbert spaces via the empirical characteristic functional.Scandinavian Journal of Statistics,48, 406-428. doi: 10.1111/sjos.12470

  3. [11]

    Hofert, M. (2013). On sampling from the multivariate t distribution.The R Journal,5, 129-136. doi: 10.32614/RJ-2013-033

  4. [12]

    Kurita, E., & Seo, T. (2022). Multivariate normality test based on kurtosis with two- step monotone missing data.Journal of Multivariate Analysis,188, 104824. (50th Anniversary Jubilee Edition) doi: https://doi.org/10.1016/j.jmva.2021.104824

  5. [13]

    G., & Rohatgi, V

    Laha, R. G., & Rohatgi, V. K. (1979). Probability theory. InWiley series in probability and mathematical statistics.John Wiley & Sons, Ltd

  6. [14]

    Tan, M., Fang, H.-B., Tian, G.-L., & Wei, G. (2005). Testing multivariate normality in incomplete data of small sample size.Journal of Multivariate Analysis,93(1), 164-179. doi: https://doi.org/10.1016/j.jmva.2004.02.014

  7. [15]

    Tsatsi, A., Batsidis, A., & Economou, P. (2024). Multivariate normality tests with two- step monotone missing data: a critical review with emphasis on the different methods of handling missing values.Journal of Statistical Computation and Simulation,94(16), 3653–3677. doi: 10....

  8. [16]

    M., & Richards, D

    Yamada, T., Romer, M. M., & Richards, D. S. P. (2015). Kurtosis tests for multivariate normality with monotone incomplete data.TEST,24(3), 532–557. doi: 10.1007/s11749 -015-0427-4 25 ˇCoupek, P., Doln´ ık, V., Hl´ avka, Z., & Hlubinka, D. (2024). Fourier approach to goodness- ...

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.