REVIEW 3 major objections 5 minor 38 references
When missingness is independent of the data, the distributions of all finite-dimensional projections of the incomplete data determine the complete-data distribution, reducing every Hilbert-space testing problem here to projected finite-dime
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Incomplete observations in high-dimensional or functional spaces can be tested by averaging standard tests over finite-dimensional projections, provided missingness is independent of the data and every coordinate subset has a chance of being fully observed.
T0 review reviewed 2026-08-02 challenge →
load-bearing objection Theorem 1 is sound, but the paper's central testing claim overreaches: the reduction of composite goodness-of-fit to projected problems needs a projective identifiability property that is neither stated nor proved. the 3 major comments →
A unified approach for testing in Hilbert spaces on incomplete data
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On the paper's own terms, the discovery is Theorem 1: for a separable Hilbert space H of arbitrary dimension, with I the binary-coordinate missingness variable and ⊙ the Hadamard product (componentwise multiplication that blanks out unobserved coordinates), P^X = P^{X'} holds if and only if P^{π(I⊙X)} = P^{π(I⊙X')} holds for every finite-dimensional coordinate projection π, under the assumptions that X and I (and X' and I) are independent and that each finite projection has positive probability of being fully observed. The proof chains two facts: finite projections determine a distribution on a separable Hilbert space, and, under independence, the distribution of I⊙X determines the distribut
What carries the argument
The central object is the distributional equivalence of Theorem 1, powered by three ingredients: the family P(H) of projections onto finite-dimensional coordinate subspaces, which is known to determine a distribution on a separable Hilbert space; the missingness process I with binary coordinates and the Hadamard product I⊙X that blanks out unobserved coordinates; and the independence condition X⊥I together with the positivity condition P(π(I)=1)>0, which permits recovery of the law of X from the law of I⊙X by the identity P(X_1≤x_1,...,X_d≤x_d) = P(I_1 X_1≤x_1,..., I_d X_d≤x_d, I_1=1,...,I_d=1)/P(I_1=1,...,I_d=1). The proof chains projection determination with this conditional recovery, and
Load-bearing premise
The load-bearing premise is that missingness is independent of the data values and that every finite coordinate set has positive probability of being fully observed, since the identity that recovers the complete-data distribution from the incomplete one rests entirely on that independence and positivity.
What would settle it
Take d=2, let I be supported on {(1,0),(0,1)} only, and set X=(U,V) with U and V independent uniform variables while X'=(U,U). Then P^X≠P^{X'}, yet for every finite projection π the distributions of π(I⊙X) and π(I⊙X') coincide because no projection ever observes both coordinates jointly—showing that the positivity condition P(π(I)=1)>0 is necessary for the theorem's equivalence.
If this is right
- Goodness-of-fit, symmetry, homogeneity, marginal homogeneity, and independence tests on incomplete Hilbert-space data can each be written as 'for all finite projections π the projected hypothesis holds, versus there exists π for which it fails'.
- A test statistic may be formed as T_n = Σ_π q(π)T_n(π) with arbitrary weights chosen by the statistician, and a bootstrap quantile of T_n* under the null can serve as the critical value.
- The same incomplete-data model covers missing entries in random vectors, monotone dropout, partially observed functions, smoothed functions, and ultra-high-dimensional vectors via a dimension d_n that tends to infinity.
- For the normality goodness-of-fit example, the characteristic function of the projected incomplete data under the null has a closed form in terms of estimated mean, covariance, and missingness probabilities, yielding a characteristic-function-based test that is available even when the dimension exceeds the sample size.
Where Pith is reading between the lines
- If the equivalence is taken as a template, any consistent finite-dimensional test for each projection—not only characteristic-function-based ones—can be averaged to yield a Hilbert-space test, so the proposal is a general reduction scheme rather than one particular test.
- The independence assumption is the practical boundary: in settings like dropout in clinical trials, where leaving the study may depend on the unobserved outcome, the recovery identity fails; a future extension would need to model the missingness mechanism or introduce imputation.
- In the ultra-high-dimensional regime, the requirement that every finite coordinate set be observed with positive probability may be unrealistic for large coordinate sets; a possible relaxation would restrict π to coordinate sets that are actually observable, sacrificing exact equivalence for an asymptotic version.
- Because the weighted statistic is an average over randomly generated projections, variance reduction and replication-sparing techniques from Monte-Carlo bootstrap testing could apply directly to make the procedure computationally feasible.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a projection-based framework for hypothesis testing on incomplete observations with values in a separable Hilbert space. Incompleteness is encoded by a Bernoulli-valued coordinate indicator I and observed value I⊙X. The main mathematical result, Theorem 1 in Section 2.4, asserts that for two random elements X and X′ with I independent of both, and under positivity of all finite-dimensional complete-case probabilities, P^X = P^{X′} if and only if P^{π(I⊙X)} = P^{π(I⊙X′)} for every finite-dimensional projection π. On this basis, Section 2.5 claims that all listed testing problems — goodness-of-fit, symmetry, homogeneity, independence — can be expressed equivalently as "∀π H0(π) vs. ∃π H1(π)", and proposes tests of the form T_n = Σ_π q(π) T_n(π) with bootstrap critical values. Section 2.6 sketches the example of goodness-of-fit for normality via a BHEP-type statistic, referring to a companion paper for details. Section 3 states that the distribution theory of the proposed statistics requires further investigation.
Significance. If the central reduction were valid, the paper would provide a clean and broadly applicable identification statement: under MCAR-type missingness and positivity, finite-dimensional projected incomplete-data laws determine the full law, so that nonparametric hypotheses can be attacked by aggregating projected tests. The proof of Theorem 1 is elementary and correct under the stated assumptions, and the explicit identification formula is a useful contribution. Credit is due for identifying the key independence and positivity conditions and for embedding several missingness patterns in one framework. However, the leap from Theorem 1 to composite parametric goodness-of-fit is not justified, the only concrete example is delegated to a companion paper, and no asymptotic properties of the proposed test statistics are established. The paper is therefore better viewed as a programmatic contribution than as a completed testing methodology, and its advertised reduction is not yet a theorem.
major comments (3)
- [Section 2.5 and Example 3] The central reduction for goodness-of-fit is not proved. Theorem 1 is an equivalence between two fixed laws: P^X=P^{X'} ⇔ ∀π P^{π(I⊙X)}=P^{π(I⊙X')}. For the composite null H0: P^X ∈ {P^X_ϑ; ϑ∈Θ}, the reverse direction requires a projective identifiability/closure property: if for every π there exists some ϑ_π with P^{π(I⊙X)} = P^{π(I⊙X)}_{ϑ_π}, then there exists a single ϑ with P^X = P^X_ϑ. This does not follow from Theorem 1 because ϑ_π may vary with π. Example 3 simply asserts the equivalence, and Section 2.6 does not repair the gap: the claim there that φ_π = ϕ_π for all π iff H0 is stated without proof and is delegated to Gaigall and Wübbolding (2026). This is a load-bearing gap for the paper's main claim that all listed testing problems reduce to families of projected tests.
- [Section 2.4 and Example 2] The theorem's two key assumptions — I ⊥ X and P(π(I)=1)>0 for every finite π — are not satisfied by some of the motivating applications. In the monotone dropout model of Example 2b, D_i is specified as iid but its independence from X_i is not stated; informative dropout is precisely a concern in the cited depression-trial setting and, if D_i depends on unobserved coordinates of X_i, the identification formula in the proof of Theorem 1 collapses. In the ultra-high-dimensional regime of Example 2d, positivity fails for any projection involving coordinates beyond d_n, since P(I_n(j)=1)=0 for j>d_n. If the intended regime is d_n→∞ with positivity only on the first d_n coordinates, this needs to be stated explicitly, and the asymptotic consequences need to be developed.
- [Section 2.6] The normality example is not fully coherent as written. The family P^{π(I⊙X)} is described as "corresponding families of k-dimensional normal distributions", but π(I⊙X) has zero coordinates with positive probability whenever some components of I are zero; its distribution under H0 is a mixture over missingness patterns, as the displayed formula for ϕ_π shows. The equivalence "φ_π = ϕ_π for all π if and only if H0" is the key step for the example, but no proof is provided here. Since this is the only concrete illustration of the general idea, the example should either be proved within the paper or explicitly labeled as a conjecture/sketch with the proof referring to a verifiable source.
minor comments (5)
- [Throughout] The notation P^{π(I⊙X)} is used both for the distribution of π(I⊙X) and, in Example 3, for the parametric family {P^{π(I⊙X)}_ϑ; ϑ∈Θ}. This ambiguity is confusing; use different notation, e.g. P_ϑ^π for the projected parametric family.
- [Section 2.4] The proof of Theorem 1 relies on the fact that finite-dimensional projections determine distributions on H, referring to "similar arguments as in Ditzhaus and Gaigall (2018)". This is standard, but since it is the core technical step of the theorem, the paper should either give a self-contained argument or a precise statement with a full reference.
- [Section 2.5] If T_n is a weighted average with weight function q that is not supported on all finite-dimensional projections, a test based on T_n may have no power against alternatives for which H1(π) holds only on projections with q(π)=0. If consistency is intended, the support condition on Q/q should be stated.
- [Section 2.6] The estimators μ̂_n, Σ̂_n, and p̂_n(a) are introduced without any concrete definition or regularity conditions. At least one explicit estimator (or a reference to the companion paper with the relevant conditions) should be given, since the test statistic depends on them.
- [Section 3 and Section 2.6] Typos: "Smilar" in Section 3 should be "Similar"; "scetch" in Section 2.6 should be "sketch".
Circularity Check
No significant circularity: Theorem 1 is proved in-text; the Section 2.5 composite-family reduction is an unproved generalization, not a circular step.
full rationale
The derivation chain is not circular. Theorem 1 is argued in-text: the projection-uniqueness fact is stated explicitly ('we obtain that the projections in P(H) determines the distribution P^X of X uniquely') and the incomplete-data identification formula is displayed in the proof ('Analogously to Gaigall and Wübbolding (2026), it is P(X(1)≤x_1,...,X(d)≤x_d)=...'), so the core equivalence P^X=P^{X'} ⇔ P^{I⊙X}=P^{I⊙X'} is not an assumed input. The BHEP illustration in Section 2.6 relies on the companion paper for the concrete test but also derives the characteristic-function identity from Theorem 1. The main caveat is non-circular: Section 2.5/Example 3 asserts that Theorem 1 reduces composite goodness-of-fit to '∀π∈P(H):H0(π) vs. ∃π∈P(H):H1(π)', but Theorem 1 only equates fixed laws; for a general parametric family the reverse direction requires a projective identifiability/closure property that is not stated or proved, and ϑ_π may vary with π. This is a logical gap/correctness risk, not a circularity, and the paper itself concedes that the test statistics' development 'requires independent investigations.' Self-citations are numerous and structural, but the load-bearing identification is displayed in the text rather than imported as an unverified black box, and the projected-family construction does not define the target conclusion into existence. I therefore find no significant circularity; the score of 1 reflects only the heavy structural reliance on the authors' prior framework, not a circular step.
Axiom & Free-Parameter Ledger
free parameters (3)
- Weight function q / probability measure Q on P(H)
- Projection-sampling distributions ν and η
- Estimators μ̂_n, Σ̂_n, p̂_n(a)
axioms (5)
- standard math Finite-rank coordinate projections determine the distribution on a separable Hilbert space; the characteristic functional determines the distribution
- domain assumption Missingness indicator I is independent of X (and of X')
- domain assumption Positivity: P(π(I)=1) > 0 for every π ∈ P(H)
- domain assumption Observations (I_i, I_i⊙X_i) are i.i.d. copies of (I, I⊙X)
- ad hoc to paper For parametric goodness-of-fit, ∀π: P^{π(I⊙X)} ∈ P^{π(I⊙X)} implies P^X ∈ P^X
Cite this review
Pith. "Pith review of A unified approach for testing in Hilbert spaces on incomplete data." pith.science (2026). https://pith.science/paper/PPUPWQHK
@misc{pith2026260713209,
author = {Pith},
title = {Pith review of: A unified approach for testing in Hilbert spaces on incomplete data},
year = {2026},
howpublished = {\url{https://pith.science/paper/PPUPWQHK}},
note = {Machine review of arXiv:2607.13209}
}
read the original abstract
We consider statistical testing on the basis of incomplete observations with values in a separable Hilbert space, where the dimension is possibly large or even infinite. The general Hilbert space setting allows various data types as they arise in modern applications, in particular high dimensional and functional data. Possible Hilbert space testing problems are goodness-of-fit, symmetry, homogeneity and independence. We present an approach for modeling incomplete data that covers several problems in practice, e.g., ultra high dimensional random vectors with missing entries or partially observed stochastic processes. We identify a specific structure (independent and identically distributed) in the incomplete data that enables the analysis of statistical procedures with the help of suitable mathematical results (e.g., laws of large numbers and central limit theorems). Additionally, a general and novel concept for testing different hypotheses in this situation is suggested and sketched for the example of testing goodness-of-fit for normality.
Reference graph
Works this paper leans on
-
[1]
Danijel G. Aleksić and Bojana Milošević. To impute or not? testing multivariate normality on incomplete dataset: revisiting the bhep test. Journal of Applied Statistics, pages 1--18, 2024. doi:10.1080/02664763.2024.2438798
arXiv 2024
-
[2]
Nail K. Bakirov, Maria L. Rizzo, and Gábor J. Székely. A multivariate nonparametric test of independence. Journal of Multivariate Analysis, 97 0 (8): 0 1742--1756, 2006. doi:10.1016/j.jmva.2005.10.005
-
[3]
A goodness-of-fit test for the compound poisson exponential model
Ludwig Baringhaus and Daniel Gaigall. A goodness-of-fit test for the compound poisson exponential model. Journal of Multivariate Analysis, 195: 0 105154, 2023. doi:10.1016/j.jmva.2022.105154
arXiv 2023
-
[4]
A consistent test for multivariate normality based on the empirical characteristic function
Ludwig Baringhaus and Norbert Henze. A consistent test for multivariate normality based on the empirical characteristic function . Metrika: International Journal for Theoretical and Applied Statistics, 35: 0 339--348, 1988. doi:10.1007/BF02613322
-
[5]
Federico A. Bugni and Joel L. Horowitz. Permutation tests for equality of distributions of functional data. Journal of Applied Econometrics, 36: 0 861--877, 2021. doi:10.1002/jae.2846
-
[6]
Federico A. Bugni, Peter Hall, Joel L. Horowitz, and George R. Neumann. Goodness-of-fit tests for functional data. The Econometrics Journal, 12: 0 S1--S18, 2009. doi:10.1111/j.1368-423X.2008.00266.x
Pith/arXiv arXiv 2009
-
[7]
A simple multiway ANOVA for functional data
Juan Antonio Cuesta-Albertos and Manuel Febrero-Bande. A simple multiway ANOVA for functional data . TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, 19: 0 537--557, 2010. doi:10.1007/s11749-010-0185-3
-
[8]
Random projections and goodness-of-fit tests in infinite-dimensional spaces
Juan Antonio Cuesta-Albertos, Ricardo Fraiman, and Thomas Ransford. Random projections and goodness-of-fit tests in infinite-dimensional spaces. Bull Braz Math Soc, 37: 0 477--501, 2006. doi:10.1007/s00574-006-0023-0
-
[9]
The random projection method in goodness of fit for functional data
Juan Antonio Cuesta-Albertos, Eustasio del Barrio , Ricardo Fraiman, and Carlos Matrán. The random projection method in goodness of fit for functional data. Computational Statistics & Data Analysis, 51: 0 4814--4831, 2007. doi:10.1016/j.csda.2006.09.007
-
[10]
On depth measures and dual statistics
Antonio Cuevas and Ricardo Fraiman. On depth measures and dual statistics. A methodology for dealing with general data . Journal of Multivariate Analysis, 100: 0 753--766, 2009. doi:10.1016/j.jmva.2008.08.002
-
[11]
To impute or to adapt? Model specification tests’ perspective
Marija Cuparić and Bojana Milošević. To impute or to adapt? Model specification tests’ perspective . Statistical Papers, 65: 0 1021--1039, 2024. doi:10.1007/s00362-023-01421-4
-
[12]
Nonparametric density estimation for intentionally corrupted functional data
Aurore Delaigle and Alexander Meister. Nonparametric density estimation for intentionally corrupted functional data. Statistica Sinica, 31 0 (4): 0 pp. 1915--1934, 2021. doi:10.5705/ss.202018.0484
arXiv 1915
-
[13]
Michael J. Detke, Curtis G. Wiltse, Craig H. Mallinckrodt, Robert K. McNamara, Mark A. Demitrack, and Istvan Bitter. Duloxetine in the acute and long-term treatment of major depressive disorder: a placebo- and paroxetine-controlled trial. European Neuropsychopharmacology, 14: 0 457--470, 2004. doi:10.1016/j.euroneuro.2004.01.002
-
[14]
New energy distances for statistical inference on infinite dimensional hilbert spaces without moment conditions
Holger Dette and Jiajun Tang. New energy distances for statistical inference on infinite dimensional hilbert spaces without moment conditions. Bernoulli, 2026. URL https://www.e-publications.org/ims/submission/BEJ/user/submissionFile/64170?confirm=122adcd9
2026
-
[15]
A consistent goodness-of-fit test for huge dimensional and functional data
Marc Ditzhaus and Daniel Gaigall. A consistent goodness-of-fit test for huge dimensional and functional data. Journal of Nonparametric Statistics, 30 0 (4): 0 834--859, 2018. doi:10.1080/10485252.2018.1486402
arXiv 2018
-
[16]
Testing marginal homogeneity in hilbert spaces with applications to stock market returns
Marc Ditzhaus and Daniel Gaigall. Testing marginal homogeneity in hilbert spaces with applications to stock market returns. TEST, 31: 0 749--770, 2022. doi:10.1007/s11749-022-00802-5
-
[17]
Tests for multivariate normality—a critical review with emphasis on weighted L^2 -statistics
Bruno Ebner and Norbert Henze. Tests for multivariate normality—a critical review with emphasis on weighted L^2 -statistics . TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, 29: 0 845--892, 2020. doi:10.1007/s11749-020-00740-0
-
[18]
T. W. Epps and Lawrence B. Pulley. A test for normality based on the empirical characteristic function. Biometrika, 70 0 (3): 0 723--726, 1983. ISSN 00063444. doi:10.2307/2336512
doi:10.2307/2336512 1983
-
[19]
Escobar, Bernadette Klotz, Beatriz E
Juan S. Escobar, Bernadette Klotz, Beatriz E. Valdes, and Gloria M. Agudelo. The gut microbiota of colombians differs from that of americans, europeans and asians. BMC Microbiology, 14: 0 311, 2014. doi:10.1186/s12866-014-0311-6
-
[20]
Rank-based two-sample tests for paired data with missing values
Youyi Fong, Ying Huang, Maria P Lemos, and M Juliana Mcelrath. Rank-based two-sample tests for paired data with missing values. Biostatistics, 19: 0 281--294, 2017. doi:10.1093/biostatistics/kxx039
-
[21]
Daniel Gaigall. Testing marginal homogeneity of a continuous bivariate distribution with possibly incomplete paired data . Metrika: International Journal for Theoretical and Applied Statistics, 83: 0 437--465, 2020. doi:10.1007/s00184-019-00742-5
-
[22]
On a new approach to the multi-sample goodness-of-fit problem
Daniel Gaigall. On a new approach to the multi-sample goodness-of-fit problem. Communications in Statistics - Simulation and Computation, 50: 0 2971--2989, 2021. doi:10.1080/03610918.2019.1618472
arXiv 2021
-
[23]
On the applicability of several tests to models with not identically distributed random effects
Daniel Gaigall. On the applicability of several tests to models with not identically distributed random effects. Statistics, 57: 0 300--327, 2023. doi:10.1080/02331888.2023.2193748
arXiv 2023
-
[24]
On the number of replications in resampling tests and monte carlo simulation studies
Daniel Gaigall and Julian Gerstenberg. On the number of replications in resampling tests and monte carlo simulation studies. The American Statistician, pages 1--20, 2026. doi:10.1080/00031305.2025.2612197
Pith/arXiv arXiv 2026
-
[25]
A goodness-of-fit test for geometric Brownian motion
Daniel Gaigall and Philipp Wübbolding. A goodness-of-fit test for geometric Brownian motion. Computational Statistics & Data Analysis, 210: 0 108196, 2025. doi:10.1016/j.csda.2025.108196
arXiv 2025
-
[26]
A BHEP test for multivariate normality on incomplete data, 2026
Daniel Gaigall and Philipp Wübbolding. A BHEP test for multivariate normality on incomplete data, 2026. URL https://arxiv.org/abs/2607.03335
Pith/arXiv arXiv 2026
-
[27]
A general approach for testing independence in hilbert spaces
Daniel Gaigall, Shunyao Wu, and Hua Liang. A general approach for testing independence in hilbert spaces. Journal of Multivariate Analysis, 206: 0 105384, 2025. doi:10.1016/j.jmva.2024.105384
arXiv 2025
-
[28]
David J. Goldstein, Yili Lu, Michael J. Detke, Curtis Wiltse, Craig Mallinckrodt, and Mark A. Demitrack. Duloxetine in the treatment of depression: a double-blind placebo-controlled comparison with paroxetine. Journal of Clinical Psychopharmacology, 24: 0 389--399, 2004. doi:10.1097/01.jcp.0000132448.65972.d9
arXiv 2004
-
[29]
A review on specification tests for models with functional data
Wenceaslao Gonz \'a lez-Manteiga. A review on specification tests for models with functional data. Spanish Journal of Statistics, pages 9--40, 2022. doi:10.37830/SJS.2022.1.02
-
[30]
A test for gaussianity in hilbert spaces via the empirical characteristic functional
Norbert Henze and María Dolores Jiménez-Gamero. A test for gaussianity in hilbert spaces via the empirical characteristic functional. Scandinavian Journal of Statistics, 48: 0 406--428, 05 2020. doi:10.1111/sjos.12470
-
[31]
Norbert Henze, Bernhard Klar, and Simos G. Meintanis. Invariant tests for symmetry about an unspecified point based on the empirical characteristic function. Journal of Multivariate Analysis, 87: 0 275--297, 2003. doi:10.1016/S0047-259X(03)00044-7
-
[32]
V. S. Koroljukand and X.V. Borovskich. Theory of u -statistics. Dordrecht: Kluwer Academic Publishers Group, 1994. doi:10.1007/978-94-017-3515-5
-
[33]
Components and completion of partially observed functional data
David Kraus. Components and completion of partially observed functional data. Journal of the Royal Statistical Society Series B, 77: 0 777--801, 2015. doi:10.1111/rssb.12087
-
[34]
R. G. Laha and V. K. Rohatgi. Probability theory. In Wiley Series in Probability and Mathematical Statistics. John Wiley & Sons, Ltd, 1979
1979
-
[35]
Meintanis, James Allison, and Leonard Santana
Simos G. Meintanis, James Allison, and Leonard Santana. Goodness-of-fit tests for semiparametric and parametric hypotheses based on the probability weighted empirical characteristic function . Statistical Papers, 57: 0 957--976, 2016. doi:10.1007/s00362-016-0760-0
-
[36]
Fourier–legendre expansion of the one-electron density matrix of ground-state two-electron atoms
Sébastien Ragot and María Belén Ruiz. Fourier–legendre expansion of the one-electron density matrix of ground-state two-electron atoms. The Journal of Chemical Physics, 129: 0 124117, 2008. doi:10.1063/1.2981526
-
[37]
Joseph P. Romano. Bootstrap and Randomization Tests of some Nonparametric Hypotheses . The Annals of Statistics, 17: 0 141 -- 159, 1989. doi:10.1214/aos/1176347007
arXiv 1989
-
[38]
Fourier approach to goodness-of-fit tests for Gaussian random processes
Petr Čoupek, Viktor Dolník, Zdeněk Hlávka, and Daniel Hlubinka. Fourier approach to goodness-of-fit tests for Gaussian random processes. Statistical Papers, 65: 0 2937--2972, 2024. doi:10.1007/s00362-023-01510-4
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.