{"id":"d916ff5f-1d38-4691-b0b1-6bbe1c783428","arxiv_id":"2607.13209","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Incomplete observations in high-dimensional or functional spaces can be tested by averaging standard tests over finite-dimensional projections, provided missingness is independent of the data and every coordinate subset has a chance of being fully observed.","lead":"This paper packages testing with missing data in high-dimensional and functional spaces into one framework: represent each partly observed point by a binary mask multiplied coordinate-by-coordinate into the data, then test by averaging standard statistics over random finite-dimensional projections. Its main theorem shows the masked data pins down the full data's distribution when missingness is random, so the framework extends to normality, symmetry, homogeneity, and independ","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unproved parametric reduction in §2.5: Theorem 1 only equates fixed laws, not membership in composite families; the ∀π formulation for goodness-of-fit is not deduced.","rationale":"The reader's weakest_assumption is the MCAR independence I⊥X. That is a real scope restriction, but the paper explicitly states this assumption in Theorem 1 and frames the framework conditionally. The more load-bearing concern is the unproved parametric reduction in Section 2.5, which affects the central claim even when the MCAR and positivity assumptions hold. The reader's rationale does mention this as weakness (1), but does not elevate it to the weakest assumption. I therefore partially disagree with the ranking: the parametric reduction is more central because it is an internal logical gap, not a modeling limitation. My recommended verdict remains CONDITIONAL (labelled UNCHANGED) since the core Theorem 1 appears sound and the gap is fixable by adding a proof or restriction to families that satisfy the projective identifiability property. The paper should either prove this property for the stated families or explicitly list it as an additional condition before claiming the unified testing reduction.","tokens_in":13347,"tokens_out":28212,"duration_ms":251096,"concrete_test":"For the Gaussian family in R^d and a missingness mechanism satisfying the positivity condition of Theorem 1 (e.g., independent Bernoulli coordinates with P(I_j=1)>0), determine whether the assertion in Example 3 holds: if for every finite subset J' there exists (μ_{J'},Σ_{J'}) such that the law of (I_j X_j)_{j∈J'} equals the projected Gaussian-mixture law, must X be Gaussian? Concretely, attempt a computational search for d=3 with a non-Gaussian distribution having Gaussian marginals (e.g., a Gaussian copula with a non-Gaussian dependence structure): compute all finite projected masked laws and check whether each belongs to the corresponding normal-mixture family with some parameter ϑ_π. If such X exists, the Section 2.5 reduction is false; if not, a proof should be supplied for at least the Gaussian case.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that every listed testing problem reduces to a family of projected tests rests on Section 2.5's assertion that 'under the stated assumptions, we deduce from Theorem 1 that our testing problems ... can be expressed equivalently in the form ∀π H0(π) vs ∃π H1(π).' But Theorem 1 is an equivalence of two fixed laws: P^X=P^X' ⇔ ∀π P^{π(I⊙X)}=P^{π(I⊙X')}. For composite goodness-of-fit H0: P^X∈{P^X_ϑ; ϑ∈Θ}, the reverse direction would require: if for every π there exists some ϑ_π with P^{π(I⊙X)} = P^{π(I⊙X)}_{ϑ_π}, then there exists a single ϑ with P^X=P^X_ϑ. This is a projective identifiability/closure property of the parametric family that is neither stated nor proved, and it does not follow from Theorem 1 because ϑ_π may vary with π. Example 3 simply introduces the projected families and asserts the equivalence. The one concrete illustration (Section 2.6) delegates the proof to Gaigall and Wübbolding (2026), which is not verifiable here. Section 3 candidly says the test statistics 'requires independent investigations,' but the reduction itself is presented as established. If the parametric family lacks this property, the proposed goodness-of-fit tests could have incorrect size/power even under perfect MCAR data. This is a logical gap in the central claim, not merely a scope limitation.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a projection-based framework for hypothesis testing on incomplete observations with values in a separable Hilbert space. Incompleteness is encoded by a Bernoulli-valued coordinate indicator I and observed value I⊙X. The main mathematical result, Theorem 1 in Section 2.4, asserts that for two random elements X and X′ with I independent of both, and under positivity of all finite-dimensional complete-case probabilities, P^X = P^{X′} if and only if P^{π(I⊙X)} = P^{π(I⊙X′)} for every finite-dimensional projection π. On this basis, Section 2.5 claims that all listed testing problems — goodness-of-fit, symmetry, homogeneity, independence — can be expressed equivalently as \"∀π H0(π) vs. ∃π H1(π)\", and proposes tests of the form T_n = Σ_π q(π) T_n(π) with bootstrap critical values. Section 2.6 sketches the example of goodness-of-fit for normality via a BHEP-type statistic, referring to a companion paper for details. Section 3 states that the distribution theory of the proposed statistics requires further investigation.","tokens_in":13642,"tokens_out":6969,"duration_ms":64202,"significance":"If the central reduction were valid, the paper would provide a clean and broadly applicable identification statement: under MCAR-type missingness and positivity, finite-dimensional projected incomplete-data laws determine the full law, so that nonparametric hypotheses can be attacked by aggregating projected tests. The proof of Theorem 1 is elementary and correct under the stated assumptions, and the explicit identification formula is a useful contribution. Credit is due for identifying the key independence and positivity conditions and for embedding several missingness patterns in one framework. However, the leap from Theorem 1 to composite parametric goodness-of-fit is not justified, the only concrete example is delegated to a companion paper, and no asymptotic properties of the proposed test statistics are established. The paper is therefore better viewed as a programmatic contribution than as a completed testing methodology, and its advertised reduction is not yet a theorem.","major_comments":[{"comment":"The central reduction for goodness-of-fit is not proved. Theorem 1 is an equivalence between two fixed laws: P^X=P^{X'} ⇔ ∀π P^{π(I⊙X)}=P^{π(I⊙X')}. For the composite null H0: P^X ∈ {P^X_ϑ; ϑ∈Θ}, the reverse direction requires a projective identifiability/closure property: if for every π there exists some ϑ_π with P^{π(I⊙X)} = P^{π(I⊙X)}_{ϑ_π}, then there exists a single ϑ with P^X = P^X_ϑ. This does not follow from Theorem 1 because ϑ_π may vary with π. Example 3 simply asserts the equivalence, and Section 2.6 does not repair the gap: the claim there that φ_π = ϕ_π for all π iff H0 is stated without proof and is delegated to Gaigall and Wübbolding (2026). This is a load-bearing gap for the paper's main claim that all listed testing problems reduce to families of projected tests.","section":"Section 2.5 and Example 3"},{"comment":"The theorem's two key assumptions — I ⊥ X and P(π(I)=1)>0 for every finite π — are not satisfied by some of the motivating applications. In the monotone dropout model of Example 2b, D_i is specified as iid but its independence from X_i is not stated; informative dropout is precisely a concern in the cited depression-trial setting and, if D_i depends on unobserved coordinates of X_i, the identification formula in the proof of Theorem 1 collapses. In the ultra-high-dimensional regime of Example 2d, positivity fails for any projection involving coordinates beyond d_n, since P(I_n(j)=1)=0 for j>d_n. If the intended regime is d_n→∞ with positivity only on the first d_n coordinates, this needs to be stated explicitly, and the asymptotic consequences need to be developed.","section":"Section 2.4 and Example 2"},{"comment":"The normality example is not fully coherent as written. The family P^{π(I⊙X)} is described as \"corresponding families of k-dimensional normal distributions\", but π(I⊙X) has zero coordinates with positive probability whenever some components of I are zero; its distribution under H0 is a mixture over missingness patterns, as the displayed formula for ϕ_π shows. The equivalence \"φ_π = ϕ_π for all π if and only if H0\" is the key step for the example, but no proof is provided here. Since this is the only concrete illustration of the general idea, the example should either be proved within the paper or explicitly labeled as a conjecture/sketch with the proof referring to a verifiable source.","section":"Section 2.6"}],"minor_comments":[{"comment":"The notation P^{π(I⊙X)} is used both for the distribution of π(I⊙X) and, in Example 3, for the parametric family {P^{π(I⊙X)}_ϑ; ϑ∈Θ}. This ambiguity is confusing; use different notation, e.g. P_ϑ^π for the projected parametric family.","section":"Throughout"},{"comment":"The proof of Theorem 1 relies on the fact that finite-dimensional projections determine distributions on H, referring to \"similar arguments as in Ditzhaus and Gaigall (2018)\". This is standard, but since it is the core technical step of the theorem, the paper should either give a self-contained argument or a precise statement with a full reference.","section":"Section 2.4"},{"comment":"If T_n is a weighted average with weight function q that is not supported on all finite-dimensional projections, a test based on T_n may have no power against alternatives for which H1(π) holds only on projections with q(π)=0. If consistency is intended, the support condition on Q/q should be stated.","section":"Section 2.5"},{"comment":"The estimators μ̂_n, Σ̂_n, and p̂_n(a) are introduced without any concrete definition or regularity conditions. At least one explicit estimator (or a reference to the companion paper with the relevant conditions) should be given, since the test statistic depends on them.","section":"Section 2.6"},{"comment":"Typos: \"Smilar\" in Section 3 should be \"Similar\"; \"scetch\" in Section 2.6 should be \"sketch\".","section":"Section 3 and Section 2.6"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim — that Theorem 1 yields a reduction of all listed testing problems, including composite parametric goodness-of-fit — is not established. The missing projective identifiability condition for composite families is a logical gap, not merely a presentational issue, and it affects the only concrete example. The authors may be able to fix this by adding conditions and proofs, but in its current form the manuscript is more a research proposal than a complete methodology paper. I would also ask the editor to consider whether the heavy reliance on two companion papers (Gaigall and Wübbolding 2025, 2026) for the only concrete application and for bootstrap procedures is appropriate for the claimed contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core identification theorem is correct. Under MCAR and positivity, the equivalence P^X = P^X' ⇔ ∀π P^{π(I⊙X)} = P^{π(I⊙X')} is proved cleanly with standard arguments. The I⊙X mask-product model is also a genuinely tidy way to unify missing entries, dropout, partially observed functions, ultra-high dimension, and smoothing under one Hilbert-space roof. Credit where due: the paper is honest that the asymptotics of the proposed test statistics are not developed (Section 3 says so explicitly), and it does not pretend to have simulations or real data analysis.\n\nThe soft spot is in Section 2.5. Theorem 1 is an equivalence between two fixed laws. The paper claims that every testing problem, including composite goodness-of-fit, reduces to ∀π H0(π) vs ∃π H1(π). For goodness-of-fit, that requires: if for every π there is some ϑ_π with P^{π(I⊙X)} = P^{π(I⊙X)}_{ϑ_π}, then there is a single ϑ with P^X = P^X_ϑ. That projective identifiability/closure property is not deduced from Theorem 1, and it does not follow because ϑ_π can vary with π. Example 3 just asserts the equivalence. This is a real logical gap in the central claim, not a minor scope limitation. If a parametric family lacks the property, the proposed tests could have wrong size or power even under perfectly MCAR data.\n\nThe other weaknesses are proportional. The MCAR assumption and positivity of every finite complete-case set are restrictive—the paper's own dropout example allows informative dropout to break the identification, and the ultra-high-dimensional example strains positivity. That is not a fatal flaw; it is a scope restriction that should be stated more carefully. The BHEP illustration is a sketch and delegates all details to a companion paper, so the concrete test is not independently assessable here. No code, data, or simulations are provided.\n\nWho is this for? Researchers working on nonparametric testing with incomplete functional or high-dimensional data who want an organizing framework. The paper is a coherent template, but the testing concept is unvalidated in the composite case. It deserves a serious referee—the theorem is worth checking into the literature—but the reduction claim needs either a proof for the relevant parametric families or a clear restriction to problems where the equivalence does hold. I would send it to peer review with the expectation of major revision.","headline":"Theorem 1 is sound, but the paper's central testing claim overreaches: the reduction of composite goodness-of-fit to projected problems needs a projective identifiability property that is neither stated nor proved.","tokens_in":14203,"tokens_out":1867,"would_cite":false,"duration_ms":20109,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62G09"],"pacs":[],"model":"deepseek-v4-flash","headline":"When missingness is independent of the data, the distributions of all finite-dimensional projections of the incomplete data determine the complete-data distribution, reducing every Hilbert-space testing problem here to projected finite-dime","keywords":["bootstrap","functional data","goodness-of-fit","high dimensional data","Hilbert space","hypothesis testing","projection","incomplete data"],"falsifier":"Take d=2, let I be supported on {(1,0),(0,1)} only, and set X=(U,V) with U and V independent uniform variables while X'=(U,U). Then P^X≠P^{X'}, yet for every finite projection π the distributions of π(I⊙X) and π(I⊙X') coincide because no projection ever observes both coordinates jointly—showing that the positivity condition P(π(I)=1)>0 is necessary for the theorem's equivalence.","tokens_in":13116,"feed_emoji":"📊","tokens_out":11149,"duration_ms":91423,"temperature":0.7,"pith_summary":"The paper establishes an equivalence that makes hypothesis testing possible when data live in a separable Hilbert space—covering high-dimensional vectors and functional data—but are only partially observed. The central theorem says two complete-data distributions are equal if and only if, for every projection onto a finite coordinate set, the distributions of the projected incomplete observations are equal, provided the missingness indicator is independent of the data and every finite coordinate set has positive probability of being fully observed. This reduces the testing problems of goodness-of-fit, symmetry, homogeneity, and independence to checking a family of finite-dimensional projected hypotheses, and motivates test statistics that average per-projection statistics with weights chosen by the statistician. The authors propose bootstrap quantiles as critical values and sketch a concrete characteristic-function-based normality test on incomplete data.","feed_headline":"Projections recover full distributions from incomplete Hilbert-space data","feed_subtitle":"One theorem reduces goodness-of-fit, symmetry, homogeneity, and independence tests on missing data to finite dimensions.","key_machinery":"The central object is the distributional equivalence of Theorem 1, powered by three ingredients: the family P(H) of projections onto finite-dimensional coordinate subspaces, which is known to determine a distribution on a separable Hilbert space; the missingness process I with binary coordinates and the Hadamard product I⊙X that blanks out unobserved coordinates; and the independence condition X⊥I together with the positivity condition P(π(I)=1)>0, which permits recovery of the law of X from the law of I⊙X by the identity P(X_1≤x_1,...,X_d≤x_d) = P(I_1 X_1≤x_1,..., I_d X_d≤x_d, I_1=1,...,I_d=1)/P(I_1=1,...,I_d=1). The proof chains projection determination with this conditional recovery, and","core_discovery":"On the paper's own terms, the discovery is Theorem 1: for a separable Hilbert space H of arbitrary dimension, with I the binary-coordinate missingness variable and ⊙ the Hadamard product (componentwise multiplication that blanks out unobserved coordinates), P^X = P^{X'} holds if and only if P^{π(I⊙X)} = P^{π(I⊙X')} holds for every finite-dimensional coordinate projection π, under the assumptions that X and I (and X' and I) are independent and that each finite projection has positive probability of being fully observed. The proof chains two facts: finite projections determine a distribution on a separable Hilbert space, and, under independence, the distribution of I⊙X determines the distribut","pith_inferences":["If the equivalence is taken as a template, any consistent finite-dimensional test for each projection—not only characteristic-function-based ones—can be averaged to yield a Hilbert-space test, so the proposal is a general reduction scheme rather than one particular test.","The independence assumption is the practical boundary: in settings like dropout in clinical trials, where leaving the study may depend on the unobserved outcome, the recovery identity fails; a future extension would need to model the missingness mechanism or introduce imputation.","In the ultra-high-dimensional regime, the requirement that every finite coordinate set be observed with positive probability may be unrealistic for large coordinate sets; a possible relaxation would restrict π to coordinate sets that are actually observable, sacrificing exact equivalence for an asymptotic version.","Because the weighted statistic is an average over randomly generated projections, variance reduction and replication-sparing techniques from Monte-Carlo bootstrap testing could apply directly to make the procedure computationally feasible."],"forward_implications":["Goodness-of-fit, symmetry, homogeneity, marginal homogeneity, and independence tests on incomplete Hilbert-space data can each be written as 'for all finite projections π the projected hypothesis holds, versus there exists π for which it fails'.","A test statistic may be formed as T_n = Σ_π q(π)T_n(π) with arbitrary weights chosen by the statistician, and a bootstrap quantile of T_n* under the null can serve as the critical value.","The same incomplete-data model covers missing entries in random vectors, monotone dropout, partially observed functions, smoothed functions, and ultra-high-dimensional vectors via a dimension d_n that tends to infinity.","For the normality goodness-of-fit example, the characteristic function of the projected incomplete data under the null has a closed form in terms of estimated mean, covariance, and missingness probabilities, yielding a characteristic-function-based test that is available even when the dimension exceeds the sample size."],"fun_headline_variants":["Finite projections unlock Hilbert-space tests on partial data","Unified testing for incomplete Hilbert data via projections","Missing data in Hilbert spaces? Reduce to finite projections","All tests on partial Hilbert data reduce to finite projections"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that missingness is independent of the data values and that every finite coordinate set has positive probability of being fully observed, since the identity that recovers the complete-data distribution from the incomplete one rests entirely on that independence and positivity.","fun_headline_variants_meta":{"raw":{"variants":["Finite projections unlock Hilbert-space tests on partial data","Unified testing for incomplete Hilbert data via projections","Missing data in Hilbert spaces? Reduce to finite projections","All tests on partial Hilbert data reduce to finite projections"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000164,"raw_usage":{"total_tokens":1051,"prompt_tokens":677,"completion_tokens":374,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":421,"completion_tokens_details":{"reasoning_tokens":321}},"tokens_in":421,"tokens_out":374,"duration_ms":31122,"temperature":1.0,"reasoning_tokens":321,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T05:52:51.774507+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take d=2, let I be supported on {(1,0),(0,1)} only, and set X=(U,V) with U and V independent uniform variables while X'=(U,U). Then P^X≠P^{X'}, yet for every finite projection π the distributions of π(I⊙X) and π(I⊙X') coincide because no projection ever observes both coordinates jointly—showing that the positivity condition P(π(I)=1)>0 is necessary for the theorem's equivalence.","supporting_citations":[],"review_version":1}