REVIEW 3 major objections 6 minor 18 references
Filters must meet data on enough frequencies to maximize separation
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-08 18:02 UTC pith:NUVBHZGR
load-bearing objection Scattering network filter design criteria derived from geometric measure theory; lower bounds need polynomial parametrization assumption the 3 major comments →
Separation Capacity of Scattering Networks on Low-Dimensional Datasets
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central mechanism is a pair of bounds on the s-separation capacity SC_s of a feature extractor on a countably H^s-rectifiable set. The lower bound is governed by the rank of the differential of the feature map on approximate tangent spaces (local geometry), while the upper bound is governed by the rank of a second-moment matrix formed by pushing forward the Hausdorff measure through the feature map (global geometry). For scattering networks specifically, the upper bound translates into a requirement on the spectral overlap between filters and data, and the lower bound translates into a conditioning requirement on matrices that couple the filter frame to the parametrization of the data. A
What carries the argument
The key objects are: (1) countably H^s-rectifiable sets as a model for low-dimensional data; (2) Cover's s-separation capacity as the measure of classification power; (3) approximate tangent spaces T_f E of rectifiable sets; (4) the second-moment matrix C_Phi = integral of y y^T d Phi_sharp(H^s restricted to E); (5) the vector Veronese map v_{M,d}, which encodes the effect of the monomial nonlinearity of degree d; (6) the matrix A_Lambda constructed from filter Fourier coefficients, which couples the filters to the Veronese lift; and (7) the matrices T_{hat Psi_j}^H A_{Lambda,chi} T_{hat Psi_j}, whose conditioning controls the lower bound.
Load-bearing premise
The lower bounds, which yield the actionable conditioning-based design criteria, require the bi-Lipschitz parametrizations of the rectifiable data set to be polynomial maps. Many natural low-dimensional data structures (smooth manifolds with non-algebraic parametrizations, fractal sets) may not admit such polynomial bi-Lipschitz parametrizations, limiting the scope of the design recommendations.
What would settle it
If one could exhibit a rectifiable dataset E with a non-polynomial bi-Lipschitz parametrization where the conditioning-based design criterion fails to correlate with actual separation capacity, the practical applicability of the lower-bound design recommendations would be called into question.
If this is right
- Filter design for scattering networks on low-dimensional data can be guided by two checkable conditions: spectral coverage of the data spectrum and numerical conditioning of data-dependent matrices, rather than by empirical tuning alone.
- For sparse-signal models, the design criterion reduces to minimizing a restricted isometry constant, directly connecting scattering network filter design to the compressed sensing literature.
- The upper bound's group-theoretic condition (spectral supports not in a coset of a proper subgroup) provides a concrete, checkable necessary condition for a filter bank to achieve maximum separation capacity.
- The framework applies to any Lipschitz feature extractor on rectifiable sets, so the geometric bounds could in principle be applied to feature extractors beyond scattering networks.
Where Pith is reading between the lines
- The polynomial parametrization assumption required for the lower bounds may exclude natural data manifolds with non-algebraic structure (e.g., smooth manifolds parametrized by transcendental maps). Extending the lower-bound argument to general bi-Lipschitz parametrizations would broaden the applicability of the conditioning-based design criterion.
- The framework treats the nonlinearity degree d and network depth n_d as fixed; the interaction between depth and the conditioning of the coupling matrices is not fully explored and could reveal depth-dependent tradeoffs in separation capacity.
- The separation capacity as defined measures linear separability of the network's output features. A natural extension would investigate whether the geometric bounds transfer to nonlinear classifiers applied on top of the scattering features.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies the separation capacity of pooling-free scattering networks with fixed monomial nonlinearities on data modeled as countably H^s-rectifiable sets. The authors first establish general bounds on the s-separation capacity of feature extractors: a lower bound in terms of the rank of the differential restricted to approximate tangent spaces (Lemma 2.7) and an upper bound via a second-moment matrix (Lemma 2.9), with an exact formula under real-analyticity (Corollary 2.9.1). These are then particularized to scattering networks for two data models: sparse signals and polynomially parametrized rectifiable sets. The resulting design criteria are: (i) filter spectral supports should not be contained in a coset of a proper subgroup (Theorem 3.3), and (ii) matrices coupling the filter frame to the data geometry should be well-conditioned (Theorems 3.5, 3.8 and Corollaries 3.5.1, 3.8.1). The proofs proceed through a clean chain from the subspace characterization (Lemma 2.2) through geometric measure theory tools to concrete matrix conditions.
Significance. The paper provides a mathematically rigorous bridge between geometric measure theory and the separation capacity of scattering networks, yielding actionable filter-design criteria expressed in terms of well-conditioning of data-dependent matrices. The derivation chain is carefully executed: the area formula application in Corollary 2.9.1, the Veronese map factorization in Lemma 3.2, and the restricted isometry constant formulation in Theorem 3.5 are notable technical contributions. The reduction of the sparse-signal lower bound to an RIP condition on (T^H_{bXi} A_{Lambda,chi} T_{bXi})^{1/2} is a clean, falsifiable criterion. The upper-bound design criterion (spectral supports not in a coset of a proper subgroup) is parameter-free in the sense that it depends only on the group structure of Z/MZ and the support sets. The paper is self-contained in its proof structure, though it relies on two companion papers [7, 9] for background results on Cover-type separation capacity.
major comments (3)
- Section 3.3, Theorem 3.8: The lower bound for general rectifiable sets depends on the assumption that each bi-Lipschitz parametrization psi_j is a polynomial of degree n_j. This assumption is load-bearing: the factorization v_{M,d}(hat{psi}_j(x)) = T_{hat{psi}_j} w_{s,n_j d}(x) in the proof of Theorem 3.8 requires the Veronese lift of hat{psi}_j to land in a finite-dimensional monomial space. If psi_j is not polynomial, T_{hat{psi}_j} is not a finite-dimensional object and the trace bounds in (3.23) break down. The paper does not discuss which natural data models satisfy this polynomial assumption, nor does it acknowledge that smooth manifolds with non-algebraic parametrizations (a primary motivating example per the introduction's reference to [1, 2]) are excluded. This is a scope limitation on the central actionable design criterion (Corollary 3.8.1). The authors should either (a) add a
- Section 3.3, Theorem 3.7: The upper bound also relies on the polynomial assumption to bound dim_C(span_C(v_{M,d}(F_M(E_j)))) <= C(s + n_j d, n_j d) via the degree of x -> v_{M,d}(F_M(psi_j(x))). While the upper bound is less directly actionable, the same scope concern applies. The paper should clarify whether the polynomial assumption is essential for the upper bound as well, or whether a more general bound (e.g., in terms of the Hausdorff dimension of the Fourier support of psi_j(K_j)) could replace it.
- Theorem 3.5 and Corollary 3.5.1: The lower bound for sparse signals involves the term inf_{j in J_S} 2(Tr(C_{S,j}))^2 / Tr((C_{S,j})^2), which depends on the parametrization domains K_{S,j} but not on the frame Xi. The paper does not discuss whether this term can be bounded below in terms of s and d alone, or whether it can be arbitrarily small for adversarial choices of K_{S,j}. If the latter, the practical utility of the design criterion (minimizing delta_{s,d}(Xi; Lambda, chi)) is diminished, as the overall lower bound could still be small. A brief remark on the behavior of this term would strengthen the result.
minor comments (6)
- The notation T_{hat{psi}_j} in Theorem 3.8 uses a hat on psi_j, but the text in Section 3.3 defines psi_j without a hat. The hat presumably denotes the Fourier transform, but this should be stated explicitly.
- In the proof of Lemma 3.2, the matrix A is defined with the condition supp(alpha) subseteq H_{lambda,S}, but in the subsequent definition of A_lambda (footnote 2 on page 11), this condition is dropped. The relationship between A and A_lambda should be stated more explicitly to avoid confusion.
- Theorem 3.3: The upper bound includes a term 4 min_{S} [...] rather than 2 min_{S} [...]. The factor of 2 relative to the real-valued case is presumably due to the complex-to-real identification, but this should be noted explicitly, perhaps with a reference to equation (3.3).
- In the proof of Theorem 3.8(b), the chain of inequalities bounding Tr(G_j A_Lambda G_j A_Lambda) uses the Loewner ordering notation (preceq, succeq) without explicitly defining it. While standard, a brief note would aid readability.
- Reference [9] (arXiv:2607.01010) is cited for the identity SC_s(Phi) = min_{j in J} SC_s(Phi|_{psi_j(K_j)}) used in Corollary 2.9.1, and for the identity in (2.2). Since these are central to the derivation, the reader would benefit from a brief statement of the relevant results from [9] rather than relying entirely on the companion paper.
- The abstract states 'no pooling' but Section 3.1 defines the general scattering network with pooling operators P_n. The restriction to P_n = Id is stated in Section 3.1 but could be mentioned in the abstract for precision.
Simulated Author's Rebuttal
We thank the referee for a careful reading and for identifying a genuine scope limitation in Section 3.3 that we will address in revision. The referee's three major comments are interconnected: the first two concern the polynomial parametrization assumption in Theorems 3.7 and 3.8, and the third concerns the parametrization-domain-dependent term in the sparse-signal lower bound. We agree that the polynomial assumption is load-bearing and that the manuscript does not adequately discuss its scope or which data models satisfy it. We will add a dedicated remark addressing this. On the third comment, we will add discussion of the behavior of the trace-ratio term, including both the lower bound it admits and the adversarial cases the referee identifies.
read point-by-point responses
-
Referee: Section 3.3, Theorem 3.8: The lower bound depends on the polynomial parametrization assumption. The paper does not discuss which natural data models satisfy this, nor that smooth manifolds with non-algebraic parametrizations are excluded. The authors should either (a) add a discussion of scope, or (b) extend to non-polynomial parametrizations.
Authors: The referee is correct that the polynomial assumption is load-bearing for Theorem 3.8. Specifically, the factorization v_{M,d}(hat{psi}_j(x)) = T_{hat{psi}_j} w_{s,n_j d}(x) in the proof requires the Veronese lift of hat{psi}_j to land in a finite-dimensional monomial space, which fails for non-polynomial psi_j. We acknowledge that this excludes smooth manifolds with non-algebraic parametrizations, which the introduction's reference to [1, 2] might suggest are covered. We will add a dedicated remark in Section 3.3 explicitly stating this scope limitation, clarifying that the results of Theorem 3.8 and Corollary 3.8.1 apply to polynomially parametrized rectifiable sets (which include sparse-signal models as a special case, as well as algebraic varieties and their finite unions). Regarding option (b), extending to non-polynomial parametrizations would require replacing the finite-dimensional matrix T_{hat{psi}_j} with an operator acting on an infinite-dimensional function space, and the trace bounds in (3.23) would no longer apply in their current form. We believe this extension is beyond the scope of the present paper, but we will note it as a direction for future work. We will also adjust the introduction to avoid implying that general smooth submanifolds are covered by the polynomial results. revision: yes
-
Referee: Section 3.3, Theorem 3.7: The upper bound also relies on the polynomial assumption to bound dim_C(span_C(v_{M,d}(F_M(E_j)))) <= C(s + n_j d, n_j d). The paper should clarify whether the polynomial assumption is essential for the upper bound, or whether a more general bound could replace it.
Authors: The referee correctly identifies that the bound dim_C(span_C(v_{M,d}(F_M(E_j)))) <= C(s + n_j d, n_j d) in Lemma 3.6 relies on x -> v_{M,d}(F_M(psi_j(x))) being a multivariate polynomial of degree at most n_j d, which requires psi_j to be polynomial. For the upper bound, however, the situation is somewhat different from the lower bound. The term |H_{d,lambda,psi_j} cap supp(hat{chi})| in (3.21) does not depend on the polynomial assumption and remains valid for general bi-Lipschitz parametrizations. Only the combinatorial term C(s + n_j d, n_j d) requires polynomial structure. For a general (non-polynomial) bi-Lipschitz parametrization, the Veronese lift v_{M,d}(hat{psi}_j(x)) need not lie in any finite-dimensional subspace, so the combinatorial bound breaks down. A replacement bound in terms of the Hausdorff dimension of the Fourier support of psi_j(K_j) is an interesting possibility, but it would require a fundamentally different proof technique, as the current argument proceeds through finite-dimensional polynomial degree counting. We will add a remark clarifying that the polynomial assumption is essential for the combinatorial term in the upper bound, while the spectral-support term is not, and that extending the upper bound to non-polynomial parametrizations is left open. revision: yes
-
Referee: Theorem 3.5 and Corollary 3.5.1: The term inf_{j in J_S} 2(Tr(C_{S,j}))^2 / Tr((C_{S,j})^2) depends on the parametrization domains K_{S,j} but not on the frame Xi. The paper does not discuss whether this term can be bounded below in terms of s and d alone, or whether it can be arbitrarily small for adversarial choices of K_{S,j}.
Authors: The referee raises a valid point. The term 2(Tr(C_{S,j}))^2 / Tr((C_{S,j})^2) equals 2 times the effective rank of C_{S,j}, which is at least 2 and at most 2*C(s-1+d, d). For well-behaved parametrization domains (e.g., K_{S,j} containing an open ball in R^{2s}), C_{S,j} is full-rank and the term equals 2*C(s-1+d, d), the maximum possible value. However, the referee is correct that for adversarial choices of K_{S,j} — for instance, domains concentrated near a lower-dimensional subset — the matrix C_{S,j} can become rank-deficient and the term can be arbitrarily small. We will add a remark noting both the upper bound 2*C(s-1+d, d) (achieved for full-dimensional domains) and the fact that the term can degenerate for adversarial domains. We would also note that in the typical setting where the parametrization domains are fixed by the data model (not chosen adversarially), this term is a geometric constant of the dataset, and the design criterion of minimizing delta_{s,d}(Xi; Lambda, chi) remains the actionable lever for the practitioner, as the frame Xi is the only design variable. revision: yes
Circularity Check
No significant circularity; two self-citations to companion papers are used as foundational definitions, but the central bounds and design criteria are derived independently from standard GMT results.
full rationale
The paper's central results—Lemma 2.7 (lower bound via tangent space rank), Lemma 2.9 (upper bound via second-moment matrix), and the scattering-network design criteria (Theorems 3.3, 3.5, 3.7, 3.8)—are derived from first principles using standard geometric measure theory results [3, 4, 11, 12]. The two self-citations [7, 9] are used to import the definition of separation capacity (Definition 2.1, citing [9]) and the identity SC_s(Φ) = 2·min dim(span(Φ(A))) (Eq. 2.2, citing [9]). These are foundational definitions rather than results being 'predicted' or 'derived' in this paper. The actual bounds in Lemmas 2.7 and 2.9 are proven self-containedly: Lemma 2.7 uses approximate tangent space theory [12] and does not rely on [9] for its proof; Lemma 2.9 uses a kernel/support argument independent of [9]. The particularization to scattering networks (Section 3) uses the Veronese map, area formula, and RIP theory [17, 18]—all external. The self-citation [9] provides the function-counting framework (analogous to how Cover's theorem [8] is cited), and the present paper builds GMT-based bounds on top of it. The identity in Eq. 2.2 is a definition/import, not a circular 'prediction.' The polynomial parametrization assumption in Section 3.3 is a scope limitation (correctness risk), not a circularity issue. No step in the derivation chain reduces to its own inputs by construction. The self-citations are standard companion-paper references for definitions, not load-bearing circular chains. Score 2 reflects the presence of self-cited definitions without independent verification, but the central technical content is independently derived.
Axiom & Free-Parameter Ledger
free parameters (3)
- Monomial degree d
- Network depth n_d
- Bi-Lipschitz constant t
axioms (5)
- domain assumption Data lies on a countably H^s-rectifiable set (Definition 2.4)
- domain assumption The feature extractor Φ is Lipschitz (Lemma 2.7)
- ad hoc to paper The bi-Lipschitz parametrization {ψ_j} consists of polynomials of degree n_j (Section 3.3)
- standard math The s-separation capacity equals 2·min over positive-measure subsets A of dim_R(span_R(Φ(A))) (Eq. 2.2, attributed to [9])
- domain assumption Real-analytic extension of Φ∘ψ_j to a connected open neighborhood (Corollary 2.9.1)
read the original abstract
We aim to identify scattering network architectures that maximize the separation capacity on data with low intrinsic dimension. The networks we consider employ a fixed monomial nonlinearity and no pooling, so that the only design variable is the frame generated by the network filters. For data modeled as rectifiable sets, we first characterize and bound the separation capacity of general feature extractors in terms of the geometry of the dataset. We then particularize to scattering networks and obtain two design criteria: (i) the filters should meet the data on sufficiently many frequencies, and (ii) the matrices coupling the frame to the geometry of the data should be well-conditioned.
Reference graph
Works this paper leans on
-
[1]
Representation learning: A review and new perspectives,
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,”IEEE transactions on pattern analysis and machine intelligence, vol. 35, no. 8, pp. 1798–1828, 2013
work page 2013
-
[2]
Testing the manifold hypothesis,
C. Fefferman, S. Mitter, and H. Narayanan, “Testing the manifold hypothesis,”Journal of the American Mathematical Society, vol. 29, no. 4, pp. 983–1049, 2016
work page 2016
- [3]
-
[4]
L. Ambrosio and B. Kirchheim, “Currents in metric spaces,”Acta Mathematica, vol. 185, pp. 1–80, 2000
work page 2000
-
[5]
S. Mallat, “Group invariant scattering,”Communications on Pure and Applied Mathemat- ics, vol. 65, no. 10, pp. 1331–1398, 2012. Separation Capacity of Scattering Networks on Low-Dimensional Datasets 19
work page 2012
-
[6]
A mathematical theory of deep convolutional neural net- works for feature extraction,
T. Wiatowski and H. B¨ olcskei, “A mathematical theory of deep convolutional neural net- works for feature extraction,”IEEE Transactions on Information Theory, vol. 64, no. 3, pp. 1845–1866, 2017
work page 2017
-
[7]
Separation Capacity of Scattering Networks
K. H¨ aberle and H. B¨ olcskei, “Separation capacity of scattering networks,”arXiv preprint arXiv:2606.30822, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[8]
T. M. Cover, “Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition,”IEEE Transactions on Electronic Computers, no. 3, pp. 326–334, 1965
work page 1965
-
[9]
Function-Counting Theory for Low-Dimensional Data Structures
K. H¨ aberle and H. B¨ olcskei, “Function-counting theory for low-dimensional data struc- tures,”arXiv preprint arXiv:2607.01010, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[10]
Christensen,An introduction to frames and Riesz bases
O. Christensen,An introduction to frames and Riesz bases. Springer, 2003, vol. 7
work page 2003
-
[11]
Introduction to geometric measure theory,
L. Simon, “Introduction to geometric measure theory,”Tsinghua Lectures, 2014
work page 2014
-
[12]
F. Maggi,Sets of finite perimeter and geometric variational problems: An introduction to geometric measure theory. Cambridge University Press, 2012, no. 135
work page 2012
-
[13]
The zero set of a real analytic function,
B. Mityagin, “The zero set of a real analytic function,”Mathematical Notes, vol. 107, no. 3, pp. 529–530, 2020
work page 2020
-
[14]
On a generalization of linear independence in finite-dimensional vector spaces,
M. Feinberg, “On a generalization of linear independence in finite-dimensional vector spaces,”Journal of Combinatorial Theory, Series B, vol. 30, no. 1, pp. 61–69, 1981
work page 1981
-
[15]
J. A. Gallian,Contemporary abstract algebra, 7th ed. Brooks/Cole Cengage Learning, 2010
work page 2010
-
[16]
Inequalities for the trace of matrix product,
Y. Fang, K. A. Loparo, and X. Feng, “Inequalities for the trace of matrix product,”IEEE Transactions on Automatic Control, vol. 39, no. 12, pp. 2489–2490, 1994
work page 1994
-
[17]
S. Foucart and H. Rauhut,A mathematical introduction to compressive sensing. Birkh¨ auser New York, NY, 2013
work page 2013
-
[18]
Decoding by linear programming,
E. J. Candes and T. Tao, “Decoding by linear programming,”IEEE transactions on infor- mation theory, vol. 51, no. 12, pp. 4203–4215, 2005
work page 2005
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.