Pith. sign in

REVIEW 3 major objections 6 minor 18 references

Filters must meet data on enough frequencies to maximize separation

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-08 18:02 UTC pith:NUVBHZGR

load-bearing objection Scattering network filter design criteria derived from geometric measure theory; lower bounds need polynomial parametrization assumption the 3 major comments →

arxiv 2607.06048 v1 pith:NUVBHZGR submitted 2026-07-07 stat.ML cs.ITcs.LGmath.CAmath.IT

Separation Capacity of Scattering Networks on Low-Dimensional Datasets

classification stat.ML cs.ITcs.LGmath.CAmath.IT
keywords datacapacitynetworksscatteringseparationdesignfiltersframe
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks a precise question: given a scattering network with a fixed nonlinearity and no pooling, how should you choose the convolutional filters to maximize the network's ability to separate (i.e., linearly classify) data that lives on a low-dimensional structure inside a high-dimensional space? The authors model low-dimensional data as rectifiable sets from geometric measure theory, which generalizes smooth manifolds and unions of linear subspaces (including sparse-signal models). Using Cover's separation capacity, they prove two-sided bounds that connect the network's classification power to the geometry of the data. The upper bound says the separation capacity is limited by the rank of a second-moment matrix that couples the feature extractor to the global geometry of the dataset. The lower bound says the separation capacity is at least twice the essential infimum of the rank of the differential of the feature map restricted to approximate tangent spaces of the data. When particularized to scattering networks, these bounds yield two concrete filter-design criteria. First, the spectral supports of the filters, when intersected with the data's spectrum, should not be contained in a coset of a proper subgroup of the frequency group. Second, certain filter-dependent matrices that couple the frame to the data geometry should be as well-conditioned as possible. For sparse signals, the second criterion reduces to minimizing a restricted isometry constant.

Core claim

The central mechanism is a pair of bounds on the s-separation capacity SC_s of a feature extractor on a countably H^s-rectifiable set. The lower bound is governed by the rank of the differential of the feature map on approximate tangent spaces (local geometry), while the upper bound is governed by the rank of a second-moment matrix formed by pushing forward the Hausdorff measure through the feature map (global geometry). For scattering networks specifically, the upper bound translates into a requirement on the spectral overlap between filters and data, and the lower bound translates into a conditioning requirement on matrices that couple the filter frame to the parametrization of the data. A

What carries the argument

The key objects are: (1) countably H^s-rectifiable sets as a model for low-dimensional data; (2) Cover's s-separation capacity as the measure of classification power; (3) approximate tangent spaces T_f E of rectifiable sets; (4) the second-moment matrix C_Phi = integral of y y^T d Phi_sharp(H^s restricted to E); (5) the vector Veronese map v_{M,d}, which encodes the effect of the monomial nonlinearity of degree d; (6) the matrix A_Lambda constructed from filter Fourier coefficients, which couples the filters to the Veronese lift; and (7) the matrices T_{hat Psi_j}^H A_{Lambda,chi} T_{hat Psi_j}, whose conditioning controls the lower bound.

Load-bearing premise

The lower bounds, which yield the actionable conditioning-based design criteria, require the bi-Lipschitz parametrizations of the rectifiable data set to be polynomial maps. Many natural low-dimensional data structures (smooth manifolds with non-algebraic parametrizations, fractal sets) may not admit such polynomial bi-Lipschitz parametrizations, limiting the scope of the design recommendations.

What would settle it

If one could exhibit a rectifiable dataset E with a non-polynomial bi-Lipschitz parametrization where the conditioning-based design criterion fails to correlate with actual separation capacity, the practical applicability of the lower-bound design recommendations would be called into question.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Filter design for scattering networks on low-dimensional data can be guided by two checkable conditions: spectral coverage of the data spectrum and numerical conditioning of data-dependent matrices, rather than by empirical tuning alone.
  • For sparse-signal models, the design criterion reduces to minimizing a restricted isometry constant, directly connecting scattering network filter design to the compressed sensing literature.
  • The upper bound's group-theoretic condition (spectral supports not in a coset of a proper subgroup) provides a concrete, checkable necessary condition for a filter bank to achieve maximum separation capacity.
  • The framework applies to any Lipschitz feature extractor on rectifiable sets, so the geometric bounds could in principle be applied to feature extractors beyond scattering networks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The polynomial parametrization assumption required for the lower bounds may exclude natural data manifolds with non-algebraic structure (e.g., smooth manifolds parametrized by transcendental maps). Extending the lower-bound argument to general bi-Lipschitz parametrizations would broaden the applicability of the conditioning-based design criterion.
  • The framework treats the nonlinearity degree d and network depth n_d as fixed; the interaction between depth and the conditioning of the coupling matrices is not fully explored and could reveal depth-dependent tradeoffs in separation capacity.
  • The separation capacity as defined measures linear separability of the network's output features. A natural extension would investigate whether the geometric bounds transfer to nonlinear classifiers applied on top of the scattering features.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper studies the separation capacity of pooling-free scattering networks with fixed monomial nonlinearities on data modeled as countably H^s-rectifiable sets. The authors first establish general bounds on the s-separation capacity of feature extractors: a lower bound in terms of the rank of the differential restricted to approximate tangent spaces (Lemma 2.7) and an upper bound via a second-moment matrix (Lemma 2.9), with an exact formula under real-analyticity (Corollary 2.9.1). These are then particularized to scattering networks for two data models: sparse signals and polynomially parametrized rectifiable sets. The resulting design criteria are: (i) filter spectral supports should not be contained in a coset of a proper subgroup (Theorem 3.3), and (ii) matrices coupling the filter frame to the data geometry should be well-conditioned (Theorems 3.5, 3.8 and Corollaries 3.5.1, 3.8.1). The proofs proceed through a clean chain from the subspace characterization (Lemma 2.2) through geometric measure theory tools to concrete matrix conditions.

Significance. The paper provides a mathematically rigorous bridge between geometric measure theory and the separation capacity of scattering networks, yielding actionable filter-design criteria expressed in terms of well-conditioning of data-dependent matrices. The derivation chain is carefully executed: the area formula application in Corollary 2.9.1, the Veronese map factorization in Lemma 3.2, and the restricted isometry constant formulation in Theorem 3.5 are notable technical contributions. The reduction of the sparse-signal lower bound to an RIP condition on (T^H_{bXi} A_{Lambda,chi} T_{bXi})^{1/2} is a clean, falsifiable criterion. The upper-bound design criterion (spectral supports not in a coset of a proper subgroup) is parameter-free in the sense that it depends only on the group structure of Z/MZ and the support sets. The paper is self-contained in its proof structure, though it relies on two companion papers [7, 9] for background results on Cover-type separation capacity.

major comments (3)
  1. Section 3.3, Theorem 3.8: The lower bound for general rectifiable sets depends on the assumption that each bi-Lipschitz parametrization psi_j is a polynomial of degree n_j. This assumption is load-bearing: the factorization v_{M,d}(hat{psi}_j(x)) = T_{hat{psi}_j} w_{s,n_j d}(x) in the proof of Theorem 3.8 requires the Veronese lift of hat{psi}_j to land in a finite-dimensional monomial space. If psi_j is not polynomial, T_{hat{psi}_j} is not a finite-dimensional object and the trace bounds in (3.23) break down. The paper does not discuss which natural data models satisfy this polynomial assumption, nor does it acknowledge that smooth manifolds with non-algebraic parametrizations (a primary motivating example per the introduction's reference to [1, 2]) are excluded. This is a scope limitation on the central actionable design criterion (Corollary 3.8.1). The authors should either (a) add a
  2. Section 3.3, Theorem 3.7: The upper bound also relies on the polynomial assumption to bound dim_C(span_C(v_{M,d}(F_M(E_j)))) <= C(s + n_j d, n_j d) via the degree of x -> v_{M,d}(F_M(psi_j(x))). While the upper bound is less directly actionable, the same scope concern applies. The paper should clarify whether the polynomial assumption is essential for the upper bound as well, or whether a more general bound (e.g., in terms of the Hausdorff dimension of the Fourier support of psi_j(K_j)) could replace it.
  3. Theorem 3.5 and Corollary 3.5.1: The lower bound for sparse signals involves the term inf_{j in J_S} 2(Tr(C_{S,j}))^2 / Tr((C_{S,j})^2), which depends on the parametrization domains K_{S,j} but not on the frame Xi. The paper does not discuss whether this term can be bounded below in terms of s and d alone, or whether it can be arbitrarily small for adversarial choices of K_{S,j}. If the latter, the practical utility of the design criterion (minimizing delta_{s,d}(Xi; Lambda, chi)) is diminished, as the overall lower bound could still be small. A brief remark on the behavior of this term would strengthen the result.
minor comments (6)
  1. The notation T_{hat{psi}_j} in Theorem 3.8 uses a hat on psi_j, but the text in Section 3.3 defines psi_j without a hat. The hat presumably denotes the Fourier transform, but this should be stated explicitly.
  2. In the proof of Lemma 3.2, the matrix A is defined with the condition supp(alpha) subseteq H_{lambda,S}, but in the subsequent definition of A_lambda (footnote 2 on page 11), this condition is dropped. The relationship between A and A_lambda should be stated more explicitly to avoid confusion.
  3. Theorem 3.3: The upper bound includes a term 4 min_{S} [...] rather than 2 min_{S} [...]. The factor of 2 relative to the real-valued case is presumably due to the complex-to-real identification, but this should be noted explicitly, perhaps with a reference to equation (3.3).
  4. In the proof of Theorem 3.8(b), the chain of inequalities bounding Tr(G_j A_Lambda G_j A_Lambda) uses the Loewner ordering notation (preceq, succeq) without explicitly defining it. While standard, a brief note would aid readability.
  5. Reference [9] (arXiv:2607.01010) is cited for the identity SC_s(Phi) = min_{j in J} SC_s(Phi|_{psi_j(K_j)}) used in Corollary 2.9.1, and for the identity in (2.2). Since these are central to the derivation, the reader would benefit from a brief statement of the relevant results from [9] rather than relying entirely on the companion paper.
  6. The abstract states 'no pooling' but Section 3.1 defines the general scattering network with pooling operators P_n. The restriction to P_n = Id is stated in Section 3.1 but could be mentioned in the abstract for precision.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for a careful reading and for identifying a genuine scope limitation in Section 3.3 that we will address in revision. The referee's three major comments are interconnected: the first two concern the polynomial parametrization assumption in Theorems 3.7 and 3.8, and the third concerns the parametrization-domain-dependent term in the sparse-signal lower bound. We agree that the polynomial assumption is load-bearing and that the manuscript does not adequately discuss its scope or which data models satisfy it. We will add a dedicated remark addressing this. On the third comment, we will add discussion of the behavior of the trace-ratio term, including both the lower bound it admits and the adversarial cases the referee identifies.

read point-by-point responses
  1. Referee: Section 3.3, Theorem 3.8: The lower bound depends on the polynomial parametrization assumption. The paper does not discuss which natural data models satisfy this, nor that smooth manifolds with non-algebraic parametrizations are excluded. The authors should either (a) add a discussion of scope, or (b) extend to non-polynomial parametrizations.

    Authors: The referee is correct that the polynomial assumption is load-bearing for Theorem 3.8. Specifically, the factorization v_{M,d}(hat{psi}_j(x)) = T_{hat{psi}_j} w_{s,n_j d}(x) in the proof requires the Veronese lift of hat{psi}_j to land in a finite-dimensional monomial space, which fails for non-polynomial psi_j. We acknowledge that this excludes smooth manifolds with non-algebraic parametrizations, which the introduction's reference to [1, 2] might suggest are covered. We will add a dedicated remark in Section 3.3 explicitly stating this scope limitation, clarifying that the results of Theorem 3.8 and Corollary 3.8.1 apply to polynomially parametrized rectifiable sets (which include sparse-signal models as a special case, as well as algebraic varieties and their finite unions). Regarding option (b), extending to non-polynomial parametrizations would require replacing the finite-dimensional matrix T_{hat{psi}_j} with an operator acting on an infinite-dimensional function space, and the trace bounds in (3.23) would no longer apply in their current form. We believe this extension is beyond the scope of the present paper, but we will note it as a direction for future work. We will also adjust the introduction to avoid implying that general smooth submanifolds are covered by the polynomial results. revision: yes

  2. Referee: Section 3.3, Theorem 3.7: The upper bound also relies on the polynomial assumption to bound dim_C(span_C(v_{M,d}(F_M(E_j)))) <= C(s + n_j d, n_j d). The paper should clarify whether the polynomial assumption is essential for the upper bound, or whether a more general bound could replace it.

    Authors: The referee correctly identifies that the bound dim_C(span_C(v_{M,d}(F_M(E_j)))) <= C(s + n_j d, n_j d) in Lemma 3.6 relies on x -> v_{M,d}(F_M(psi_j(x))) being a multivariate polynomial of degree at most n_j d, which requires psi_j to be polynomial. For the upper bound, however, the situation is somewhat different from the lower bound. The term |H_{d,lambda,psi_j} cap supp(hat{chi})| in (3.21) does not depend on the polynomial assumption and remains valid for general bi-Lipschitz parametrizations. Only the combinatorial term C(s + n_j d, n_j d) requires polynomial structure. For a general (non-polynomial) bi-Lipschitz parametrization, the Veronese lift v_{M,d}(hat{psi}_j(x)) need not lie in any finite-dimensional subspace, so the combinatorial bound breaks down. A replacement bound in terms of the Hausdorff dimension of the Fourier support of psi_j(K_j) is an interesting possibility, but it would require a fundamentally different proof technique, as the current argument proceeds through finite-dimensional polynomial degree counting. We will add a remark clarifying that the polynomial assumption is essential for the combinatorial term in the upper bound, while the spectral-support term is not, and that extending the upper bound to non-polynomial parametrizations is left open. revision: yes

  3. Referee: Theorem 3.5 and Corollary 3.5.1: The term inf_{j in J_S} 2(Tr(C_{S,j}))^2 / Tr((C_{S,j})^2) depends on the parametrization domains K_{S,j} but not on the frame Xi. The paper does not discuss whether this term can be bounded below in terms of s and d alone, or whether it can be arbitrarily small for adversarial choices of K_{S,j}.

    Authors: The referee raises a valid point. The term 2(Tr(C_{S,j}))^2 / Tr((C_{S,j})^2) equals 2 times the effective rank of C_{S,j}, which is at least 2 and at most 2*C(s-1+d, d). For well-behaved parametrization domains (e.g., K_{S,j} containing an open ball in R^{2s}), C_{S,j} is full-rank and the term equals 2*C(s-1+d, d), the maximum possible value. However, the referee is correct that for adversarial choices of K_{S,j} — for instance, domains concentrated near a lower-dimensional subset — the matrix C_{S,j} can become rank-deficient and the term can be arbitrarily small. We will add a remark noting both the upper bound 2*C(s-1+d, d) (achieved for full-dimensional domains) and the fact that the term can degenerate for adversarial domains. We would also note that in the typical setting where the parametrization domains are fixed by the data model (not chosen adversarially), this term is a geometric constant of the dataset, and the design criterion of minimizing delta_{s,d}(Xi; Lambda, chi) remains the actionable lever for the practitioner, as the frame Xi is the only design variable. revision: yes

Circularity Check

0 steps flagged

No significant circularity; two self-citations to companion papers are used as foundational definitions, but the central bounds and design criteria are derived independently from standard GMT results.

full rationale

The paper's central results—Lemma 2.7 (lower bound via tangent space rank), Lemma 2.9 (upper bound via second-moment matrix), and the scattering-network design criteria (Theorems 3.3, 3.5, 3.7, 3.8)—are derived from first principles using standard geometric measure theory results [3, 4, 11, 12]. The two self-citations [7, 9] are used to import the definition of separation capacity (Definition 2.1, citing [9]) and the identity SC_s(Φ) = 2·min dim(span(Φ(A))) (Eq. 2.2, citing [9]). These are foundational definitions rather than results being 'predicted' or 'derived' in this paper. The actual bounds in Lemmas 2.7 and 2.9 are proven self-containedly: Lemma 2.7 uses approximate tangent space theory [12] and does not rely on [9] for its proof; Lemma 2.9 uses a kernel/support argument independent of [9]. The particularization to scattering networks (Section 3) uses the Veronese map, area formula, and RIP theory [17, 18]—all external. The self-citation [9] provides the function-counting framework (analogous to how Cover's theorem [8] is cited), and the present paper builds GMT-based bounds on top of it. The identity in Eq. 2.2 is a definition/import, not a circular 'prediction.' The polynomial parametrization assumption in Section 3.3 is a scope limitation (correctness risk), not a circularity issue. No step in the derivation chain reduces to its own inputs by construction. The self-citations are standard companion-paper references for definitions, not load-bearing circular chains. Score 2 reflects the presence of self-cited definitions without independent verification, but the central technical content is independently derived.

Axiom & Free-Parameter Ledger

3 free parameters · 5 axioms · 0 invented entities

The paper introduces no new physical entities, particles, forces, or dimensions. All mathematical objects (rectifiable sets, approximate tangent spaces, frames, Veronese maps, restricted isometry constants) are standard in their respective literatures. The matrices A_Λ, G_{S,j}, C_{S,j}, T_{ξ̂}, M_j are defined constructions, not postulated entities. The free parameters (d, n_d, t) are design choices or data properties, not fitted constants.

free parameters (3)
  • Monomial degree d
    The nonlinearity ρ(z) = z^d is fixed but d is a free design parameter; the bounds depend on d throughout (e.g., binomial coefficients (s-1+d choose d) in Lemma 3.2).
  • Network depth n_d
    The scattering network depth n_d is a free parameter appearing in the upper bounds of Theorems 3.3 and 3.7.
  • Bi-Lipschitz constant t
    The bi-Lipschitz constant t of the parametrization appears in the lower bound of Theorem 3.8 as t^{-4s}; it is a property of the data model, not fitted, but enters as a parameter.
axioms (5)
  • domain assumption Data lies on a countably H^s-rectifiable set (Definition 2.4)
    Invoked in Section 2 to connect separation capacity to geometric measure theory. This is standard in GMT but a modeling choice for the data.
  • domain assumption The feature extractor Φ is Lipschitz (Lemma 2.7)
    Required for the tangent-space lower bound; scattering networks with frame bounds and Lipschitz nonlinearities satisfy this.
  • ad hoc to paper The bi-Lipschitz parametrization {ψ_j} consists of polynomials of degree n_j (Section 3.3)
    Introduced to make the lower bounds tractable via the Veronese decomposition. Not required for upper bounds. Restricts the class of rectifiable sets to those admitting polynomial parametrizations.
  • standard math The s-separation capacity equals 2·min over positive-measure subsets A of dim_R(span_R(Φ(A))) (Eq. 2.2, attributed to [9])
    This is the foundational identity from Cover-type function-counting theory, attributed to the companion paper [9] but proved self-containedly in Lemma 2.2.
  • domain assumption Real-analytic extension of Φ∘ψ_j to a connected open neighborhood (Corollary 2.9.1)
    Required for the exact formula in Corollary 2.9.1 via the zero-set property of real-analytic functions [13]. Scattering networks with monomial nonlinearities satisfy this.

pith-pipeline@v1.1.0-glm · 22965 in / 3544 out tokens · 346742 ms · 2026-07-08T18:02:58.995594+00:00 · methodology

0 comments
read the original abstract

We aim to identify scattering network architectures that maximize the separation capacity on data with low intrinsic dimension. The networks we consider employ a fixed monomial nonlinearity and no pooling, so that the only design variable is the frame generated by the network filters. For data modeled as rectifiable sets, we first characterize and bound the separation capacity of general feature extractors in terms of the geometry of the dataset. We then particularize to scattering networks and obtain two design criteria: (i) the filters should meet the data on sufficiently many frequencies, and (ii) the matrices coupling the frame to the geometry of the data should be well-conditioned.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

18 extracted references · 18 canonical work pages · 2 internal anchors

  1. [1]

    Representation learning: A review and new perspectives,

    Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,”IEEE transactions on pattern analysis and machine intelligence, vol. 35, no. 8, pp. 1798–1828, 2013

  2. [2]

    Testing the manifold hypothesis,

    C. Fefferman, S. Mitter, and H. Narayanan, “Testing the manifold hypothesis,”Journal of the American Mathematical Society, vol. 29, no. 4, pp. 983–1049, 2016

  3. [3]

    Federer,Geometric measure theory

    H. Federer,Geometric measure theory. Springer, 2014

  4. [4]

    Currents in metric spaces,

    L. Ambrosio and B. Kirchheim, “Currents in metric spaces,”Acta Mathematica, vol. 185, pp. 1–80, 2000

  5. [5]

    Group invariant scattering,

    S. Mallat, “Group invariant scattering,”Communications on Pure and Applied Mathemat- ics, vol. 65, no. 10, pp. 1331–1398, 2012. Separation Capacity of Scattering Networks on Low-Dimensional Datasets 19

  6. [6]

    A mathematical theory of deep convolutional neural net- works for feature extraction,

    T. Wiatowski and H. B¨ olcskei, “A mathematical theory of deep convolutional neural net- works for feature extraction,”IEEE Transactions on Information Theory, vol. 64, no. 3, pp. 1845–1866, 2017

  7. [7]

    Separation Capacity of Scattering Networks

    K. H¨ aberle and H. B¨ olcskei, “Separation capacity of scattering networks,”arXiv preprint arXiv:2606.30822, 2026

  8. [8]

    Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition,

    T. M. Cover, “Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition,”IEEE Transactions on Electronic Computers, no. 3, pp. 326–334, 1965

  9. [9]

    Function-Counting Theory for Low-Dimensional Data Structures

    K. H¨ aberle and H. B¨ olcskei, “Function-counting theory for low-dimensional data struc- tures,”arXiv preprint arXiv:2607.01010, 2026

  10. [10]

    Christensen,An introduction to frames and Riesz bases

    O. Christensen,An introduction to frames and Riesz bases. Springer, 2003, vol. 7

  11. [11]

    Introduction to geometric measure theory,

    L. Simon, “Introduction to geometric measure theory,”Tsinghua Lectures, 2014

  12. [12]

    Maggi,Sets of finite perimeter and geometric variational problems: An introduction to geometric measure theory

    F. Maggi,Sets of finite perimeter and geometric variational problems: An introduction to geometric measure theory. Cambridge University Press, 2012, no. 135

  13. [13]

    The zero set of a real analytic function,

    B. Mityagin, “The zero set of a real analytic function,”Mathematical Notes, vol. 107, no. 3, pp. 529–530, 2020

  14. [14]

    On a generalization of linear independence in finite-dimensional vector spaces,

    M. Feinberg, “On a generalization of linear independence in finite-dimensional vector spaces,”Journal of Combinatorial Theory, Series B, vol. 30, no. 1, pp. 61–69, 1981

  15. [15]

    J. A. Gallian,Contemporary abstract algebra, 7th ed. Brooks/Cole Cengage Learning, 2010

  16. [16]

    Inequalities for the trace of matrix product,

    Y. Fang, K. A. Loparo, and X. Feng, “Inequalities for the trace of matrix product,”IEEE Transactions on Automatic Control, vol. 39, no. 12, pp. 2489–2490, 1994

  17. [17]

    Foucart and H

    S. Foucart and H. Rauhut,A mathematical introduction to compressive sensing. Birkh¨ auser New York, NY, 2013

  18. [18]

    Decoding by linear programming,

    E. J. Candes and T. Tao, “Decoding by linear programming,”IEEE transactions on infor- mation theory, vol. 51, no. 12, pp. 4203–4215, 2005