REVIEW 3 major objections 24 references
Weight-level symmetry readout sees only what the positional encoding can exactly lift.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 04:50 UTC pith:U5L5MXE5
load-bearing objection Clean PE-dependent upper bound on weight-level symmetry readout, with complementary constructions and matching 2-D experiments; soft only on the untested symmetry-sufficiency of the Gram observables. the 3 major comments →
Observable- and Positional-Encoding-Dependent Symmetry Readout from Neural Network Weights
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For positional-encoding-equipped neural fields, the symmetry visible from weights is never the true symmetry group itself; it is an observable symmetry set obeying the exact hierarchy G_obs^exact ⊆ G_lift^exact(φ) ∩ G_true. A true geometric symmetry is therefore structurally invisible to weight-level observables whenever the positional encoding cannot exactly lift the corresponding transformation into feature space.
What carries the argument
The exact observability hierarchy (Proposition 1): G_obs^exact(θ; φ, Φ) ⊆ G_lift^exact(φ) ∩ G_true, where G_lift^exact(φ) is the set of input transformations that the positional encoding can realize exactly as linear maps on its feature space. The hierarchy supplies an upper bound on every weight-level structural readout.
Load-bearing premise
That the Gram-matrix readouts used in the experiments already imply functional invariance of the network, and that approximate lifts remain faithful enough outside the exact regime for relative scores to diagnose the hierarchy.
What would settle it
Train the same MLP architecture on a D3-symmetric signed-distance function with DyadicAxisPE and measure the shallowest prefix Gram under 120-degree rotation; if that score systematically drops to the same low level obtained for D4 shapes under 90-degree rotation, the claimed structural suppression is false.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that, for positional-encoding-equipped neural fields, post-hoc weight-level symmetry readout recovers not G_true but an observable set G_obs(θ;ϕ,Φ) constrained by the PE and the chosen observable. It formalizes an exact hierarchy G_obs^exact ⊆ G_lift^exact(ϕ) ∩ G_true (Proposition 1), proves exact liftability bounds for DyadicAxisPE (D4), TriAxisPE (D6⊃D3), and RFF (Z2) in Lemma 1, and tests the prediction with ReLU MLPs trained on 2D SDFs under Gram-type observables (especially prefix Gram L0). Empirically, DyadicAxisPE yields D4-aligned dips but suppresses D3, TriAxisPE lowers D3/D6 scores, and RFF mainly responds at π, consistent with the PE liftability bounds; appendices cover layers L0–L4, translation, false-positive checks, and a PE×observable grid.
Significance. If the hierarchy and PE-dependent patterns hold, the work supplies a useful structural caveat for post-hoc weight-space analysis of INRs: PE design is not only an approximation choice but also a hard upper bound on which geometric symmetries can appear in weight-level observables. The contribution is concrete rather than purely conceptual: Lemma 1 and Proposition 1 are proved under stated assumptions, Procrustes residuals reach machine precision exactly where lifts are predicted, and the experimental grid (multiple shapes per group, 5 seeds, full-layer and translation appendices, false-positive checks) produces falsifiable PE-dependent profiles. This is a solid, well-scoped addition to the literature on weight-space models, symmetry discovery, and PE representation theory (e.g., GRAPE), even if the practical impact is currently limited to 2D SDF MLPs and Gram-type readouts.
major comments (3)
- Definition 2 (symmetry-sufficient observable) is load-bearing for Proposition 1’s containment G_obs^exact ⊆ G_true, yet the main experiments use prefix Gram L0 / Weff Gram without verifying that Φ(Tgθ)=Φ(θ) implies functional invariance fθ(g^{-1}x)=fθ(x). Appendix B correctly distinguishes functional vs structural symmetry, but the paper should either (i) prove or empirically check sufficiency for the Gram family under the stated PE lifts, or (ii) restate Proposition 1 as a structural upper bound only and avoid claiming implication for functional G_true without that check.
- §3.5 and the operational regime: outside G_lift^exact the paper replaces ρ by the Procrustes ˆρ and reports relative Sop profiles (e.g., Fig. 3–4, A3). Relative profiles are informative, but the manuscript sometimes presents them as diagnosing the hierarchy itself. Please state more sharply that Dop is an operational diagnostic, not a theorem-level extension of Proposition 1, and quantify how large rP can be before relative score orderings become unreliable (especially for RFF and TriAxis D4 probes).
- Scope of the central empirical claim: the cleanest PE-dependent pattern is shown for prefix Gram L0 on a small set of discrete groups (D3/D4/D6). Table A1 and Fig. A4–A7 show that deeper prefixes and Weff degrade sensitivity; activation/output scores in Appendix E largely erase PE differences. The main text should more carefully bound the claim to weight-prefix Gram under structured PEs, rather than suggesting a general PE design principle for all post-hoc weight-level readouts.
Circularity Check
No significant circularity: hierarchy is definitional organization of exact-lift and symmetry-sufficient observables, with independent PE constructions and empirical tests of the predicted bounds.
full rationale
The central hierarchy (Proposition 1) follows immediately from the paper's own definitions of G_exact_obs (restricted to exact lifts where Phi is invariant), G_exact_lift (PE algebraic equivariance), and symmetry-sufficiency (Definition 2), plus a standard triangle-inequality convergence argument; the Appendix A proof is elementary and does not smuggle external results. Lemma 1 constructs or rules out linear lifts for the three fixed a-priori PEs by direct trigonometric identities and separability (no fitting). Experiments then measure operational Gram scores on independently trained models and check consistency with those liftability bounds; no parameter is fitted to data and re-labeled a prediction, and the sole external PE taxonomy citation (GRAPE) has non-overlapping authors. The mild definitional character of the hierarchy is ordinary for a formalization paper and does not force the empirical PE-dependent patterns. Score remains near zero.
Axiom & Free-Parameter Ledger
free parameters (3)
- detection threshold ε =
0.05 (sensitivity 0.02–0.10)
- PE output dimension / octave counts =
48
- MLP depth/width and training schedule =
5×128, 2000 epochs
axioms (4)
- domain assumption A PE admits an exact linear lift ρ(g) on feature space iff ϕ(gx)=ρ(g)ϕ(x); only then does the first-layer weight transform W0↦W0ρ(g^{-1}) realize Tg exactly.
- domain assumption General-position assumption on RFF frequencies: the set {ωi} is not closed under the tested rotations/reflections except sign reversal.
- ad hoc to paper Symmetry-sufficient observable (Def. 2): Φ(Tgθ)=Φ(θ) implies functional invariance of fθ under g.
- standard math Standard facts of orthogonal Procrustes, Frobenius norms, dihedral groups D_n, and addition formulas for sine/cosine.
invented entities (3)
-
Observable symmetry set G_obs(θ;φ,Φ) and its exact version G_obs^exact
no independent evidence
-
Exact-lift group G_lift^exact(φ)
no independent evidence
-
Operational detection set Dop and operational score Sop via Procrustes ˆρ
no independent evidence
read the original abstract
Post-hoc analysis of trained neural network weights often seeks to recover geometric structure directly from the parameters. We show that, for positional-encoding-equipped neural fields, the symmetry visible from weights is not the true symmetry group itself, but an observable symmetry set determined by the trained parameters, the positional encoding (PE), and readout observable. We formulate this dependence through an exact observability hierarchy, $G_{\mathrm{obs}}^{\mathrm{exact}} \subseteq G_{\mathrm{lift}}^{\mathrm{exact}}(\phi) \cap G_{\mathrm{true}}$, where $G_{\mathrm{lift}}^{\mathrm{exact}}(\phi)$ is the set of input transformations that the PE can exactly lift to the feature space. The hierarchy implies that even when a target function has a geometric symmetry, that symmetry may be structurally invisible to weight-level observables if the PE does not represent the corresponding transformation. We test this prediction using MLPs trained on two-dimensional signed distance functions with multiple shape symmetry groups, positional encodings, and Gram-based observables. The results show a consistent PE-dependent pattern: DyadicAxisPE supports $D_4$-sensitive readout but structurally suppresses $D_3$ rotations, TriAxisPE yields lower $D_3$ / $D_6$ readout scores under the tested Gram observables by replacing coordinate axes with three 120-degree-separated axes, and random Fourier features mainly exhibit a $\pi$-rotation response under these readouts. These findings show that PE design affects not only approximation behavior but also which structures are accessible to post-hoc weight-level readouts. This provides a basis for a principled observable-dependent symmetry readout.
Figures
Reference graph
Works this paper leans on
-
[1]
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo. Implicit regularization in deep matrix factorization. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2019
2019
-
[2]
Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇckovi´c
Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇckovi´c. Geometric deep learning: Grids, groups, graphs, geodesics, and gauges.arXiv preprint arXiv:2104.13478, 2021
Pith/arXiv arXiv 2021
-
[3]
Group equivariant convolutional networks
Taco Cohen and Max Welling. Group equivariant convolutional networks. InInternational Conference on Machine Learning (ICML), 2016
2016
-
[4]
Deep learning on implicit neural representations of shapes
Luca De Luigi, Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, and Luigi Di Stefano. Deep learning on implicit neural representations of shapes. In International Conference on Learning Representations (ICLR), 2023
2023
-
[5]
Automatic symmetry discovery with lie algebra convolutional network
Nima Dehmamy, Robin Walters, Yanchen Liu, Dashun Wang, and Rose Yu. Automatic symmetry discovery with lie algebra convolutional network. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2021
2021
-
[6]
Emilien Dupont, Hyunjik Kim, S. M. Ali Eslami, Danilo Jimenez Rezende, and Dan Rosenbaum. From data to functa: Your data point is a function and you can treat it like one. InInternational Conference on Machine Learning (ICML), 2022
2022
-
[7]
On the symmetries of deep learning models and their internal representations
Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge. On the symmetries of deep learning models and their internal representations. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2022
2022
-
[8]
Symmetry discovery for different data types.arXiv preprint arXiv:2410.09841, 2024
Lexiang Hu, Yikang Li, and Zhouchen Lin. Symmetry discovery for different data types.arXiv preprint arXiv:2410.09841, 2024
Pith/arXiv arXiv 2024
-
[9]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. InInterna- tional Conference on Learning Representations (ICLR), 2015
2015
-
[10]
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations (ICLR), 2017
2017
-
[11]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. InEuropean Conference on Computer Vision (ECCV), 2020
2020
-
[12]
Liegg: studying learned lie group generators
Artem Moskalev, Anna Sepliarskaia, Ivan Sosnovik, and Arnold Smeulders. Liegg: studying learned lie group generators. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2022
2022
-
[13]
Equivariant architectures for learning in deep weight spaces
Aviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya, Gal Chechik, and Haggai Maron. Equivariant architectures for learning in deep weight spaces. InInternational Conference on Machine Learning (ICML), 2023. 10
2023
-
[14]
Hamprecht, Yoshua Bengio, and Aaron Courville
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. InInternational Conference on Machine Learning (ICML), 2019
2019
-
[15]
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht. Random features for large-scale kernel machines. InAnnual Conference on Neural Information Processing Systems (NIPS), 2007
2007
-
[16]
E(n) equivariant graph neural networks
Víctor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E(n) equivariant graph neural networks. InInternational Conference on Machine Learning (ICML), 2021
2021
-
[17]
Saxe, James L
Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. InInternational Conference on Learning Representations (ICLR), 2014
2014
-
[18]
Schönemann
Peter H. Schönemann. A generalized solution of the orthogonal procrustes problem.Psychome- trika, 31(1):1–10, 1966
1966
-
[19]
Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2020
2020
-
[20]
Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T
Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2020
2020
-
[21]
General e(2)-equivariant steerable cnns
Maurice Weiler and Gabriele Cesa. General e(2)-equivariant steerable cnns. InAnnual Confer- ence on Neural Information Processing Systems (NeurIPS), 2019
2019
-
[22]
Group representational position embedding
Yifan Zhang, Zixiang Chen, Yifeng Liu, Zhen Qin, Huizhuo Yuan, Kangping Xu, Yang Yuan, Quanquan Gu, and Andrew Chi-Chih Yao. Group representational position embedding. In International Conference on Learning Representations (ICLR), 2026
2026
-
[23]
Symmetry in neural network parameter spaces.Transac- tions on Machine Learning Research, 2026
Bo Zhao, Robin Walters, and Rose Yu. Symmetry in neural network parameter spaces.Transac- tions on Machine Learning Research, 2026
2026
-
[24]
Parameter symmetry potentially unifies deep learning theory.arXiv preprint arXiv:2502.05300, 2025
Liu Ziyin, Yizhou Xu, Tomaso Poggio, and Isaac Chuang. Parameter symmetry potentially unifies deep learning theory.arXiv preprint arXiv:2502.05300, 2025. 11 A Proofs of Lemma 1 and Proposition 1 Proof of Lemma 1 (Exact liftability of structured PEs). We verify each claim in this section. For each PE, we construct (or show non-existence of) a linear mapρ(g...
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.