Pith. sign in

REVIEW 3 major objections 24 references

Weight-level symmetry readout sees only what the positional encoding can exactly lift.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 04:50 UTC pith:U5L5MXE5

load-bearing objection Clean PE-dependent upper bound on weight-level symmetry readout, with complementary constructions and matching 2-D experiments; soft only on the untested symmetry-sufficiency of the Gram observables. the 3 major comments →

arxiv 2607.03108 v1 pith:U5L5MXE5 submitted 2026-07-03 cs.LG cs.CV

Observable- and Positional-Encoding-Dependent Symmetry Readout from Neural Network Weights

classification cs.LG cs.CV
keywords positional encodingsymmetry readoutneural fieldsweight-space observablesexact liftabilityGram matriximplicit neural representations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Post-hoc analysis of trained neural fields often hopes to recover geometric symmetry groups straight from the weights. This paper shows that what appears is not the true symmetry of the target function, but an observable set fixed by three things at once: the trained parameters, the positional encoding, and the chosen readout. It formulates an exact hierarchy: the symmetries that can be read out are contained inside the intersection of the true symmetries and the transformations the positional encoding can lift exactly into feature space. Consequently a real geometric symmetry can be structurally invisible from weights whenever the encoding cannot represent the corresponding transformation. Experiments with MLPs trained on two-dimensional signed-distance functions confirm the prediction across several shape groups, encodings, and Gram-matrix observables: axis-aligned dyadic encodings expose D4 structure while suppressing D3 rotations, three-axis encodings reverse that preference, and random Fourier features mainly register 180-degree rotations. The result reframes positional encoding as a design choice that decides not only what a network can approximate but also which internal structures remain accessible after training.

Core claim

For positional-encoding-equipped neural fields, the symmetry visible from weights is never the true symmetry group itself; it is an observable symmetry set obeying the exact hierarchy G_obs^exact ⊆ G_lift^exact(φ) ∩ G_true. A true geometric symmetry is therefore structurally invisible to weight-level observables whenever the positional encoding cannot exactly lift the corresponding transformation into feature space.

What carries the argument

The exact observability hierarchy (Proposition 1): G_obs^exact(θ; φ, Φ) ⊆ G_lift^exact(φ) ∩ G_true, where G_lift^exact(φ) is the set of input transformations that the positional encoding can realize exactly as linear maps on its feature space. The hierarchy supplies an upper bound on every weight-level structural readout.

Load-bearing premise

That the Gram-matrix readouts used in the experiments already imply functional invariance of the network, and that approximate lifts remain faithful enough outside the exact regime for relative scores to diagnose the hierarchy.

What would settle it

Train the same MLP architecture on a D3-symmetric signed-distance function with DyadicAxisPE and measure the shallowest prefix Gram under 120-degree rotation; if that score systematically drops to the same low level obtained for D4 shapes under 90-degree rotation, the claimed structural suppression is false.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper argues that, for positional-encoding-equipped neural fields, post-hoc weight-level symmetry readout recovers not G_true but an observable set G_obs(θ;ϕ,Φ) constrained by the PE and the chosen observable. It formalizes an exact hierarchy G_obs^exact ⊆ G_lift^exact(ϕ) ∩ G_true (Proposition 1), proves exact liftability bounds for DyadicAxisPE (D4), TriAxisPE (D6⊃D3), and RFF (Z2) in Lemma 1, and tests the prediction with ReLU MLPs trained on 2D SDFs under Gram-type observables (especially prefix Gram L0). Empirically, DyadicAxisPE yields D4-aligned dips but suppresses D3, TriAxisPE lowers D3/D6 scores, and RFF mainly responds at π, consistent with the PE liftability bounds; appendices cover layers L0–L4, translation, false-positive checks, and a PE×observable grid.

Significance. If the hierarchy and PE-dependent patterns hold, the work supplies a useful structural caveat for post-hoc weight-space analysis of INRs: PE design is not only an approximation choice but also a hard upper bound on which geometric symmetries can appear in weight-level observables. The contribution is concrete rather than purely conceptual: Lemma 1 and Proposition 1 are proved under stated assumptions, Procrustes residuals reach machine precision exactly where lifts are predicted, and the experimental grid (multiple shapes per group, 5 seeds, full-layer and translation appendices, false-positive checks) produces falsifiable PE-dependent profiles. This is a solid, well-scoped addition to the literature on weight-space models, symmetry discovery, and PE representation theory (e.g., GRAPE), even if the practical impact is currently limited to 2D SDF MLPs and Gram-type readouts.

major comments (3)
  1. Definition 2 (symmetry-sufficient observable) is load-bearing for Proposition 1’s containment G_obs^exact ⊆ G_true, yet the main experiments use prefix Gram L0 / Weff Gram without verifying that Φ(Tgθ)=Φ(θ) implies functional invariance fθ(g^{-1}x)=fθ(x). Appendix B correctly distinguishes functional vs structural symmetry, but the paper should either (i) prove or empirically check sufficiency for the Gram family under the stated PE lifts, or (ii) restate Proposition 1 as a structural upper bound only and avoid claiming implication for functional G_true without that check.
  2. §3.5 and the operational regime: outside G_lift^exact the paper replaces ρ by the Procrustes ˆρ and reports relative Sop profiles (e.g., Fig. 3–4, A3). Relative profiles are informative, but the manuscript sometimes presents them as diagnosing the hierarchy itself. Please state more sharply that Dop is an operational diagnostic, not a theorem-level extension of Proposition 1, and quantify how large rP can be before relative score orderings become unreliable (especially for RFF and TriAxis D4 probes).
  3. Scope of the central empirical claim: the cleanest PE-dependent pattern is shown for prefix Gram L0 on a small set of discrete groups (D3/D4/D6). Table A1 and Fig. A4–A7 show that deeper prefixes and Weff degrade sensitivity; activation/output scores in Appendix E largely erase PE differences. The main text should more carefully bound the claim to weight-prefix Gram under structured PEs, rather than suggesting a general PE design principle for all post-hoc weight-level readouts.

Circularity Check

0 steps flagged

No significant circularity: hierarchy is definitional organization of exact-lift and symmetry-sufficient observables, with independent PE constructions and empirical tests of the predicted bounds.

full rationale

The central hierarchy (Proposition 1) follows immediately from the paper's own definitions of G_exact_obs (restricted to exact lifts where Phi is invariant), G_exact_lift (PE algebraic equivariance), and symmetry-sufficiency (Definition 2), plus a standard triangle-inequality convergence argument; the Appendix A proof is elementary and does not smuggle external results. Lemma 1 constructs or rules out linear lifts for the three fixed a-priori PEs by direct trigonometric identities and separability (no fitting). Experiments then measure operational Gram scores on independently trained models and check consistency with those liftability bounds; no parameter is fitted to data and re-labeled a prediction, and the sole external PE taxonomy citation (GRAPE) has non-overlapping authors. The mild definitional character of the hierarchy is ordinary for a formalization paper and does not force the empirical PE-dependent patterns. Score remains near zero.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 3 invented entities

The central claim rests on standard representation theory of PEs, the definition of linear lifts, and the modeling choice that Gram sandwiches capture structural symmetry; free parameters are mainly experimental thresholds and architecture choices that do not enter the hierarchy proof itself. Invented entities are definitional constructs introduced to name the observable set and the operational regime.

free parameters (3)
  • detection threshold ε = 0.05 (sensitivity 0.02–0.10)
    Hard detection uses ε≈0.05 on normalized Gram distance; primary analysis uses relative profiles, but false-positive claims and Gε_obs depend on this hand-chosen cutoff (Appendix F).
  • PE output dimension / octave counts = 48
    All structured PEs fixed to 48 dims (K=12 Dyadic, K=8 TriAxis, n=24 RFF) for controlled comparison; dimension is a design choice that affects redundancy and readout reliability (IdentityPE null case).
  • MLP depth/width and training schedule = 5×128, 2000 epochs
    Depth 5, width 128, 2000 epochs, Adam 1e-3; held fixed while PE and shape vary, but absolute score levels and residual S≈0.2 for exact lifts depend on these choices.
axioms (4)
  • domain assumption A PE admits an exact linear lift ρ(g) on feature space iff ϕ(gx)=ρ(g)ϕ(x); only then does the first-layer weight transform W0↦W0ρ(g^{-1}) realize Tg exactly.
    Stated in §3.2 and used throughout the hierarchy and Gram sandwich derivations.
  • domain assumption General-position assumption on RFF frequencies: the set {ωi} is not closed under the tested rotations/reflections except sign reversal.
    Lemma 1(iii); without it RFF could accidentally lift more angles.
  • ad hoc to paper Symmetry-sufficient observable (Def. 2): Φ(Tgθ)=Φ(θ) implies functional invariance of fθ under g.
    Required for the inclusion G_obs^exact ⊆ G_true in Proposition 1; not automatically true for every Gram construction.
  • standard math Standard facts of orthogonal Procrustes, Frobenius norms, dihedral groups D_n, and addition formulas for sine/cosine.
    Used in Lemma 1 proofs and score definitions.
invented entities (3)
  • Observable symmetry set G_obs(θ;φ,Φ) and its exact version G_obs^exact no independent evidence
    purpose: Name the symmetries that are actually stable under a chosen weight-level readout rather than the true group of the function.
    Definition 1; central object of the paper.
  • Exact-lift group G_lift^exact(φ) no independent evidence
    purpose: Upper-bound the symmetries that can ever appear as exact structural invariants of weight observables for a given PE.
    Defined in §3.4; Lemma 1 classifies it for the three PEs.
  • Operational detection set Dop and operational score Sop via Procrustes ˆρ no independent evidence
    purpose: Extend readout outside the exact-lift regime so experiments can score non-liftable transforms.
    §3.5; used for all main figures.

pith-pipeline@v1.1.0-grok45 · 32790 in / 3578 out tokens · 41764 ms · 2026-07-12T04:50:26.598948+00:00 · methodology

0 comments
read the original abstract

Post-hoc analysis of trained neural network weights often seeks to recover geometric structure directly from the parameters. We show that, for positional-encoding-equipped neural fields, the symmetry visible from weights is not the true symmetry group itself, but an observable symmetry set determined by the trained parameters, the positional encoding (PE), and readout observable. We formulate this dependence through an exact observability hierarchy, $G_{\mathrm{obs}}^{\mathrm{exact}} \subseteq G_{\mathrm{lift}}^{\mathrm{exact}}(\phi) \cap G_{\mathrm{true}}$, where $G_{\mathrm{lift}}^{\mathrm{exact}}(\phi)$ is the set of input transformations that the PE can exactly lift to the feature space. The hierarchy implies that even when a target function has a geometric symmetry, that symmetry may be structurally invisible to weight-level observables if the PE does not represent the corresponding transformation. We test this prediction using MLPs trained on two-dimensional signed distance functions with multiple shape symmetry groups, positional encodings, and Gram-based observables. The results show a consistent PE-dependent pattern: DyadicAxisPE supports $D_4$-sensitive readout but structurally suppresses $D_3$ rotations, TriAxisPE yields lower $D_3$ / $D_6$ readout scores under the tested Gram observables by replacing coordinate axes with three 120-degree-separated axes, and random Fourier features mainly exhibit a $\pi$-rotation response under these readouts. These findings show that PE design affects not only approximation behavior but also which structures are accessible to post-hoc weight-level readouts. This provides a basis for a principled observable-dependent symmetry readout.

Figures

Figures reproduced from arXiv: 2607.03108 by Naoya Chiba, Satoshi Sugiyama, Yuki Uranishi.

Figure 1
Figure 1. Figure 1: Overview of the observable-symmetry readout framework. Weight-level readout yields an [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Zero level sets of the SDFs for the 6 shapes used in the main text (red: boundary; dark: [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Prefix Gram L0 score ∆ over rotation angles for 3 shapes × 3 PEs. DyadicAxisPE shows D4-aligned dips, TriAxis gives lower scores on the D3 shape, and RFF mainly shows a π dip. Shading: ±1 std over 5 seeds. See Figure A3 and Appendix D for extended comparisons [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Prefix Gram L0 score ∆ for 3 PEs × 6 shapes. DyadicAxisPE gives the lowest π/2 scores on D4 shapes, whereas TriAxis gives the lowest 2π/3 scores on D3 shapes, though still above ε = 0.05. Results are mean ±1σ over 5 seeds. 4.2.2 Suppression of D3 Rotation Responses Under DyadicAxisPE Below, we examine a positive example (DyadicAxisPE’s π/2) and a structural limitation (D3 non￾detection) in contrast. Since … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 3 linked inside Pith

  1. [1]

    Implicit regularization in deep matrix factorization

    Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo. Implicit regularization in deep matrix factorization. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2019

  2. [2]

    Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇckovi´c

    Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veliˇckovi´c. Geometric deep learning: Grids, groups, graphs, geodesics, and gauges.arXiv preprint arXiv:2104.13478, 2021

  3. [3]

    Group equivariant convolutional networks

    Taco Cohen and Max Welling. Group equivariant convolutional networks. InInternational Conference on Machine Learning (ICML), 2016

  4. [4]

    Deep learning on implicit neural representations of shapes

    Luca De Luigi, Adriano Cardace, Riccardo Spezialetti, Pierluigi Zama Ramirez, Samuele Salti, and Luigi Di Stefano. Deep learning on implicit neural representations of shapes. In International Conference on Learning Representations (ICLR), 2023

  5. [5]

    Automatic symmetry discovery with lie algebra convolutional network

    Nima Dehmamy, Robin Walters, Yanchen Liu, Dashun Wang, and Rose Yu. Automatic symmetry discovery with lie algebra convolutional network. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2021

  6. [6]

    Emilien Dupont, Hyunjik Kim, S. M. Ali Eslami, Danilo Jimenez Rezende, and Dan Rosenbaum. From data to functa: Your data point is a function and you can treat it like one. InInternational Conference on Machine Learning (ICML), 2022

  7. [7]

    On the symmetries of deep learning models and their internal representations

    Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge. On the symmetries of deep learning models and their internal representations. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2022

  8. [8]

    Symmetry discovery for different data types.arXiv preprint arXiv:2410.09841, 2024

    Lexiang Hu, Yikang Li, and Zhouchen Lin. Symmetry discovery for different data types.arXiv preprint arXiv:2410.09841, 2024

  9. [9]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. InInterna- tional Conference on Learning Representations (ICLR), 2015

  10. [10]

    SGDR: Stochastic gradient descent with warm restarts

    Ilya Loshchilov and Frank Hutter. SGDR: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations (ICLR), 2017

  11. [11]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. InEuropean Conference on Computer Vision (ECCV), 2020

  12. [12]

    Liegg: studying learned lie group generators

    Artem Moskalev, Anna Sepliarskaia, Ivan Sosnovik, and Arnold Smeulders. Liegg: studying learned lie group generators. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2022

  13. [13]

    Equivariant architectures for learning in deep weight spaces

    Aviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya, Gal Chechik, and Haggai Maron. Equivariant architectures for learning in deep weight spaces. InInternational Conference on Machine Learning (ICML), 2023. 10

  14. [14]

    Hamprecht, Yoshua Bengio, and Aaron Courville

    Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred A. Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. InInternational Conference on Machine Learning (ICML), 2019

  15. [15]

    Random features for large-scale kernel machines

    Ali Rahimi and Benjamin Recht. Random features for large-scale kernel machines. InAnnual Conference on Neural Information Processing Systems (NIPS), 2007

  16. [16]

    E(n) equivariant graph neural networks

    Víctor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E(n) equivariant graph neural networks. InInternational Conference on Machine Learning (ICML), 2021

  17. [17]

    Saxe, James L

    Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. InInternational Conference on Learning Representations (ICLR), 2014

  18. [18]

    Schönemann

    Peter H. Schönemann. A generalized solution of the orthogonal procrustes problem.Psychome- trika, 31(1):1–10, 1966

  19. [19]

    Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2020

  20. [20]

    Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T

    Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. InAnnual Conference on Neural Information Processing Systems (NeurIPS), 2020

  21. [21]

    General e(2)-equivariant steerable cnns

    Maurice Weiler and Gabriele Cesa. General e(2)-equivariant steerable cnns. InAnnual Confer- ence on Neural Information Processing Systems (NeurIPS), 2019

  22. [22]

    Group representational position embedding

    Yifan Zhang, Zixiang Chen, Yifeng Liu, Zhen Qin, Huizhuo Yuan, Kangping Xu, Yang Yuan, Quanquan Gu, and Andrew Chi-Chih Yao. Group representational position embedding. In International Conference on Learning Representations (ICLR), 2026

  23. [23]

    Symmetry in neural network parameter spaces.Transac- tions on Machine Learning Research, 2026

    Bo Zhao, Robin Walters, and Rose Yu. Symmetry in neural network parameter spaces.Transac- tions on Machine Learning Research, 2026

  24. [24]

    Parameter symmetry potentially unifies deep learning theory.arXiv preprint arXiv:2502.05300, 2025

    Liu Ziyin, Yizhou Xu, Tomaso Poggio, and Isaac Chuang. Parameter symmetry potentially unifies deep learning theory.arXiv preprint arXiv:2502.05300, 2025. 11 A Proofs of Lemma 1 and Proposition 1 Proof of Lemma 1 (Exact liftability of structured PEs). We verify each claim in this section. For each PE, we construct (or show non-existence of) a linear mapρ(g...