Pith. sign in

REVIEW 1 cited by

Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02362 v1 pith:LDBGX2EQ submitted 2024-02-04 cs.LG cs.AIhep-thphysics.comp-ph

Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE

classification cs.LG cs.AIhep-thphysics.comp-ph
keywords neuralsymmetriesgaugelearningnetworksodesfeedforwardmachine
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Understanding the inner workings of neural networks, including transformers, remains one of the most challenging puzzles in machine learning. This study introduces a novel approach by applying the principles of gauge symmetries, a key concept in physics, to neural network architectures. By regarding model functions as physical observables, we find that parametric redundancies of various machine learning models can be interpreted as gauge symmetries. We mathematically formulate the parametric redundancies in neural ODEs, and find that their gauge symmetries are given by spacetime diffeomorphisms, which play a fundamental role in Einstein's theory of gravity. Viewing neural ODEs as a continuum version of feedforward neural networks, we show that the parametric redundancies in feedforward neural networks are indeed lifted to diffeomorphisms in neural ODEs. We further extend our analysis to transformer models, finding natural correspondences with neural ODEs and their gauge symmetries. The concept of gauge symmetries sheds light on the complex behavior of deep learning models through physics and provides us with a unifying perspective for analyzing various machine learning architectures.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries

    cs.LG 2026-07 conditional novelty 6.0

    A training-time rotation loss learns quantization-friendly internal bases for LLMs, lowering LLaMA-2 7B W4A4 perplexity from 8.22 to 6.73.