Pith. sign in

REVIEW 4 cited by

Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.05409 v1 pith:4M2DLY3C submitted 2025-05-08 cs.LG

Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It

classification cs.LG
keywords sharpnesssymmetriescorrelationgeneralizationtransformersexistingfullygeodesic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The concept of sharpness has been successfully applied to traditional architectures like MLPs and CNNs to predict their generalization. For transformers, however, recent work reported weak correlation between flatness and generalization. We argue that existing sharpness measures fail for transformers, because they have much richer symmetries in their attention mechanism that induce directions in parameter space along which the network or its loss remain identical. We posit that sharpness must account fully for these symmetries, and thus we redefine it on a quotient manifold that results from quotienting out the transformer symmetries, thereby removing their ambiguities. Leveraging tools from Riemannian geometry, we propose a fully general notion of sharpness, in terms of a geodesic ball on the symmetry-corrected quotient manifold. In practice, we need to resort to approximating the geodesics. Doing so up to first order yields existing adaptive sharpness measures, and we demonstrate that including higher-order terms is crucial to recover correlation with generalization. We present results on diagonal networks with synthetic data, and show that our geodesic sharpness reveals strong correlation for real-world transformers on both text and image classification tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Link between Shock-wave Theory and Symmetry-reduced Stochastic Gradient Descent for Artificial Neural Networks

    cs.LG 2026-06 unverdicted novelty 7.0

    Symmetry-quotiented SGD dynamics in neural networks satisfy viscous Hamilton-Jacobi and Burgers-type equations after local-entropy coarse-graining, with rigorous shock formation under a gradient-field assumption.

  2. Generalizing the Geometry of Model Merging Through Frechet Averages

    cs.LG 2026-04 unverdicted novelty 7.0

    Model merging is reframed as Fréchet averaging on manifolds whose geometry respects architectural symmetries, generalizing Fisher merging and enabling better LoRA merges.

  3. Generalizing the Geometry of Model Merging Through Frechet Averages

    cs.LG 2026-04 unverdicted novelty 7.0

    Model merging is generalized as Fréchet averaging on symmetry-invariant manifolds, containing Fisher merging as a special case and offering a new approach for LoRA adapters.

  4. Toward Manifest Relationality in Transformers via Symmetry Reduction

    cs.LG 2026-02 conditional novelty 6.0

    Transformer attention and parameter optimization can be rewritten on symmetry-reduced relational variables (Gram matrices and invariant parameter composites), removing coordinate redundancies by construction.