Pith. sign in

REVIEW 4 cited by

A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.03025 v2 pith:P6PJJSPY submitted 2023-02-06 cs.LG cs.AImath.RT

classification cs.LGcs.AImath.RT
keywords networkslearnuniversalityalgorithmcircuitsgroupcompositionengineering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Universality is a key hypothesis in mechanistic interpretability -- that different models learn similar features and circuits when trained on similar tasks. In this work, we study the universality hypothesis by examining how small neural networks learn to implement group composition. We present a novel algorithm by which neural networks may implement composition for any finite group via mathematical representation theory. We then show that networks consistently learn this algorithm by reverse engineering model logits and weights, and confirm our understanding using ablations. By studying networks of differing architectures trained on various groups, we find mixed evidence for universality: using our algorithm, we can completely characterize the family of circuits and features that networks learn on this task, but for a given network the precise circuits learned -- as well as the order they develop -- are arbitrary.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 11 citations worldwide. Full citation record

  1. Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters

    cs.LG 2026-07 accept novelty 6.0 of 10

    In a fully tractable 12K Llama-style model, grokking is a conditional fragile phase transition gated by coverage (tracking modulus more than structure), weight decay, and floating-point reduction order, so evidence mu...

  2. Emergent Symbolic Structure in Health Foundation Models: Extraction, Alignment, and Cross-Modal Transfer

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Independently trained PPG and accelerometer health foundation models share a linearly alignable subspace in which health-condition classifiers transfer with >95% of in-domain AUC.

  3. Circuit Stability Characterizes Language Model Generalization

    cs.CL 2025-05 reject novelty 6.0 of 10

    Circuit stability, measured as rank correlation between soft circuits across subtasks, is proposed as a predictor of language model generalization.

  4. Semantic Convergence: Investigating Shared Representations Across Scaled LLMs

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Gemma-2-2B and Gemma-2-9B align most strongly on SAE-derived features in middle layers, with preliminary evidence for shared multi-token concept subspaces.

Pith tools