Pith. sign in

REVIEW 1 cited by

Transformer Dynamics: A neuroscientific approach to interpretability of large language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12131 v1 pith:KQMDMWAG submitted 2025-02-17 cs.AI

classification cs.AI
keywords layersdynamicalmodelssystemsacrossactivationsdynamicsindividual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As artificial intelligence models have exploded in scale and capability, understanding of their internal mechanisms remains a critical challenge. Inspired by the success of dynamical systems approaches in neuroscience, here we propose a novel framework for studying computations in deep learning systems. We focus on the residual stream (RS) in transformer models, conceptualizing it as a dynamical system evolving across layers. We find that activations of individual RS units exhibit strong continuity across layers, despite the RS being a non-privileged basis. Activations in the RS accelerate and grow denser over layers, while individual units trace unstable periodic orbits. In reduced-dimensional spaces, the RS follows a curved trajectory with attractor-like dynamics in the lower layers. These insights bridge dynamical systems theory and mechanistic interpretability, establishing a foundation for a "neuroscience of AI" that combines theoretical rigor with large-scale data analysis to advance our understanding of modern neural networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Analysis of Residual-Stream Geometry Across Transformer Depth

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Across six instruction-tuned transformers, residual-stream layer transitions follow a model-specific, condition-stable depth curve: large early and late updates, a quiet middle, near-flat rotation, and a rising final ...

Pith tools