REVIEW 3 cited by
One-layer transformers fail to solve the induction heads task
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
One-layer transformers fail to solve the induction heads task
read the original abstract
A simple communication complexity argument proves that no one-layer transformer can solve the induction heads task unless its size is exponentially larger than the size sufficient for a two-layer transformer.
Forward citations
Cited by 3 Pith papers
-
Indexing: the Beginning and the End
Causal-complexity bounds show RNNs, SSMs, and masked linear attention need ω(1) layers for right-hand indexing, while a one-layer softmax transformer solves it; when the index is first, a one-layer RNN suffices.
-
Eigenvalues as a Metric for Memory Dynamics in Sequence Models
Eigenvalue spectra of attention and SSM dynamics show consistent signatures of memory retention and selective forgetting that align with task requirements.
-
Fast attention mechanisms: a tale of parallelism
ANNA, a hashing-based sub-quadratic attention mechanism, provably preserves standard attention's MPC expressiveness and can simulate low-rank attention, while being simulable by MPC with near-linear machines.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.