Pith. sign in

REVIEW 2 cited by

A phase transition between positional and semantic learning in a solvable model of dot-product attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03902 v2 pith:WV5JWCN6 submitted 2024-02-06 cs.LG

classification cs.LG
keywords attentionsemanticdot-productmechanismmodelpositionalattendingcharacterization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many empirical studies have provided evidence for the emergence of algorithmic mechanisms (abilities) in the learning of language models, that lead to qualitative improvements of the model capabilities. Yet, a theoretical characterization of how such mechanisms emerge remains elusive. In this paper, we take a step in this direction by providing a tight theoretical analysis of the emergence of semantic attention in a solvable model of dot-product attention. More precisely, we consider a non-linear self-attention layer with trainable tied and low-rank query and key matrices. In the asymptotic limit of high-dimensional data and a comparably large number of training samples we provide a tight closed-form characterization of the global minimum of the non-convex empirical loss landscape. We show that this minimum corresponds to either a positional attention mechanism (with tokens attending to each other based on their respective positions) or a semantic attention mechanism (with tokens attending to each other based on their meaning), and evidence an emergent phase transition from the former to the latter with increasing sample complexity. Finally, we compare the dot-product attention layer to a linear positional baseline, and show that it outperforms the latter using the semantic mechanism provided it has access to sufficient data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Physics of Skill Learning

    cs.LG 2025-01 conditional novelty 6.0 of 10

    The paper introduces Geometry, Resource, and Domino models that reproduce the sequential Domino effect in skill learning and link it to scaling laws, optimizers, and modularity.

  2. Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph

    cs.LG 2024-12 conditional novelty 3.0 of 10

    A didactic review showing that multi-layer autoregressive Transformers can infer position from the causal mask and context alone, so explicit positional encodings are unnecessary beyond one layer.

Pith tools