Pith. sign in

REVIEW 24 cited by

Why do LLMs attend to the first token?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02732 v4 pith:CJBGWDZR submitted 2025-04-03 cs.CL

Why do LLMs attend to the first token?

classification cs.CL
keywords attentionllmsattendfirstmanypatternsprovidessink
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large Language Models (LLMs) tend to attend heavily to the first token in the sequence -- creating a so-called attention sink. Many works have studied this phenomenon in detail, proposing various ways to either leverage or alleviate it. Attention sinks have been connected to quantisation difficulties, security issues, and streaming attention. Yet, while many works have provided conditions in which they occur or not, a critical question remains shallowly answered: Why do LLMs learn such patterns and how are they being used? In this work, we argue theoretically and empirically that this mechanism provides a method for LLMs to avoid over-mixing, connecting this to existing lines of work that study mathematically how information propagates in Transformers. We conduct experiments to validate our theoretical intuitions and show how choices such as context length, depth, and data packing influence the sink behaviour. We hope that this study provides a new practical perspective on why attention sinks are useful in LLMs, leading to a better understanding of the attention patterns that form during training.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

    cs.CV 2026-07 conditional novelty 7.0

    Chat-template tokens in LLM-conditioned DiTs act as implicit semantic registers: they absorb object identity from image latents and maintain it, while direct prompt-reading heads are causally inert.

  2. Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference

    cs.CL 2026-07 conditional novelty 7.0

    Heavy-hitter KV cache eviction over-retains structural delimiters and keys on schema-dense inputs; a role-conditional reallocation of SnapKV's score recovers most of the accuracy collapse.

  3. Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test

    cs.LG 2026-06 unverdicted novelty 7.0

    In 160M and 290M parameter models, a new residual-stream split into scratch and protected channels causes massive activations to re-emerge in the protected decode channel, more concentrated on the start token.

  4. ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection

    cs.LG 2026-05 conditional novelty 7.0

    ASAP amortizes Sinkhorn-based doubly-stochastic attention by learning a parametric map from 1D potentials to the Sinkhorn dual and reconstructing the plan via two-sided entropic c-transform, delivering 5.3x faster inf...

  5. A Mechanistic Analysis of Looped Reasoning Language Models

    cs.LG 2026-04 unverdicted novelty 7.0

    Looped LLMs converge to distinct cyclic fixed points per layer, repeating feedforward-style inference stages across recurrences.

  6. Perceptrons and localization of attention's mean-field landscape

    cs.LG 2026-01 unverdicted novelty 7.0

    In the mean-field limit of attention with perceptron blocks, critical points of the energy landscape are generically atomic and localized on subsets of the unit sphere.

  7. Selective KV Cache Protection for Noise-Resilient LLM Inference on Analog Compute-In-Memory Systems

    cs.AR 2026-07 conditional novelty 6.0

    A hybrid analog-digital KV cache that protects sink and recent tokens digitally, while keeping the bulk in analog CIM, lowers LLM perplexity under hardware noise from 33.9 to 11.95, near the clean 11.06 baseline.

  8. Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

    cs.LG 2026-05 unverdicted novelty 6.0

    Contribution Weights combine attention, value magnitude, and directional alignment to measure token influence more faithfully than attention alone, and show attention sinks actively suppress information via a convex s...

  9. Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models

    cs.AI 2026-05 unverdicted novelty 6.0

    An attention-guided RL reward combined with diverse persuasion strategies produces higher attack success rates against large reasoning models than prior jailbreak methods.

  10. Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling

    cs.CV 2026-05 unverdicted novelty 6.0

    Decouples semantic and spatial tokens in NVS transformers to resolve representation ambiguity, yielding consistent gains with near-zero added latency.

  11. SLASH the Sink: Sharpening Structural Attention Inside LLMs

    cs.AI 2026-05 unverdicted novelty 6.0

    LLMs spontaneously build graph topology inside their attention layers but attention sinks suppress it; SLASH redistributes attention to restore structural understanding without training.

  12. SLASH the Sink: Sharpening Structural Attention Inside LLMs

    cs.AI 2026-05 unverdicted novelty 6.0

    SLASH is a plug-and-play attention redistribution technique that counters attention sinks to enhance LLMs' intrinsic graph topology reconstruction without any training or fine-tuning.

  13. The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity

    cs.LG 2026-05 unverdicted novelty 6.0

    Attention sinks arise from variance discrepancy in self-attention value aggregation, amplified by super neurons and first-token dimension disparity, and can be mitigated by head-wise RMSNorm to accelerate pre-training...

  14. What Makes Position Zero Special? A Mechanistic Study of Position Zero Attention Sinks in LLMs

    cs.LG 2026-02 conditional novelty 6.0

    Position-zero attention sinks in transformers emerge from causal-masking asymmetry: position zero attends only to itself, and an MLP then amplifies its representation into a stable, high-norm 'sink'.

  15. Small Initialization Matters for Large Language Models

    cs.AI 2026-06 unverdicted novelty 5.0

    Reducing parameter initialization scale in LLMs improves pretraining and reasoning by inducing a low-to-high complexity developmental trajectory in weights.

  16. P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8

    cs.AR 2026-06 unverdicted novelty 5.0

    Forward KV iteration in FP8 attention produces P-collapse under attention sink; reverse iteration with S=256 removes it and is optimal among bit-exact scales.

  17. Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization

    cs.LG 2026-06 unverdicted novelty 5.0

    Massive spikes in LLMs are identified as rigid vector biases preserved in rotational stability zones; INSERTQUANT clamps them with template vectors to achieve spike-free PTQ with SOTA parity on LLMs and generalization...

  18. OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning

    cs.CV 2026-05 unverdicted novelty 5.0

    OccamToken replaces absolute token ranking with register-anchored relative evidence testing to enable adaptive, high-ratio visual token pruning in VLMs while preserving most accuracy.

  19. ASAP: Attention Sink Anchored Pruning

    cs.LG 2026-05 unverdicted novelty 5.0

    ASAP prunes tokens in ViTs by anchoring on attention sinks modeled as lazy random walks, using cumulative transition matrices and radial diffusion clustering to compress redundancy while preserving accuracy.

  20. SLASH the Sink: Sharpening Structural Attention Inside LLMs

    cs.AI 2026-05 unverdicted novelty 5.0

    SLASH redistributes attention in LLMs to amplify their spontaneous internal reconstruction of graph topologies, yielding gains on graph and molecular tasks.

  21. Exploring Motion-Language Alignment for Text-driven Motion Generation

    cs.CV 2026-04 unverdicted novelty 5.0

    MLA-Gen advances text-driven motion synthesis by aligning global motion patterns with fine-grained text semantics and mitigating attention sink effects via new masking techniques.

  22. Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse

    cs.CL 2026-02 unverdicted novelty 5.0

    Attention sinks forge native MoE mechanisms in attention layers that cause head collapse, addressed by sink-aware training with auxiliary load balancing.

  23. Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse

    cs.CL 2026-02 conditional novelty 5.0

    Attention-sink weight is recast as an implicit MoE router per head, motivating a sink-aware head-balancing loss that yields small, consistent benchmark gains across three attention variants but rests on a definitional...

  24. Information-Regularized Attention for Visual-Centric Reasoning

    cs.CV 2026-07 unverdicted novelty 4.0

    IRA is a stochastic attention mechanism that regulates visual information injection in VLMs to yield smoother embedding trajectories and reduced attention sinks.