{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:CJBGWDZRGL24P6W2S6USAY6M6K","short_pith_number":"pith:CJBGWDZR","schema_version":"1.0","canonical_sha256":"12426b0f3132f5c7fada97a92063ccf2b0aa5b78e4fa40651db52d65df754e19","source":{"kind":"arxiv","id":"2504.02732","version":4},"attestation_state":"computed","paper":{"title":"Why do LLMs attend to the first token?","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"\\'Alvaro Arroyo, Christos Perivolaropoulos, Federico Barbero, Michael Bronstein, Petar Veli\\v{c}kovi\\'c, Razvan Pascanu, Xiangming Gu","submitted_at":"2025-04-03T16:17:55Z","abstract_excerpt":"Large Language Models (LLMs) tend to attend heavily to the first token in the sequence -- creating a so-called attention sink. Many works have studied this phenomenon in detail, proposing various ways to either leverage or alleviate it. Attention sinks have been connected to quantisation difficulties, security issues, and streaming attention. Yet, while many works have provided conditions in which they occur or not, a critical question remains shallowly answered: Why do LLMs learn such patterns and how are they being used? In this work, we argue theoretically and empirically that this mechanis"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2504.02732","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2025-04-03T16:17:55Z","cross_cats_sorted":[],"title_canon_sha256":"8b41f0d8a8bbd12dd2910e94cf9057f8f01b6f7c3911cd602e61efcba2aacc32","abstract_canon_sha256":"152672d7f2da2723a358e5a25a1461305e047c311974f0e50e9df637cf909b2a"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:48:48.431591Z","signature_b64":"7zohNvM19WGK0hVG75iTzAfmfZzylVmLs4GeGUwxpY8JXKKGXWtQG8ezpyiu/9gPkVMBGdI5qORttXwtyr3YDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"12426b0f3132f5c7fada97a92063ccf2b0aa5b78e4fa40651db52d65df754e19","last_reissued_at":"2026-07-05T11:48:48.431110Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:48:48.431110Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Why do LLMs attend to the first token?","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"\\'Alvaro Arroyo, Christos Perivolaropoulos, Federico Barbero, Michael Bronstein, Petar Veli\\v{c}kovi\\'c, Razvan Pascanu, Xiangming Gu","submitted_at":"2025-04-03T16:17:55Z","abstract_excerpt":"Large Language Models (LLMs) tend to attend heavily to the first token in the sequence -- creating a so-called attention sink. Many works have studied this phenomenon in detail, proposing various ways to either leverage or alleviate it. Attention sinks have been connected to quantisation difficulties, security issues, and streaming attention. Yet, while many works have provided conditions in which they occur or not, a critical question remains shallowly answered: Why do LLMs learn such patterns and how are they being used? In this work, we argue theoretically and empirically that this mechanis"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2504.02732","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2504.02732/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2504.02732","created_at":"2026-07-05T11:48:48.431171+00:00"},{"alias_kind":"arxiv_version","alias_value":"2504.02732v4","created_at":"2026-07-05T11:48:48.431171+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2504.02732","created_at":"2026-07-05T11:48:48.431171+00:00"},{"alias_kind":"pith_short_12","alias_value":"CJBGWDZRGL24","created_at":"2026-07-05T11:48:48.431171+00:00"},{"alias_kind":"pith_short_16","alias_value":"CJBGWDZRGL24P6W2","created_at":"2026-07-05T11:48:48.431171+00:00"},{"alias_kind":"pith_short_8","alias_value":"CJBGWDZR","created_at":"2026-07-05T11:48:48.431171+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":20,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.20743","citing_title":"Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2606.17945","citing_title":"Small Initialization Matters for Large Language Models","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00434","citing_title":"Information-Regularized Attention for Visual-Centric Reasoning","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2606.06521","citing_title":"P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02288","citing_title":"Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.29657","citing_title":"OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07604","citing_title":"Contribution Weights: A Geometrical Analysis of Self-Attention Transformers","ref_index":143,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07604","citing_title":"Contribution Weights: A Geometrical Analysis of Self-Attention Transformers","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22372","citing_title":"ASAP: Attention Sink Anchored Pruning","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10503","citing_title":"SLASH the Sink: Sharpening Structural Attention Inside LLMs","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18599","citing_title":"Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19485","citing_title":"Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2601.21366","citing_title":"Perceptrons and localization of attention's mean-field landscape","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2602.01203","citing_title":"Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12879","citing_title":"ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2604.02973","citing_title":"Exploring Motion-Language Alignment for Text-driven Motion Generation","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10503","citing_title":"SLASH the Sink: Sharpening Structural Attention Inside LLMs","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10503","citing_title":"SLASH the Sink: Sharpening Structural Attention Inside LLMs","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06611","citing_title":"The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2604.11791","citing_title":"A Mechanistic Analysis of Looped Reasoning Language Models","ref_index":5,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K","json":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K.json","graph_json":"https://pith.science/api/pith-number/CJBGWDZRGL24P6W2S6USAY6M6K/graph.json","events_json":"https://pith.science/api/pith-number/CJBGWDZRGL24P6W2S6USAY6M6K/events.json","paper":"https://pith.science/paper/CJBGWDZR"},"agent_actions":{"view_html":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K","download_json":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K.json","view_paper":"https://pith.science/paper/CJBGWDZR","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2504.02732&json=true","fetch_graph":"https://pith.science/api/pith-number/CJBGWDZRGL24P6W2S6USAY6M6K/graph.json","fetch_events":"https://pith.science/api/pith-number/CJBGWDZRGL24P6W2S6USAY6M6K/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K/action/timestamp_anchor","attest_storage":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K/action/storage_attestation","attest_author":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K/action/author_attestation","sign_citation":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K/action/citation_signature","submit_replication":"https://pith.science/pith/CJBGWDZRGL24P6W2S6USAY6M6K/action/replication_record"}},"created_at":"2026-07-05T11:48:48.431171+00:00","updated_at":"2026-07-05T11:48:48.431171+00:00"}