Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:18:35.310454Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.12149.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:18:35.310454Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a3e4f8ee-7b0c-442d-a198-3843e0ce60c1 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Systematic Outliers in Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3112966e-095e-4839-aad6-95f7a8001afb · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Qwen3-Coder-Next Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ebc945-e599-412a-8cab-0aaef105baa3 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 81e1fd05-f38e-46e3-8916-68448e4f8859 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus No Language Left Behind: Scaling Human-Centered Machine Translation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d39d2e4-ab2e-4fa0-807e-4b3fc554fa30 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus The Zamba2 Suite: Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5a2055-c360-44c3-9db5-5998aee115e3 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Summer is warm. Winter is cold
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7ee92d5a-5642-4717-b0ee-d0a87341525a · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bb65d9d-9b7b-4688-a475-873bb44e2299 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Jamba: A Hybrid Transformer-Mamba Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a302ee69-6b44-45a7-adb1-7f83887a5629 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Openceres: When open information extrac- tion meets the semi-structured web
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fafa363d-1695-48a8-bb41-1f8f3d0a9ef1 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Pointer Sentinel Mixture Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db1afc5-9de9-47e2-a091-0ae47a10a8e1 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0995e174-6174-4423-8e46-36b4fc2fec4d · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Kvsink: Understanding and enhancing the preservation of attention sinks in kv cache quantization for llms
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5d2fe05f-6c5f-44c8-a6b7-39cba0231592 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Massive Activations in Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acc87cef-edfd-409a-922c-52d187060696 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus The spike, the sparse and the sink: Anatomy of massive activations and attention sinks.arXiv preprint arXiv:2603.05498,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14dddc32-3e14-45a6-8519-315543d69a32 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Retentive Network: A Successor to Transformer for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f92748c3-ef65-44b7-b624-05681a460ac4 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Kimi Linear: An Expressive, Efficient Attention Architecture
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c1c30f-70c4-4998-bb6c-5a09dc94086e · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Efficient streaming language models with attention sinks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2c8b2b09-17d6-415c-a2f3-ee992ed281fc · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Exploring layer-wise information effectiveness for post- training quantization in small language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 84c599d5-2be7-49f9-a228-95e0a16947ef · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Qwen3 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffbc3440-38ca-4ee5-981b-2cf3dcb09cf8 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Gated Linear Attention Transformers with Hardware-Efficient Training
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b90a2a02-babf-40a6-bf41-50c4cc34ca8c · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Gated delta networks: Improving mamba2 with delta rule
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 27f9d213-eb58-45fd-8c30-31fe6e1d74e4 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Beyond outliers: A data-free layer-wise mixed- precision quantization approach driven by numerical and structural dual-sensitivity.Under review, 2026a
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e1b6f645-94c0-43a0-9a7d-a478ae67af47 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Summer is warm. Winter is cold
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation be2d2464-8e88-4fd1-aa93-5d11a0f77b69 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Pronounced MAs concentrate at attention-sink positions, particu- larly the initial token, “Summer,” and the first period
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2cfc315c-9fcc-4277-802e-3c3332daa3c4 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Summer is warm. Winter is cold
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f27c287a-5beb-4819-a9db-1393a63cfd01 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus A Systematic Analysis of Hybrid Linear Attention
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18b9ebf3-aa6b-4d70-8e6d-9ca211bb544f · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Training Verifiers to Solve Math Word Problems
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 857cc832-55c3-4e70-9512-ac5beb451c75 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10db5d10-1117-4b87-81fd-f27751b72f8d · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18342080-f70e-42a4-8dfd-c51c76f33186 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus A discourse-aware attention model for abstractive summarization of long documents
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 95ea3654-2dba-4d5b-bcf2-7d3ea87be02f · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Hidden dynamics of massive activations in transformer training.arXiv preprint arXiv:2508.03616,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5391b3-9799-4432-a250-c305d3960c86 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a8fe65-24aa-4ba7-b896-2672290f103c · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde16a2e-2a9b-48ae-b816-d593fcaf3994 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Just read twice: closing the recall gap for recurrent language models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32abbfe9-7ab3-4e07-bba0-b54781aa3997 · outbound
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.