Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T07:02:53.740597Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 100 of 159 outbound references and 100 inbound Pith citation observations for arXiv:2402.17762.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T07:02:53.740597Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T06:06:57.523801Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 159 outbound references displayed
External citation measurements
8
pith, observed 2026-08-05T02:28:24.338817Z
Observation bf7665b4-72b0-475b-945a-bff163c79dc9 · outbound
Massive Activations in Large Language Models Exploring Length Generalization in Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 37f480b5-0480-42a1-b37a-6e38239fffa5 · outbound
Massive Activations in Large Language Models Computational complexity: a modern approach
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 03c6b057-00e2-4fc8-9c27-4221672679bc · outbound
Massive Activations in Large Language Models End-to-end Algorithm Synthesis with Recurrent Networks: Logical Extrapolation Without Overthinking
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2aaa1d0a-f51b-4518-b377-c470626ae7eb · outbound
Massive Activations in Large Language Models Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f19f603-6130-40d2-99df-198aea02f32b · outbound
Massive Activations in Large Language Models Mix Barrington
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a93deff3-640e-4f57-b521-d1f692cdba26 · outbound
Massive Activations in Large Language Models Mix Barrington and Denis Thérien
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 283eca06-15f0-40e0-8499-fda607fdc194 · outbound
Massive Activations in Large Language Models On the ability and limitations of transformers to recognize formal languages
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0c657333-c417-484d-ac96-250f846e1ef2 · outbound
Massive Activations in Large Language Models Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c55341c7-b0cc-47f5-8a5d-7ef88b3854bb · outbound
Massive Activations in Large Language Models Unbounded fan-in circuits and associative functions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bfcd1834-5c72-4e04-a434-ce9dd23258e6 · outbound
Massive Activations in Large Language Models Decision transformer: Reinforcement learning via sequence modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa7b790d-53b2-4273-8abe-1a9e95f6fe11 · outbound
Massive Activations in Large Language Models Evaluating Large Language Models Trained on Code
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c8f06064-a61b-424e-a186-62fbc5164034 · outbound
Massive Activations in Large Language Models Finite-automaton aperiodicity is PSPACE -complete
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9e437545-8606-4bf4-9adc-1faa13903d5d · outbound
Massive Activations in Large Language Models The algebraic theory of context-free languages
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 20bf5ead-4004-41af-b2ca-25e534e3039a · outbound
Massive Activations in Large Language Models Conditional Positional Encodings for Vision Transformers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f851dedd-900e-4e83-a2a1-709204939889 · outbound
Massive Activations in Large Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 24cac180-9e4e-4bee-ac4b-61df7b3946d3 · outbound
Massive Activations in Large Language Models Approximation by superpositions of a sigmoidal function
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 12758b5f-19f0-4843-bbc3-47b592a800ec · outbound
Massive Activations in Large Language Models Depth separation for neural networks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5760ebfe-4b0f-4557-8a55-fee9a6865c21 · outbound
Massive Activations in Large Language Models Learning parities with neural networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3dd2a47f-bb65-4b1e-80b2-0a91578c0491 · outbound
Massive Activations in Large Language Models Universal transformers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9587bdfb-4457-41ac-a217-9877c405cf5a · outbound
Massive Activations in Large Language Models Neural Networks and the Chomsky Hierarchy
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1438c5d8-e3cb-47f3-9666-1f75af0df165 · outbound
Massive Activations in Large Language Models Patti, Jayson Lynch, Avi Shporer, Nakul Verma, Eugene Wu, and Gilbert Strang
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a92effb4-23f9-4782-9fba-952d5d7acee8 · outbound
Massive Activations in Large Language Models How can self-attention networks recognize D yck-n languages? In Findings of the Association for Computational Linguistics: EMNLP
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e9d71054-971f-43cd-8c8e-c8e76299aaca · outbound
Massive Activations in Large Language Models Inductive biases and variable creation in self-attention mechanisms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 585d67d4-f497-4b16-9b8e-0013a444403c · outbound
Massive Activations in Large Language Models Computational Holonomy Decomposition of Transformation Semigroups
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b054a14d-3746-4d23-a476-43ba27cb5b7e · outbound
Massive Activations in Large Language Models Automata, languages, and machines
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c77fa33-cbbf-43de-9e27-e19eb874af7a · outbound
Massive Activations in Large Language Models The power of depth for feedforward neural networks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f9d7f47e-4285-43e0-b52e-33b1c7728c0e · outbound
Massive Activations in Large Language Models A mathematical framework for transformer circuits
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a053446c-b77f-4ea0-ad7e-7a608e4877e8 · outbound
Massive Activations in Large Language Models Saxe, and Michael Sipser
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 233d2411-6057-4706-b4b1-bc989a61be87 · outbound
Massive Activations in Large Language Models Shortcut learning in deep neural networks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e81aff0e-10ba-4a11-a325-4f2807ad109b · outbound
Massive Activations in Large Language Models Looped Transformers as Programmable Computers
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c302ae6d-b427-43a3-af3c-1aadcee5c87f · outbound
Massive Activations in Large Language Models Reliably learning the R e LU in polynomial time
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8c69106f-3e51-4e1c-ae02-abd60c37a366 · outbound
Massive Activations in Large Language Models Adaptive Computation Time for Recurrent Neural Networks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cf40a2ef-ccda-41ad-a4f0-50d4fc44a1c9 · outbound
Massive Activations in Large Language Models Neural Turing Machines
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 68d22231-dbe0-49a6-89f7-56ece1c3ee35 · outbound
Massive Activations in Large Language Models Non-Autoregressive Neural Machine Translation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 127c3f46-08d3-411f-8224-66699687de2f · outbound
Massive Activations in Large Language Models Dream to Control: Learning Behaviors by Latent Imagination
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7a56b030-39cc-4fea-ab80-80276c58a064 · outbound
Massive Activations in Large Language Models Theoretical limitations of self-attention in neural sequence models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 653efb2c-3ef8-4d44-b276-5311aa9a4380 · outbound
Massive Activations in Large Language Models Transformer Language Models without Positional Encodings Still Learn Positional Information
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 74df9dd5-5ddd-4e7d-8d12-57ef312aad7a · outbound
Massive Activations in Large Language Models Deep residual learning for image recognition
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 04632b0f-3c44-4673-9256-cd26b6f18a59 · outbound
Massive Activations in Large Language Models Towards lower bounds on the depth of R e LU neural networks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e1e15af2-b2fb-41a1-8fb0-2ddd27a60ec8 · outbound
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 47e79d1b-50f8-485b-a623-06a975162c1d · outbound
Massive Activations in Large Language Models Multilayer feedforward networks are universal approximators
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 96fb3a48-c9fa-48c0-8c27-01efe73efd82 · outbound
Massive Activations in Large Language Models Universal Language Model Fine-tuning for Text Classification
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 68e668c7-e742-40bc-b1f2-6369c30a19c3 · outbound
Massive Activations in Large Language Models Block-Recurrent Transformers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 76e8e6fe-dfd2-497f-aa61-85680fe996a3 · outbound
Massive Activations in Large Language Models Offline reinforcement learning as one big sequence modeling problem
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation db9ec858-602d-451c-9fa5-9eeb85bb7ab9 · outbound
Massive Activations in Large Language Models Finetuning Pretrained Transformers into RNNs
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3a7f9b6d-e7bb-42ed-84d1-7cf952c15b8a · outbound
Massive Activations in Large Language Models Rethinking Positional Encoding in Language Pre-training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c7d13e69-2620-49e1-a2fe-d5a3176b760e · outbound
Massive Activations in Large Language Models The number of semigroups of order n
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dbde73a0-d7f7-456d-b251-9bd3520677a7 · outbound
Massive Activations in Large Language Models Finite permutation groups with large abelian quotients
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6b88597f-d4af-4ca5-945d-44bbbca1878d · outbound
Massive Activations in Large Language Models Produit complet des groupes de permutations et probleme d’extension de groupes II
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 59756881-3cb6-4eeb-bdf4-42e54e602ae0 · outbound
Massive Activations in Large Language Models Algebraic theory of machines, I : P rime decomposition theorem for finite semigroups and machines
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e5a57c1f-6aa7-4811-9b0c-a7a2fa2bccab · outbound
Massive Activations in Large Language Models Deep Learning for Symbolic Mathematics
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b8f8724a-2a4f-438c-84cc-792030f1fa64 · outbound
Massive Activations in Large Language Models FractalNet: Ultra-Deep Neural Networks without Residuals
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d17ad449-261b-4f31-82de-492cb1bfac53 · outbound
Massive Activations in Large Language Models On the ability of neural nets to express distributions
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3576dd43-7ecb-4202-8aff-56b9e07acc0d · outbound
Massive Activations in Large Language Models Competition-Level Code Generation with AlphaCode
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation faf9b645-badb-4687-901e-ea787e9f8423 · outbound
Massive Activations in Large Language Models Decoupled Weight Decay Regularization
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5485814b-ba2f-4b17-b144-ae447a2d7e27 · outbound
Massive Activations in Large Language Models On the K rohn- R hodes cascaded decomposition theorem
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b007f765-2385-4b31-acc8-14c145a6d39d · outbound
Massive Activations in Large Language Models On the cascaded decomposition of automata, its complexity and its application to logic ( D raft)
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 31dce5d4-1a0d-4c37-b2fe-367d9443d2f6 · outbound
Massive Activations in Large Language Models Threshold circuits for iterated matrix product and powering
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 46f67443-bc99-49d0-9490-7fc1ac0ea102 · outbound
Massive Activations in Large Language Models Saturated Transformers are Constant-Depth Threshold Circuits
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c4a750c0-6672-4e3c-8023-4b52e399a5d5 · outbound
Massive Activations in Large Language Models Transformers are Sample-Efficient World Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 15333cc0-9d04-498e-a6de-c19c7116fe3c · outbound
Massive Activations in Large Language Models Lower bounds over Boolean inputs for deep neural networks with ReLU gates
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 29ea9dea-ceda-4589-bec7-1f83577a80d0 · outbound
Massive Activations in Large Language Models A mechanistic interpretability analysis of grokking
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 794952e7-ffaa-4cff-a572-5994a2a6b0a1 · outbound
Massive Activations in Large Language Models Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6c66c989-11d2-4082-8819-e90c99765ee8 · outbound
Massive Activations in Large Language Models Identifying good directions to escape the NTK regime and efficiently learn low-degree plus sparse polynomials
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation efcc52a1-63a1-4fb6-8cf7-19a34046c046 · outbound
Massive Activations in Large Language Models Investigating the Limitations of Transformers with Simple Arithmetic Tasks
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 23fb748d-b973-4201-b7e2-abcd1725ceee · outbound
Massive Activations in Large Language Models Show Your Work: Scratchpads for Intermediate Computation with Language Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 919bd7a7-9e06-4204-bfb6-2675656d8e5d · outbound
Massive Activations in Large Language Models The complexity of M arkov decision processes
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ce8c7133-1e27-47e1-9a13-3036070ed7b5 · outbound
Massive Activations in Large Language Models Py T orch: An imperative style, high-performance deep learning library
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 22a9f215-6e25-4c69-8d0c-ad6e8d24b5b9 · outbound
Massive Activations in Large Language Models Attention is turing complete
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b66351a4-83ca-44e6-9307-ccbce2b6a2c7 · outbound
Massive Activations in Large Language Models Deep contextualized word representations
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9bec790d-b069-45cc-93ed-ef94672a6338 · outbound
Massive Activations in Large Language Models Generative Language Modeling for Automated Theorem Proving
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ed3a336f-a298-430f-9e4a-681e03d47ff8 · outbound
Massive Activations in Large Language Models Train short, test long: Attention with linear biases enables input length extrapolation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation daed02a6-0f65-43b5-a264-a14fb8bf4b36 · outbound
Massive Activations in Large Language Models Language models are unsupervised multitask learners
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 155f3d6b-cc80-4e4f-99d5-9b797f5a52fc · outbound
Massive Activations in Large Language Models Reif and Stephen R
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fd95f318-987b-4d3b-bb37-7cfd3b71e1a8 · outbound
Massive Activations in Large Language Models Applications of automata theory and algebra: via the mathematical theory of complexity to biology, physics, psychology, philosophy, and games
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bfd3ffa3-c957-4cf2-b584-4ccace65d90b · outbound
Massive Activations in Large Language Models Can contrastive learning avoid shortcut solutions? Advances in Neural Information Processing Systems
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6dddc637-081d-4bd6-ae71-da93130e882a · outbound
Massive Activations in Large Language Models Depth separations in neural networks: what is actually being separated? In Conference on Learning Theory, pages 2664--2666
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8e9b267b-53e0-4d96-ad42-5d0d1a8f979b · outbound
Massive Activations in Large Language Models Programming puzzles
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 96d86380-8c48-4056-a082-7be0ac86f010 · outbound
Massive Activations in Large Language Models On finite monoids having only trivial subgroups
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e6c89032-ba93-44b7-bf40-f51c2da5b9ce · outbound
Massive Activations in Large Language Models Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5ad8215-382b-41ae-b4a3-f8c1a7ef8088 · outbound
Massive Activations in Large Language Models On the computational power of neural nets
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fc8285ad-55a1-47c2-8e7e-0b914092f37c · outbound
Massive Activations in Large Language Models Benefits of depth in neural networks
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cf386873-078d-4130-bbdf-4b4f0b19ec21 · outbound
Massive Activations in Large Language Models BERT Rediscovers the Classical NLP Pipeline
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 90c7731a-54c6-40fe-b20d-fcd8dadc4053 · outbound
Massive Activations in Large Language Models WaveNet: A Generative Model for Raw Audio
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e4eb709b-3634-4a27-a53b-41b0fba961c9 · outbound
Massive Activations in Large Language Models Attention is all you need
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6592dcfc-0acb-49cc-aaaf-d1c573cfb465 · outbound
Massive Activations in Large Language Models Hechtman, and Jonathon Shlens
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bd94ffee-00a1-43de-ab48-9479bc180edb · outbound
Massive Activations in Large Language Models Visualizing Attention in Transformer-Based Language Representation Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 06c1682f-2943-4fdf-9c88-a1c241b5c5a9 · outbound
Massive Activations in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ec2f7afa-eb70-4ad6-8534-45b117f270dc · outbound
Massive Activations in Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5722265a-b603-46c0-a7cc-2bd47d81cc22 · outbound
Massive Activations in Large Language Models Thinking like T ransformers
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e1b26e7d-8922-46c7-80b0-b7f77f5fe92f · outbound
Massive Activations in Large Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4cd7e10a-3919-4403-bdf1-9bc8c6bee864 · outbound
Massive Activations in Large Language Models Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7e01e40b-bf8d-40be-9e15-c29aab96af9b · outbound
Massive Activations in Large Language Models A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c1b399f-6259-4e54-ab8a-c104186c7f9a · outbound
Massive Activations in Large Language Models How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e01ed2c6-4416-4824-9291-6aa1e0219a6f · outbound
Massive Activations in Large Language Models Papadimitriou, and Karthik Narasimhan
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 808e24e5-1ade-43bc-8f4e-b2beb6c0c6af · outbound
Massive Activations in Large Language Models Mastering atari games with limited data
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 49649fc0-c2c0-4f38-8b26-6ed9cea981d7 · outbound
Massive Activations in Large Language Models Cascade synthesis of finite-state machines
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e2fd5998-2df1-4a8b-8738-4969b2513d5b · outbound
Massive Activations in Large Language Models Pointer Value Retrieval: A new benchmark for understanding the limits of neural network generalization
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 693bd688-01ea-445e-b965-5fee4316af7c · outbound
Massive Activations in Large Language Models How does mixup help with robustness and generalization? In International Conference on Learning Representations, 2021 b
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 609d4f79-b35b-426f-80a3-9e7d0721a608 · outbound
Massive Activations in Large Language Models Unveiling Transformers with LEGO: a synthetic reasoning task
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0f564b89-e117-4001-8b73-a4d58afb5382 · inbound
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling Massive Activations in Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 09dd05df-56e3-4318-9adf-f4a805d11c3b · inbound
Scaling and evaluating sparse autoencoders Massive Activations in Large Language Models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b383fa3f-ff15-4c4f-8dea-dbca88241704 · inbound
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision Massive Activations in Large Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6ba80cee-860c-4379-af58-912b3c3f9c9f · inbound
When Attention Sink Emerges in Language Models: An Empirical View Massive Activations in Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 12916e7a-8ac9-475f-9935-e4e8a7a1d8fc · inbound
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training Massive Activations in Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd48f046-aa25-41ac-bf2b-ac475226fc33 · inbound
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Massive Activations in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fdcb51-0b0b-4fd4-a04a-5cb8a1238b9b · inbound
DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation Massive Activations in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53fad53-d3fa-47da-a067-5aa36e15bb28 · inbound
TinyFusion: Diffusion Transformers Learned Shallow Massive Activations in Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b425e016-3674-438a-b52d-afd409cf979b · inbound
MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic Segmentation Massive Activations in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e72f8a0-3c79-43c0-a4c8-d86e78caafb8 · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management Massive Activations in Large Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54ef8b7-9dd2-4ba5-a3cb-254468d55c91 · inbound
Leveraging Registers in Vision Transformers for Robust Adaptation Massive Activations in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 036eb6bc-19f3-4c1e-879b-2877d18b95b1 · inbound
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Massive Activations in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91fcc470-d158-4f3a-9e65-32c8a9810006 · inbound
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations Massive Activations in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceeaf3e5-1f44-4d4d-8e3c-c1ce36297e03 · inbound
An Inquiry into Datacenter TCO for LLM Inference with FP8 Massive Activations in Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b3edb7-d041-4212-bc15-c788e160fcef · inbound
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Massive Activations in Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ac3d208-f294-4d75-aeee-520b786000be · inbound
Systematic Outliers in Large Language Models Massive Activations in Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5341b35f-be30-4e0b-9ff7-fad88cf968c5 · inbound
TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks Massive Activations in Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2872e6-71dd-45b3-9642-d1d671307c1f · inbound
Quantifying Memory Utilization with Effective State-Size Massive Activations in Large Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45cac19a-40ed-437a-98ad-ca2e1c43b986 · inbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Massive Activations in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a506b6ff-1d76-45c4-9e54-27ef1cced7a3 · inbound
ICQuant: Index Coding enables Low-bit LLM Quantization Massive Activations in Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc90ad85-5426-4828-a92f-99e0c69fa85e · inbound
MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Massive Activations in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d42df7f-ea28-419d-9a80-f2464a4d062b · inbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Massive Activations in Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ee3f5c68-3595-42c2-95cf-a2be8e62329c · inbound
Polysemy of Synthetic Neurons Towards a New Type of Explanatory Categorical Vector Spaces Massive Activations in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2c8c5c-774a-4a46-9200-b00f149dacea · inbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Massive Activations in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb0f94bf-7c2a-4275-9518-11fbd28bcacf · inbound
Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Massive Activations in Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8bbd1dc-631f-4ba0-bba4-95b283e3cd07 · inbound
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Massive Activations in Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0146fe6b-aa47-49dd-86b2-f7f12d72334e · inbound
Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI Massive Activations in Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c67f3b-c46e-4aa9-80f0-31229390a4ca · inbound
Rethinking the Outlier Distribution in Large Language Models: An In-depth Study Massive Activations in Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc63c28-99f4-47d7-b4b6-d19a1ba1c380 · inbound
FPTQuant: Function-Preserving Transforms for LLM Quantization Massive Activations in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8cfbe6d-2231-4c8b-8721-906e352e75b4 · inbound
PEVLM: Parallel Encoding for Vision-Language Models Massive Activations in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e1d3a8a-d15c-409d-bba9-fc01f777ec28 · inbound
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models Massive Activations in Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9597c0d-7cd2-468a-9765-fd31b87c8fbe · inbound
Characterization and Mitigation of Training Instabilities in Microscaling Formats Massive Activations in Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328b74aa-c4eb-4829-b374-fc4130bcb1e3 · inbound
DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Massive Activations in Large Language Models
Reference 130
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac939149-f285-4c96-b1e9-8d3269bf9ef8 · inbound
Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers Massive Activations in Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 127db34a-44aa-44c1-9ba1-042b06b07865 · inbound
Mitigating Spurious Correlations in Weakly Supervised Semantic Segmentation via Cross-architecture Consistency Regularization Massive Activations in Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a28053-910a-4330-9e3e-a3bb82c42a71 · inbound
Discriminating Distal Ischemic Stroke from Seizure-Induced Stroke Mimics Using Dynamic Susceptibility Contrast MRI Massive Activations in Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef4d47a-fef7-4492-9b73-f80f95bab925 · inbound
CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models Massive Activations in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c8d797-2c9b-4f86-b5a7-7b1ba4d2cee5 · inbound
Prophecy: Inferring Formal Properties from Neuron Activations Massive Activations in Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cba88781-4174-4f78-8112-203773de3080 · inbound
CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs Massive Activations in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 11853ee6-3d40-4068-ac60-7334cef886f2 · inbound
MiMo-V2-Flash Technical Report Massive Activations in Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7ff386ff-d10b-413c-842a-21c3972864b8 · inbound
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers Massive Activations in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3020cc14-71ed-4415-8184-1da2152ecdd9 · inbound
Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Massive Activations in Large Language Models
Reference 291
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ffc5e7d6-2ead-4d3b-9349-b6a1fc54ba35 · inbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Massive Activations in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e7f4d679-9ca7-453f-ba8c-a03d7c8ba61f · inbound
Efficient Reasoning on the Edge Massive Activations in Large Language Models
Reference 122
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b6b7d8-2d88-4b30-a58a-3d7207d95271 · inbound
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models Massive Activations in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6dee6136-6b0a-4d77-b30f-d58da2d1b15a · inbound
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models Massive Activations in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80a1876-0fdb-4220-b322-79369b779401 · inbound
Noise Steering for Controlled Text Generation: Improving Diversity and Reading-Level Fidelity in Arabic Educational Story Generation Massive Activations in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f49cec62-35b2-4572-bea6-adc1f4d780ae · inbound
OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension Massive Activations in Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b1dc4a5b-5ac9-4a65-81c5-c0d9c08aab5e · inbound
Graph-Guided Adaptive Channel Elimination for KV Cache Compression Massive Activations in Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 83b3ef74-a5a2-466f-b0c2-4c4937fa1a99 · inbound
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization Massive Activations in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 95e5c181-0bb2-476c-87e6-8d1b341ab6f0 · inbound
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs Massive Activations in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2dd3e400-3304-40f7-b222-803c265efc7a · inbound
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing Massive Activations in Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 55732636-5021-41b6-9195-eab9969b97c2 · inbound
Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay Massive Activations in Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9b70fc32-94e4-4282-b233-b2fae0127f2d · inbound
Taming Outlier Tokens in Diffusion Transformers Massive Activations in Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d6009a94-3323-4e32-b7bd-c6b32bf871cc · inbound
HyperLens: Quantifying Cognitive Effort in LLMs with Fine-grained Confidence Trajectory Massive Activations in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3cfaea57-1004-4290-98c1-8905cd8367e0 · inbound
A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Massive Activations in Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ae2b1b10-7ccd-43af-b80f-0655ff6f4e04 · inbound
A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Massive Activations in Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 922095c1-efa6-420a-a1c3-694242266ca4 · inbound
Attention Sinks in Diffusion Transformers: A Causal Analysis Massive Activations in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cb3946b7-ea77-4c40-9bbb-e6d4d25819e1 · inbound
Attention Sinks in Diffusion Transformers: A Causal Analysis Massive Activations in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ceba6819-4be6-47be-b841-f47b7fbf3bcf · inbound
Attention Sinks in Diffusion Transformers: A Causal Analysis Massive Activations in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ea74b8f5-9e0e-4191-8a50-34700be059a3 · inbound
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination Massive Activations in Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c9d3649-708d-45a2-b29f-a57bcd5c25e4 · inbound
Registers Matter for Pixel-Space Diffusion Transformers Massive Activations in Large Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 80abfa05-8b46-4a34-94b2-b50624b91b6d · inbound
Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise Massive Activations in Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e08a1f29-2709-4f9f-bf97-79d48ac48c83 · inbound
A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions Massive Activations in Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f08fd51c-3415-4236-a430-80a793d5822d · inbound
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register Massive Activations in Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 394692bc-6410-4042-a7aa-acc161029fe6 · inbound
OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Massive Activations in Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 44d160ba-f95a-4d6e-bfb4-cbcc6e847c31 · inbound
Steered Generation via Gradient-Based Optimization on Sparse Query Features Massive Activations in Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 27532da8-eaba-4e7b-b0c6-7be961a8d607 · inbound
A Simple Plug-in for Improving Eviction-Based KV Cache Compression Massive Activations in Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5a9ecc30-18ea-4fd4-93e1-ef7e3d804773 · inbound
Multi-Gate Residuals Massive Activations in Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 09f478f2-335a-4ce6-80da-eb1f509e88f4 · inbound
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models Massive Activations in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 18d3ee7d-4d24-41b5-af02-b14fbeb4c824 · inbound
Rethinking the Role of Tensor Decompositions in Post-Training LLM Compression Massive Activations in Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 94ed6142-bd48-4a66-8ca7-1bb3d67694f2 · inbound
When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models Massive Activations in Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6a09064c-2530-4d0f-ae0e-1305d741deac · inbound
Dominant-Layer ZO: A Single Layer Dominates Zeroth-Order Fine-Tuning of LLMs Massive Activations in Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ab8c87e8-10b9-4d2b-a08b-368bcd703a83 · inbound
Dead Directions: Geometric Singular Learning Massive Activations in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c3a4e679-9f32-4593-819f-9832b2a8718b · inbound
P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8 Massive Activations in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 02f250f4-fe71-41c5-8299-34fea20461bb · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Massive Activations in Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 51a33f4e-b687-467c-8c08-3ba3eebdd8ca · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Massive Activations in Large Language Models
Reference 172
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5d824933-ff7a-4041-842b-d0eaf2b5bec4 · inbound
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Massive Activations in Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 172cd7cd-07e9-484b-8f53-863739c41580 · inbound
ICA Lens: Interpreting Language Models Without Training Another Dictionary Massive Activations in Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2e7efc17-490e-4eca-a6db-881506a87601 · inbound
Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models Massive Activations in Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 762dfa80-5f30-456e-a465-a0ac29fb917e · inbound
DynamicPTQ: Mitigating Activation Quantization Collapse via Residual-Stream Dynamics Massive Activations in Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a4d8e077-dc61-44be-a07e-303fd8770aca · inbound
MiniMax Sparse Attention Massive Activations in Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5782f62-4696-4b0c-b7ce-bd0696b564f7 · inbound
Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications Massive Activations in Large Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d2115d05-ec7e-42a5-b466-24f950eb2ad8 · inbound
Algebraic Dead Directions in LayerNorm Transformers: A Forward-Pass-Only Diagnostic at LLM Scale Massive Activations in Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 16a81b1e-96f9-4708-948f-ae046188af58 · inbound
Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test Massive Activations in Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cb690365-a9ba-4e4a-bfc6-16a759ba42a4 · inbound
Demystifying Numerical Instability in LLM Inference: Achieving Reproducible Inference for Mission-Critical Tasks with HEAL Massive Activations in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 46bd5bef-62cc-4abc-848b-9ef34ed5d7e1 · inbound
SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference Massive Activations in Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b028e290-b7d9-4a7b-bcc9-987a888f41fd · inbound
Information-Regularized Attention for Visual-Centric Reasoning Massive Activations in Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4d21c1c0-3162-443d-9e09-7db8d820e31f · inbound
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation Massive Activations in Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ab7a8d-73d7-4e49-a1f5-600ae34915e1 · inbound
Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale Massive Activations in Large Language Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 70a0c0dc-55c0-4f0d-b0c4-7fe95e173d1c · inbound
Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator Massive Activations in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c206bbc-5ba5-4cdc-b4de-c344a95c3fc6 · inbound
Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization Massive Activations in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbaf008b-e25d-4a52-ae7d-8becdb250a27 · inbound
Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering Massive Activations in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6794ba7-bd94-4812-94aa-56d7ae62c82a · inbound
Phasor Attention: Mean Root Square Normalization for Phase Manifold Preservation Massive Activations in Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a443176-c429-41cb-85b5-64995baa358c · inbound
Estimating Rare Events in Language Models with Proper Evaluation Massive Activations in Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39f12d3-f021-4699-939d-3efa67c21e5e · inbound
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Massive Activations in Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6459f884-0819-411f-a870-8fb057a89af0 · inbound
Reference Feature Atlases for Mechanistic Auditing of Language Models Massive Activations in Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4374f278-588b-48ab-808a-386da25b46e2 · inbound
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization Massive Activations in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ff366e-1770-45e8-ac9e-fe0d045b1de9 · inbound
Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs Massive Activations in Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8096263e-d7c3-4f93-8069-42b2d96aabc1 · inbound
Hidden Language Consistency Phenomena in Reasoning LLMs Massive Activations in Large Language Models
Reference 210
Source-reported events for the cited work
Unavailable: canonical work link unavailable.