Pith. sign in

Paper Citation Record · LEDGER

Massive Activations in Large Language Models

As of 24 August 2026, this Paper Citation Record lists 100 of 159 outbound references and 100 inbound Pith citation observations for arXiv:2402.17762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.17762 v2

Coverage vector

measured 100 of 159 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T07:02:53.740597Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 100 of 102 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:06:57.523801Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 159 outbound references displayed

  • verified exact30
  • verified fuzzy57
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation bf7665b4-72b0-475b-945a-bff163c79dc9 · outbound

This paper cites Exploring Length Generalization in Large Language Models.

Massive Activations in Large Language Models Exploring Length Generalization in Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.818893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:5a9a44fc43bdc4e70861e7e314de4538daa24e7d34666e0473beaa96a14d064d

Observation 37f480b5-0480-42a1-b37a-6e38239fffa5 · outbound

This paper cites Computational complexity: a modern approach.

Massive Activations in Large Language Models Computational complexity: a modern approach

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.081423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:4a1bee3f27ee6b182620a9a6e682c5756970f34012dd4a74b73f8007ec4a1c6d

Observation 03c6b057-00e2-4fc8-9c27-4221672679bc · outbound

This paper cites End-to-end Algorithm Synthesis with Recurrent Networks: Logical Extrapolation Without Overthinking.

Massive Activations in Large Language Models End-to-end Algorithm Synthesis with Recurrent Networks: Logical Extrapolation Without Overthinking

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.822887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:c8097252584d2061b87c251b067a944fc6db098d53d91ccf05122834db928a0e

Observation 2aaa1d0a-f51b-4518-b377-c470626ae7eb · outbound

This paper cites Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit.

Massive Activations in Large Language Models Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.826534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:9142902a2aa86d907e15568ec457cef40d92ab846e0ff8fcabe230f296503d8a

Observation 0f19f603-6130-40d2-99df-198aea02f32b · outbound

This paper cites Mix Barrington.

Massive Activations in Large Language Models Mix Barrington

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.087857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:4d3048a0643b5c8ab5d5a1135ec9629315b9adabcb700c360067d733b65f8566

Observation a93deff3-640e-4f57-b521-d1f692cdba26 · outbound

This paper cites Mix Barrington and Denis Thérien.

Massive Activations in Large Language Models Mix Barrington and Denis Thérien

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.089754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:1b6e2ef9fffb027ac3cc78b27623a56213cac7a2995f68695e76231f028d5cf7

Observation 283eca06-15f0-40e0-8499-fda607fdc194 · outbound

This paper cites On the ability and limitations of transformers to recognize formal languages.

Massive Activations in Large Language Models On the ability and limitations of transformers to recognize formal languages

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.091861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:345fe13ec070d3456f85de9418fb732dd28aa69aa880a17c605a0e88a24ffde3

Observation 0c657333-c417-484d-ac96-250f846e1ef2 · outbound

This paper cites Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges.

Massive Activations in Large Language Models Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.829983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:c42d8b66f9c5954ebb264f7c405e8fd3ae723597d2a9959654b7ed1955a2a221

Observation c55341c7-b0cc-47f5-8a5d-7ef88b3854bb · outbound

This paper cites Unbounded fan-in circuits and associative functions.

Massive Activations in Large Language Models Unbounded fan-in circuits and associative functions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.096303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:7601c68e4fb83f352c9e73b65eac20bbcfba946c2231edaf58c69ff2e9070698

Observation bfcd1834-5c72-4e04-a434-ce9dd23258e6 · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

Massive Activations in Large Language Models Decision transformer: Reinforcement learning via sequence modeling

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.098393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:5e471053cdf098d8eb7fda163d0d05533126b5d1ff2866d3b739d5f84d178281

Observation aa7b790d-53b2-4273-8abe-1a9e95f6fe11 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Massive Activations in Large Language Models Evaluating Large Language Models Trained on Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.833214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:ad3e5b688438e90fd89a5d3309ec843052bdf7509c89aeab28ffde91e724e882

Observation c8f06064-a61b-424e-a186-62fbc5164034 · outbound

This paper cites Finite-automaton aperiodicity is PSPACE -complete.

Massive Activations in Large Language Models Finite-automaton aperiodicity is PSPACE -complete

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.102743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:f5d6ad7408055625bf031ef296bb006823cdd226f2cd4c9d8c55f1e5cfc46653

Observation 9e437545-8606-4bf4-9adc-1faa13903d5d · outbound

This paper cites The algebraic theory of context-free languages.

Massive Activations in Large Language Models The algebraic theory of context-free languages

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.104694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:a96d312207050c48c84e16150eb2132c2e5c8a74c05bcb29b7329c55d6cf001d

Observation 20bf5ead-4004-41af-b2ca-25e534e3039a · outbound

This paper cites Conditional Positional Encodings for Vision Transformers.

Massive Activations in Large Language Models Conditional Positional Encodings for Vision Transformers

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.836536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:0616867189f0922bcfcc34249adc707dfcd6a23c67fc82fa8f4df9a86518a835

Observation f851dedd-900e-4e83-a2a1-709204939889 · outbound

This paper cites an unresolved cited work.

Massive Activations in Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-16T07:02:54.108357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:47bb74301f404799fb9598afc0b98b612b2e293d7379dcb1f50ab2eaeb21cea3

Observation 24cac180-9e4e-4bee-ac4b-61df7b3946d3 · outbound

This paper cites Approximation by superpositions of a sigmoidal function.

Massive Activations in Large Language Models Approximation by superpositions of a sigmoidal function

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.110385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:5834ab79a7b1f8c62076dc61224213eeaf23b1cd3288363025406c321019d6ca

Observation 12758b5f-19f0-4843-bbc3-47b592a800ec · outbound

This paper cites Depth separation for neural networks.

Massive Activations in Large Language Models Depth separation for neural networks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.112303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:ec447fd411e63817e3aedb8d1d4134a99b0d88ad055a1164fb97d7970344a675

Observation 5760ebfe-4b0f-4557-8a55-fee9a6865c21 · outbound

This paper cites Learning parities with neural networks.

Massive Activations in Large Language Models Learning parities with neural networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.114180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:970368724538001eeb90c0a36880353467e885c06477470dc5434cd5b93ce3f2

Observation 3dd2a47f-bb65-4b1e-80b2-0a91578c0491 · outbound

This paper cites Universal transformers.

Massive Activations in Large Language Models Universal transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.115941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:7d2157fa010424bdb059419570652b29a7ec7374a5fcd1341396fafaf9898f80

Observation 9587bdfb-4457-41ac-a217-9877c405cf5a · outbound

This paper cites Neural Networks and the Chomsky Hierarchy.

Massive Activations in Large Language Models Neural Networks and the Chomsky Hierarchy

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.839680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:19c30c4f798dca02bdd9dcc46ea5acbfe32d15f5ea479843578f074760b259cc

Observation 1438c5d8-e3cb-47f3-9666-1f75af0df165 · outbound

This paper cites Patti, Jayson Lynch, Avi Shporer, Nakul Verma, Eugene Wu, and Gilbert Strang.

Massive Activations in Large Language Models Patti, Jayson Lynch, Avi Shporer, Nakul Verma, Eugene Wu, and Gilbert Strang

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.119816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:c82db7ec907c482fe5df59a3d5d813266b3231401129482b87c752d4392300d3

Observation a92effb4-23f9-4782-9fba-952d5d7acee8 · outbound

This paper cites How can self-attention networks recognize D yck-n languages? In Findings of the Association for Computational Linguistics: EMNLP.

Massive Activations in Large Language Models How can self-attention networks recognize D yck-n languages? In Findings of the Association for Computational Linguistics: EMNLP

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.121875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:c5d63cafaca469d9bc4eb63ef916e02274561475a33a708bbbcca14aa5b44893

Observation e9d71054-971f-43cd-8c8e-c8e76299aaca · outbound

This paper cites Inductive biases and variable creation in self-attention mechanisms.

Massive Activations in Large Language Models Inductive biases and variable creation in self-attention mechanisms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.123803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:f0ba39030658d42ce068ca855945d98e15e7b500750747a5085438f2e043f386

Observation 585d67d4-f497-4b16-9b8e-0013a444403c · outbound

This paper cites Computational Holonomy Decomposition of Transformation Semigroups.

Massive Activations in Large Language Models Computational Holonomy Decomposition of Transformation Semigroups

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.842805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:d49090a6d15d88bc436bf300d912b66e3b4d428ee2cb09c168072cb940a4fa9a

Observation b054a14d-3746-4d23-a476-43ba27cb5b7e · outbound

This paper cites Automata, languages, and machines.

Massive Activations in Large Language Models Automata, languages, and machines

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.127815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:976b5181b5916841f3c737722dfb46b037c95535648659901e9f347633db494f

Observation 3c77fa33-cbbf-43de-9e27-e19eb874af7a · outbound

This paper cites The power of depth for feedforward neural networks.

Massive Activations in Large Language Models The power of depth for feedforward neural networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.130003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:b023db412eb8cc3c8c6b937c2e9a4833f99dcd67184cb3b443c82e1459e57373

Observation f9d7f47e-4285-43e0-b52e-33b1c7728c0e · outbound

This paper cites A mathematical framework for transformer circuits.

Massive Activations in Large Language Models A mathematical framework for transformer circuits

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.132652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:f629a3a6317d38b94ab90a60df4d904cd2c750d962a2674a34ee867425a53ae3

Observation a053446c-b77f-4ea0-ad7e-7a608e4877e8 · outbound

This paper cites Saxe, and Michael Sipser.

Massive Activations in Large Language Models Saxe, and Michael Sipser

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.134628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:3bf62f95058b7b943633a8a5aaaaccbb896af6ad691d4d3e4d8c8ad19bc261b6

Observation 233d2411-6057-4706-b4b1-bc989a61be87 · outbound

This paper cites Shortcut learning in deep neural networks.

Massive Activations in Large Language Models Shortcut learning in deep neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.137666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:6084e09e3d98038f0da7c1cc2217fe86fc1bc3fcd8139af77f20630fa5c3340c

Observation e81aff0e-10ba-4a11-a325-4f2807ad109b · outbound

This paper cites Looped Transformers as Programmable Computers.

Massive Activations in Large Language Models Looped Transformers as Programmable Computers

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.846090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:514742dc5eba454bfaec663ae122d06d1b17ae9040f043a8b6cd7a67e85b1b29

Observation c302ae6d-b427-43a3-af3c-1aadcee5c87f · outbound

This paper cites Reliably learning the R e LU in polynomial time.

Massive Activations in Large Language Models Reliably learning the R e LU in polynomial time

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.146224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:cd7f83a3969896564982db81f8040a388a8b70f8dc721dc2e47fc3aa210f49fe

Observation 8c69106f-3e51-4e1c-ae02-abd60c37a366 · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

Massive Activations in Large Language Models Adaptive Computation Time for Recurrent Neural Networks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.849372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:7d376bc6117605408c06272228e459608b22cbae44a646586240fb563ce35a62

Observation cf40a2ef-ccda-41ad-a4f0-50d4fc44a1c9 · outbound

This paper cites Neural Turing Machines.

Massive Activations in Large Language Models Neural Turing Machines

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.852409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:5a9d40f6dc4c2d128970b7402eee277178fbc71585dff43f69537931781d5c42

Observation 68d22231-dbe0-49a6-89f7-56ece1c3ee35 · outbound

This paper cites Non-Autoregressive Neural Machine Translation.

Massive Activations in Large Language Models Non-Autoregressive Neural Machine Translation

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:02:53.855519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:a7cfd9ede8e848c6de46d58f763159e2a828b543ac91099ea5ca32f42a0ef7fa

Observation 127c3f46-08d3-411f-8224-66699687de2f · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Massive Activations in Large Language Models Dream to Control: Learning Behaviors by Latent Imagination

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.858762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:3272deca808bb82a76c14956fbdceaeb88157ec036204e340413d2178a298a72

Observation 7a56b030-39cc-4fea-ab80-80276c58a064 · outbound

This paper cites Theoretical limitations of self-attention in neural sequence models.

Massive Activations in Large Language Models Theoretical limitations of self-attention in neural sequence models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.158633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:1a8bcc486ba0ad60c3066084180943bcaa80a5a33186614c5e1a14740765ae55

Observation 653efb2c-3ef8-4d44-b276-5311aa9a4380 · outbound

This paper cites Transformer Language Models without Positional Encodings Still Learn Positional Information.

Massive Activations in Large Language Models Transformer Language Models without Positional Encodings Still Learn Positional Information

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.862054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:377664c0057fa3f21364daf31c19359da53930630be9f1988d2ff7b87c4455f9

Observation 74df9dd5-5ddd-4e7d-8d12-57ef312aad7a · outbound

This paper cites Deep residual learning for image recognition.

Massive Activations in Large Language Models Deep residual learning for image recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.181422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:64bfbb06363f133b53d24f34806e18c6b7fbd9d84dd81adf5da18302b41917ef

Observation 04632b0f-3c44-4673-9256-cd26b6f18a59 · outbound

This paper cites Towards lower bounds on the depth of R e LU neural networks.

Massive Activations in Large Language Models Towards lower bounds on the depth of R e LU neural networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.184891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:b1587392327a947f8bbe270f2c849e965e2b2a133e99428303c290da3ba803df

Observation e1e15af2-b2fb-41a1-8fb0-2ddd27a60ec8 · outbound

This paper cites Steele Jr.

Massive Activations in Large Language Models Steele Jr

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.187933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:fb8d5dbf123be4fdba7307c6676eda6bd5284816d6b1387b066acb4e25832ce8

Observation 47e79d1b-50f8-485b-a623-06a975162c1d · outbound

This paper cites Multilayer feedforward networks are universal approximators.

Massive Activations in Large Language Models Multilayer feedforward networks are universal approximators

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.191135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:6ccd961728d8d862af02bd69c17d210185275bd94dedef7a1bec5327f816984e

Observation 96fb3a48-c9fa-48c0-8c27-01efe73efd82 · outbound

This paper cites Universal Language Model Fine-tuning for Text Classification.

Massive Activations in Large Language Models Universal Language Model Fine-tuning for Text Classification

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.865528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:8f4736376e4ea812dd8fd31297cc4454fefb9c5e43fc8f40e220c287f47c846c

Observation 68e668c7-e742-40bc-b1f2-6369c30a19c3 · outbound

This paper cites Block-Recurrent Transformers.

Massive Activations in Large Language Models Block-Recurrent Transformers

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.868612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:dc54f871ee50eb42c020d7ad089a921b4716924aa089fa904cdd49a9646d22a6

Observation 76e8e6fe-dfd2-497f-aa61-85680fe996a3 · outbound

This paper cites Offline reinforcement learning as one big sequence modeling problem.

Massive Activations in Large Language Models Offline reinforcement learning as one big sequence modeling problem

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.200656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:5d643e040d82a37d4f19c6b0dd29703498ded3a0d95af64ed46d883d85216d2c

Observation db9ec858-602d-451c-9fa5-9eeb85bb7ab9 · outbound

This paper cites Finetuning Pretrained Transformers into RNNs.

Massive Activations in Large Language Models Finetuning Pretrained Transformers into RNNs

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.871604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:0e1fe4e662967122e3ea7d4e7c7dbba47ef7c97f1fbcafbcefd4b9f38c26c317

Observation 3a7f9b6d-e7bb-42ed-84d1-7cf952c15b8a · outbound

This paper cites Rethinking Positional Encoding in Language Pre-training.

Massive Activations in Large Language Models Rethinking Positional Encoding in Language Pre-training

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.874934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:a828a5aa3a13b2e7e5995b6e0fd1756be8a56d9d632dc399b985f280d6462086

Observation c7d13e69-2620-49e1-a2fe-d5a3176b760e · outbound

This paper cites The number of semigroups of order n.

Massive Activations in Large Language Models The number of semigroups of order n

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.209735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:6cefe2969a8d8c65e08b09e99808a0f564150973819582ea01c679329b13b509

Observation dbde73a0-d7f7-456d-b251-9bd3520677a7 · outbound

This paper cites Finite permutation groups with large abelian quotients.

Massive Activations in Large Language Models Finite permutation groups with large abelian quotients

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.213125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:36165af93536f1acc51d79c012d50d9b53e07fca47ba1c3064d6035dc91c7744

Observation 6b88597f-d4af-4ca5-945d-44bbbca1878d · outbound

This paper cites Produit complet des groupes de permutations et probleme d’extension de groupes II.

Massive Activations in Large Language Models Produit complet des groupes de permutations et probleme d’extension de groupes II

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.216542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:32d25f92454de9783e7465fbda6d617b63b922dbd5fe958247766084eaa0fed6

Observation 59756881-3cb6-4eeb-bdf4-42e54e602ae0 · outbound

This paper cites Algebraic theory of machines, I : P rime decomposition theorem for finite semigroups and machines.

Massive Activations in Large Language Models Algebraic theory of machines, I : P rime decomposition theorem for finite semigroups and machines

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.219974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:45d50f7013bf0d3ff8584167f9aa53650c9a304b4cce6a6a7554fe26c19e6df1

Observation e5a57c1f-6aa7-4811-9b0c-a7a2fa2bccab · outbound

This paper cites Deep Learning for Symbolic Mathematics.

Massive Activations in Large Language Models Deep Learning for Symbolic Mathematics

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.877938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:680981cfd0396bd91ee6299a0d0580f5d860c301ccdf152664b688df82191faf

Observation b8f8724a-2a4f-438c-84cc-792030f1fa64 · outbound

This paper cites FractalNet: Ultra-Deep Neural Networks without Residuals.

Massive Activations in Large Language Models FractalNet: Ultra-Deep Neural Networks without Residuals

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.880895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:b841204d7b0056da805f753ac15731d7a4fea860ceee93f93d0ad5f4334ffe1a

Observation d17ad449-261b-4f31-82de-492cb1bfac53 · outbound

This paper cites On the ability of neural nets to express distributions.

Massive Activations in Large Language Models On the ability of neural nets to express distributions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.230282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:30cab5ef32c98a3b19d31bf9cec93c5fa4a34d66e1d007e4cc052a0e5f78287b

Observation 3576dd43-7ecb-4202-8aff-56b9e07acc0d · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

Massive Activations in Large Language Models Competition-Level Code Generation with AlphaCode

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:43:48.088484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:fade8a9c0fcb849dbff44ad9115c3590c8cf625acf3edf02f8bb6121f97dbe78

Observation faf9b645-badb-4687-901e-ea787e9f8423 · outbound

This paper cites Decoupled Weight Decay Regularization.

Massive Activations in Large Language Models Decoupled Weight Decay Regularization

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.887986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:af706e36888fa0d79d3174948ad549c3a173587182cf9f97108383e6cc87959d

Observation 5485814b-ba2f-4b17-b144-ae447a2d7e27 · outbound

This paper cites On the K rohn- R hodes cascaded decomposition theorem.

Massive Activations in Large Language Models On the K rohn- R hodes cascaded decomposition theorem

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.239532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:838924329fee6a9e614acc5ad5ed8fcb8b6d61de0e310abddcfe3e0aa745927b

Observation b007f765-2385-4b31-acc8-14c145a6d39d · outbound

This paper cites On the cascaded decomposition of automata, its complexity and its application to logic ( D raft).

Massive Activations in Large Language Models On the cascaded decomposition of automata, its complexity and its application to logic ( D raft)

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.242291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:7c2193a626200935c4bf9affc321fc6a1c3f9a8bdd7aa09eaf20383f257e5aec

Observation 31dce5d4-1a0d-4c37-b2fe-367d9443d2f6 · outbound

This paper cites Threshold circuits for iterated matrix product and powering.

Massive Activations in Large Language Models Threshold circuits for iterated matrix product and powering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.245277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:6fbf97031730646dae6a01f1196aadcfb976eed9e7e801bdabf018dc5d135305

Observation 46f67443-bc99-49d0-9490-7fc1ac0ea102 · outbound

This paper cites Saturated Transformers are Constant-Depth Threshold Circuits.

Massive Activations in Large Language Models Saturated Transformers are Constant-Depth Threshold Circuits

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.891005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:57381a45bcb3f64e689721171193f0a2775b41ccd716d105f02da8c81e781a2e

Observation c4a750c0-6672-4e3c-8023-4b52e399a5d5 · outbound

This paper cites Transformers are Sample-Efficient World Models.

Massive Activations in Large Language Models Transformers are Sample-Efficient World Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.894196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:a20c9b2d46ac9965f25b0422d9161b1e40f092da5da957ec2f80d8177fc68a68

Observation 15333cc0-9d04-498e-a6de-c19c7116fe3c · outbound

This paper cites Lower bounds over Boolean inputs for deep neural networks with ReLU gates.

Massive Activations in Large Language Models Lower bounds over Boolean inputs for deep neural networks with ReLU gates

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.897343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:35f43a637b48010b8db42936cb8329df1510f3987a76ab8e8d7379cd3ecefe5d

Observation 29ea9dea-ceda-4589-bec7-1f83577a80d0 · outbound

This paper cites A mechanistic interpretability analysis of grokking.

Massive Activations in Large Language Models A mechanistic interpretability analysis of grokking

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.256013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:f38267d33f6fee0b4992ddd8e5f246b3fb6c54e5147fe8090efbd865fc6e4bd1

Observation 794952e7-ffaa-4cff-a572-5994a2a6b0a1 · outbound

This paper cites an unresolved cited work.

Massive Activations in Large Language Models Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-16T07:02:54.258406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:308de10de05f750e3eef7526ac6f10f463de0c34cbdbac34e863fcb9563c88cf

Observation 6c66c989-11d2-4082-8819-e90c99765ee8 · outbound

This paper cites Identifying good directions to escape the NTK regime and efficiently learn low-degree plus sparse polynomials.

Massive Activations in Large Language Models Identifying good directions to escape the NTK regime and efficiently learn low-degree plus sparse polynomials

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.900523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:fd13535aafb0293a8d4d7d41c9a39cfbd867f0c3895263d4b4040539ecd31318

Observation efcc52a1-63a1-4fb6-8cf7-19a34046c046 · outbound

This paper cites Investigating the Limitations of Transformers with Simple Arithmetic Tasks.

Massive Activations in Large Language Models Investigating the Limitations of Transformers with Simple Arithmetic Tasks

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.903787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:698c0694922706888c02f35dd9c1fa88583577d0b525be88597bc51a7a359a8f

Observation 23fb748d-b973-4201-b7e2-abcd1725ceee · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Massive Activations in Large Language Models Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.906902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:7b1d5b036b67ca0f697ff535966259db814e30e1e597cf7124954a80ccc76f95

Observation 919bd7a7-9e06-4204-bfb6-2675656d8e5d · outbound

This paper cites The complexity of M arkov decision processes.

Massive Activations in Large Language Models The complexity of M arkov decision processes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.267563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:ce2ec3f5895816e11e6e1996c34dfe8b075ab541fe9e2c2ad18c02b6283dce99

Observation ce8c7133-1e27-47e1-9a13-3036070ed7b5 · outbound

This paper cites Py T orch: An imperative style, high-performance deep learning library.

Massive Activations in Large Language Models Py T orch: An imperative style, high-performance deep learning library

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.270094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:2d862bb8d4b302ab4bbc899f831652e8925d23516e3e08128c1510056935f7fc

Observation 22a9f215-6e25-4c69-8d0c-ad6e8d24b5b9 · outbound

This paper cites Attention is turing complete.

Massive Activations in Large Language Models Attention is turing complete

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.272370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:d181252dec025725ce9d855502b6cae6cb2553cfa48f4e93a5364a17cdc248f3

Observation b66351a4-83ca-44e6-9307-ccbce2b6a2c7 · outbound

This paper cites Deep contextualized word representations.

Massive Activations in Large Language Models Deep contextualized word representations

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T07:02:53.910338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:61bb48ace4af150f7789a97e671b4e16ff3f0a543442bda5784da0b33ca2de8c

Observation 9bec790d-b069-45cc-93ed-ef94672a6338 · outbound

This paper cites Generative Language Modeling for Automated Theorem Proving.

Massive Activations in Large Language Models Generative Language Modeling for Automated Theorem Proving

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:18:10.906922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:63bea8581f7b2695ba90b601fb452f933a291e0c2c14a6df326cbfe091903e48

Observation ed3a336f-a298-430f-9e4a-681e03d47ff8 · outbound

This paper cites Train short, test long: Attention with linear biases enables input length extrapolation.

Massive Activations in Large Language Models Train short, test long: Attention with linear biases enables input length extrapolation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.281823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:4cfbc224ac18e574f6e8e7ba276d956a18cb616c6545c00d745d75612cda863e

Observation daed02a6-0f65-43b5-a264-a14fb8bf4b36 · outbound

This paper cites Language models are unsupervised multitask learners.

Massive Activations in Large Language Models Language models are unsupervised multitask learners

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.283909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:0965179f3408f892f84d8b516d8528ff058863aa14e471c6142d008de53765e7

Observation 155f3d6b-cc80-4e4f-99d5-9b797f5a52fc · outbound

This paper cites Reif and Stephen R.

Massive Activations in Large Language Models Reif and Stephen R

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.286539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:68f2ae3e9ca8a1dcc94e46fc1af2cd11bfa68d650152b5447c90f350ea18dfc0

Observation fd95f318-987b-4d3b-bb37-7cfd3b71e1a8 · outbound

This paper cites Applications of automata theory and algebra: via the mathematical theory of complexity to biology, physics, psychology, philosophy, and games.

Massive Activations in Large Language Models Applications of automata theory and algebra: via the mathematical theory of complexity to biology, physics, psychology, philosophy, and games

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.288966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:7351ea3bd3a65453f6157c2081ca5ba1fc6aff94be2bda8fb201537da251c3c9

Observation bfd3ffa3-c957-4cf2-b584-4ccace65d90b · outbound

This paper cites Can contrastive learning avoid shortcut solutions? Advances in Neural Information Processing Systems.

Massive Activations in Large Language Models Can contrastive learning avoid shortcut solutions? Advances in Neural Information Processing Systems

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.291209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:2fef3ff4be9f8363da222d91be4564693dab235af0ec74d1f02205b3b8d16d36

Observation 6dddc637-081d-4bd6-ae71-da93130e882a · outbound

This paper cites Depth separations in neural networks: what is actually being separated? In Conference on Learning Theory, pages 2664--2666.

Massive Activations in Large Language Models Depth separations in neural networks: what is actually being separated? In Conference on Learning Theory, pages 2664--2666

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.293406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:fd88949ffee4bb597c2337247bde4157abebefe343bd702c56c61ad7e8b94001

Observation 8e9b267b-53e0-4d96-ad42-5d0d1a8f979b · outbound

This paper cites Programming puzzles.

Massive Activations in Large Language Models Programming puzzles

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.295467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:718596ce95dfa0964c7bd470296a892064187b8b739a2a6f584e4c18fb3ca807

Observation 96d86380-8c48-4056-a082-7be0ac86f010 · outbound

This paper cites On finite monoids having only trivial subgroups.

Massive Activations in Large Language Models On finite monoids having only trivial subgroups

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.297634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:97ea7224a25b734d0a66543c11ae7045a164c95d24852f3ee7dd2723777f642b

Observation e6c89032-ba93-44b7-bf40-f51c2da5b9ce · outbound

This paper cites Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks.

Massive Activations in Large Language Models Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.076502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:ba2a9a5787debc6a7fbf44088f84de89c3763a71612373c7daf077feb955a026

Observation f5ad8215-382b-41ae-b4a3-f8c1a7ef8088 · outbound

This paper cites On the computational power of neural nets.

Massive Activations in Large Language Models On the computational power of neural nets

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.078726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:b328148bffa4d565db9663c24e55e3cf205aac4dae977498fcd1a512f7af883b

Observation fc8285ad-55a1-47c2-8e7e-0b914092f37c · outbound

This paper cites Benefits of depth in neural networks.

Massive Activations in Large Language Models Benefits of depth in neural networks

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.083609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:faeee444e46787af9a2c448b44254fc606b438c13b383fe53f7755011c57d9f5

Observation cf386873-078d-4130-bbdf-4b4f0b19ec21 · outbound

This paper cites BERT Rediscovers the Classical NLP Pipeline.

Massive Activations in Large Language Models BERT Rediscovers the Classical NLP Pipeline

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:53.916830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:7132b07b301094b0b73e106a67762c2dc6e2beacac840d1e73c5ccf377a25df4

Observation 90c7731a-54c6-40fe-b20d-fcd8dadc4053 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Massive Activations in Large Language Models WaveNet: A Generative Model for Raw Audio

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.919878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:237d8ce80fc0e1faa151f861e4bad0a609a6613f0b3f57433cb00196097afac3

Observation e4eb709b-3634-4a27-a53b-41b0fba961c9 · outbound

This paper cites Attention is all you need.

Massive Activations in Large Language Models Attention is all you need

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.100746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:dd8f473f879f91edb15b732007dee610693b96e8c717530af59b9e981a202fda

Observation 6592dcfc-0acb-49cc-aaaf-d1c573cfb465 · outbound

This paper cites Hechtman, and Jonathon Shlens.

Massive Activations in Large Language Models Hechtman, and Jonathon Shlens

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.106570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:23da385588692f9c0e66f03e7f11c8239dc31c186ba9349573ffed2c53678a97

Observation bd94ffee-00a1-43de-ab48-9479bc180edb · outbound

This paper cites Visualizing Attention in Transformer-Based Language Representation Models.

Massive Activations in Large Language Models Visualizing Attention in Transformer-Based Language Representation Models

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.923110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:01de06cf0ac6a77d9ea9389e7874ae4279a1720930142d6ba50845f0c9dd7ae3

Observation 06c1682f-2943-4fdf-9c88-a1c241b5c5a9 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Massive Activations in Large Language Models Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.926612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:06f7d920fee4c1dc5185199566d8d9b4a2985cb4edee533b421298f5656cb592

Observation ec2f7afa-eb70-4ad6-8534-45b117f270dc · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Massive Activations in Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.929799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:9abb95df68cb57e3dc4f1681f3756008c9a3f57b77d8ddf3850c2df19c5cd5b1

Observation 5722265a-b603-46c0-a7cc-2bd47d81cc22 · outbound

This paper cites Thinking like T ransformers.

Massive Activations in Large Language Models Thinking like T ransformers

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.148694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:b5b5bebc0bf1cb3fad542d5d456722fe0eea6618b39a8b688ba2e34fa41c6c8c

Observation e1b26e7d-8922-46c7-80b0-b7f77f5fe92f · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Massive Activations in Large Language Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.933024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:52f1275a07510fd5036a5724ad673a11a1ea2acc90f2a3de1124662eb524ab89

Observation 4cd7e10a-3919-4403-bdf1-9bc8c6bee864 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Massive Activations in Large Language Models Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:53.936699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:e98b424780568cb0ee6a87365fd0b00de9fca854df6b38cc9cbd68247f989ff9

Observation 7e01e40b-bf8d-40be-9e15-c29aab96af9b · outbound

This paper cites A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond.

Massive Activations in Large Language Models A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.940376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:20dd3277f8786676c93d37df796aab339c0d4a3b47d579f25a6faf9785212b29

Observation 3c1b399f-6259-4e54-ab8a-c104186c7f9a · outbound

This paper cites How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks.

Massive Activations in Large Language Models How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.944970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:dd44cfc2e2262544a7067d5c97fe250f1596a95089cc68f8b620802d5e492ec7

Observation e01ed2c6-4416-4824-9291-6aa1e0219a6f · outbound

This paper cites Papadimitriou, and Karthik Narasimhan.

Massive Activations in Large Language Models Papadimitriou, and Karthik Narasimhan

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.194251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:302717eb260582f8e9f6163de85e09d448f2c6a49b5b2d0376a19c606a7bf370

Observation 808e24e5-1ade-43bc-8f4e-b2beb6c0c6af · outbound

This paper cites Mastering atari games with limited data.

Massive Activations in Large Language Models Mastering atari games with limited data

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.197478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:b6723215559709eab0214e8e3c979844da8b80fb3c880dad014cc1fdb973366f

Observation 49649fc0-c2c0-4f38-8b26-6ed9cea981d7 · outbound

This paper cites Cascade synthesis of finite-state machines.

Massive Activations in Large Language Models Cascade synthesis of finite-state machines

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.203741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:1d942369eebb26d85a0de297a2f9dc9a32898d9db6af7b1d5a767f3d6ed0509f

Observation e2fd5998-2df1-4a8b-8738-4969b2513d5b · outbound

This paper cites Pointer Value Retrieval: A new benchmark for understanding the limits of neural network generalization.

Massive Activations in Large Language Models Pointer Value Retrieval: A new benchmark for understanding the limits of neural network generalization

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.948453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:f990db2afe255673bf04295155746ee98e97526ef00f4194723d7d15ad3c6980

Observation 693bd688-01ea-445e-b965-5fee4316af7c · outbound

This paper cites How does mixup help with robustness and generalization? In International Conference on Learning Representations, 2021 b.

Massive Activations in Large Language Models How does mixup help with robustness and generalization? In International Conference on Learning Representations, 2021 b

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T07:02:54.223517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:aae043b5a29882ee0d78ac9bbff6d09c7a8f2de2db289d43943ca99ff4f10cf4

Observation 609d4f79-b35b-426f-80a3-9e7d0721a608 · outbound

This paper cites Unveiling Transformers with LEGO: a synthetic reasoning task.

Massive Activations in Large Language Models Unveiling Transformers with LEGO: a synthetic reasoning task

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:53.952149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T07:02:53.740597Z digest=sha256:312aad196ebefa198524cb5669fe3159a2b95d6d3c20f0959f22e87274bbfe4e

Pith citing papers

Observation 0f564b89-e117-4001-8b73-a4d58afb5382 · inbound

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling cites this paper.

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling Massive Activations in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T09:58:29.057357Z digest=sha256:37787128852f1a79de81449d7239c430cc3b179a6f05ef89c01fe9cba30f13ea

Observation 09dd05df-56e3-4318-9adf-f4a805d11c3b · inbound

Scaling and evaluating sparse autoencoders cites this paper.

Scaling and evaluating sparse autoencoders Massive Activations in Large Language Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T17:47:23.089288Z digest=sha256:4bad71db2c5f8c012c5326cb0ce3808770f62cc494635318ac55256304c49901

Observation b383fa3f-ff15-4c4f-8dea-dbca88241704 · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision Massive Activations in Large Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:45:36.442032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:f106e78ad549f7b84af5c89cc195abbd357b784bccd422547bd6e1f59ebfd5c3

Observation 6ba80cee-860c-4379-af58-912b3c3f9c9f · inbound

When Attention Sink Emerges in Language Models: An Empirical View cites this paper.

When Attention Sink Emerges in Language Models: An Empirical View Massive Activations in Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:41:03.745535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T17:41:03.674759Z digest=sha256:9debf0401552c20f059fe8f3c7d3c4283a720db0bb81979e34dc04c317353ad6

Observation 12916e7a-8ac9-475f-9935-e4e8a7a1d8fc · inbound

When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training cites this paper.

When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training Massive Activations in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T16:30:21.317662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:30:21.317662Z digest=sha256:313114204babe33d534171870f64d161374e0adddafa9fb9200290f44d6163b2

Observation fd48f046-aa25-41ac-bf2b-ac475226fc33 · inbound

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps cites this paper.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Massive Activations in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.393582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.393582Z digest=sha256:c1c5dd87819d2e8562216f72e6602203bae13c7f4c4179ad8eb6c454183ac0b7

Observation 99fdcb51-0b0b-4fd4-a04a-5cb8a1238b9b · inbound

DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation cites this paper.

DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation Massive Activations in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:23.900797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:20:23.900797Z digest=sha256:79652dd1539850988a0d0af2f2541e592efc4ad4838eb711abf08ccbf77e9f24

Observation a53fad53-d3fa-47da-a067-5aa36e15bb28 · inbound

TinyFusion: Diffusion Transformers Learned Shallow cites this paper.

TinyFusion: Diffusion Transformers Learned Shallow Massive Activations in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:40:19.584749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:40:19.584749Z digest=sha256:88342051eef61db27d0839ed50d5274d17f9feb6a75b9a453f36e57f1e304bfe

Observation b425e016-3674-438a-b52d-afd409cf979b · inbound

MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic Segmentation cites this paper.

MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic Segmentation Massive Activations in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:23:56.974512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:23:56.974512Z digest=sha256:c1785a05200186257a14b3ebd388995122ef80e49a841aade7197957d218411e

Observation 6e72f8a0-3c79-43c0-a4c8-d86e78caafb8 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Massive Activations in Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:46.996829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:46.996829Z digest=sha256:240461eead77bcd52b23391f6edc2d8399a558c7b01e343148bb6a606b1916af

Observation d54ef8b7-9dd2-4ba5-a3cb-254468d55c91 · inbound

Leveraging Registers in Vision Transformers for Robust Adaptation cites this paper.

Leveraging Registers in Vision Transformers for Robust Adaptation Massive Activations in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:37.720335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:37.720335Z digest=sha256:e0f65d8371d226e036190086882430b8236cbe2b8b12d8b96a575f4f1e3af225

Observation 036eb6bc-19f3-4c1e-879b-2877d18b95b1 · inbound

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models cites this paper.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models Massive Activations in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.668507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.668507Z digest=sha256:d9ac781c66437750c8eb96bdc3167ce815355d38e166711c61352cd356dfb3b6

Observation 91fcc470-d158-4f3a-9e65-32c8a9810006 · inbound

RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations cites this paper.

RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations Massive Activations in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:54:22.715681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:54:22.715681Z digest=sha256:0a5d50f8dfdb1806d543b69bcdb9750c828e3fdcb3600c3416ef3cae6e3a275c

Observation ceeaf3e5-1f44-4d4d-8e3c-c1ce36297e03 · inbound

An Inquiry into Datacenter TCO for LLM Inference with FP8 cites this paper.

An Inquiry into Datacenter TCO for LLM Inference with FP8 Massive Activations in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T16:47:23.623705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:47:23.623705Z digest=sha256:cd6dcd5860a0f6b02b66ccfa2ef4264a51b062c3eb8c34131abde7186b09e9f4

Observation 94b3edb7-d041-4212-bc15-c788e160fcef · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Massive Activations in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.993209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.993209Z digest=sha256:56791171d57cbee5e304117f2879de22d190786bcac852b8b0e6d10d6de546e4

Observation 2ac3d208-f294-4d75-aeee-520b786000be · inbound

Systematic Outliers in Large Language Models cites this paper.

Systematic Outliers in Large Language Models Massive Activations in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:37:37.525550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:37:37.525550Z digest=sha256:7cf8fc3ba46f03646c5752e05adeb5cb94e861bac94310cb42a94b182891f03b

Observation 5341b35f-be30-4e0b-9ff7-fad88cf968c5 · inbound

TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks cites this paper.

TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks Massive Activations in Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T06:06:57.523801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:06:57.523801Z digest=sha256:80d763288afe16f22a79857be12b058126a4eb9b10cc7bda734be696d2654d2a

Observation eb2872e6-71dd-45b3-9642-d1d671307c1f · inbound

Quantifying Memory Utilization with Effective State-Size cites this paper.

Quantifying Memory Utilization with Effective State-Size Massive Activations in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:23.672217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:58:23.672217Z digest=sha256:e88a4c4b6e6e90e7c3ab9f3b4d9a74ff83e23da74b986151731c5344fe64bdbc

Observation 45cac19a-40ed-437a-98ad-ca2e1c43b986 · inbound

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models cites this paper.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Massive Activations in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.907873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.907873Z digest=sha256:a4b57f15cbe74806f5db8578f79e2ad103906006ab6b5493e042586a3be93ff8

Observation a506b6ff-1d76-45c4-9e54-27ef1cced7a3 · inbound

ICQuant: Index Coding enables Low-bit LLM Quantization cites this paper.

ICQuant: Index Coding enables Low-bit LLM Quantization Massive Activations in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:49.755531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:43:49.755531Z digest=sha256:b7d3f76c9221325e6f4c992ec9792af9de5b2517afbc0d3d0a9153a34341bb4d

Observation bc90ad85-5426-4828-a92f-99e0c69fa85e · inbound

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design cites this paper.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Massive Activations in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.142490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.142490Z digest=sha256:31a4cbf63280a548bc33563fdf356e9cb5156724f282ab7740a136990f031e9f

Observation 8d42df7f-ea28-419d-9a80-f2464a4d062b · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Massive Activations in Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:a3e3f2756c8658052fea249ac90d78b09bb4cf4683f6caa776ccc2c6e9a00124

Observation ee3f5c68-3595-42c2-95cf-a2be8e62329c · inbound

Polysemy of Synthetic Neurons Towards a New Type of Explanatory Categorical Vector Spaces cites this paper.

Polysemy of Synthetic Neurons Towards a New Type of Explanatory Categorical Vector Spaces Massive Activations in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:04:37.521728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:04:37.521728Z digest=sha256:7e1f34b8cfae366a6326f9a3a4275acb1010abfa044079fc83c3ae20af2bbeee

Observation 6f2c8c5c-774a-4a46-9200-b00f149dacea · inbound

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference cites this paper.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Massive Activations in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.737732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.737732Z digest=sha256:cf69c18ddeefa8204a5700a6177e734e8357775f75d93a3def7d9dde8981cb11

Observation bb0f94bf-7c2a-4275-9518-11fbd28bcacf · inbound

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations cites this paper.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Massive Activations in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.759191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.759191Z digest=sha256:b52cfb6c0148ab6588f6fdc4da169b8dd36e6983eed2c20f835d959a1191543d

Observation e8bbd1dc-631f-4ba0-bba4-95b283e3cd07 · inbound

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? cites this paper.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Massive Activations in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.968312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.968312Z digest=sha256:bdc2707be02b9249803e603f43fd5232b2558d066c88744a1dff0b0ac56ed4c0

Observation 0146fe6b-aa47-49dd-86b2-f7f12d72334e · inbound

Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI cites this paper.

Making Sense of the Unsensible: Reflection, Survey, and Challenges for XAI in Large Language Models Toward Human-Centered AI Massive Activations in Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:36:55.342354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:36:55.342354Z digest=sha256:e56e6c5f3d6c63f724837aff2bae158b30350c0a732f6d4b6d178999fde1bfb3

Observation 51c67f3b-c46e-4aa9-80f0-31229390a4ca · inbound

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study cites this paper.

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study Massive Activations in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:40.241455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:30:40.241455Z digest=sha256:732b1f135ff83016c610c680ae0a06987bb50de78da7ce1ddafa3ad1278c4831

Observation dbc63c28-99f4-47d7-b4b6-d19a1ba1c380 · inbound

FPTQuant: Function-Preserving Transforms for LLM Quantization cites this paper.

FPTQuant: Function-Preserving Transforms for LLM Quantization Massive Activations in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.694378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.694378Z digest=sha256:6da9f511cdf9b02ec9065a3e5cb4d63136d20a87e4eb507eb284cae6cdf7c1ba

Observation a8cfbe6d-2231-4c8b-8721-906e352e75b4 · inbound

PEVLM: Parallel Encoding for Vision-Language Models cites this paper.

PEVLM: Parallel Encoding for Vision-Language Models Massive Activations in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:37:59.050314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:37:59.050314Z digest=sha256:4412220cc13c6cb49459d8e0851132a346f85f56604d9d89e8a6dbcf8c0eb67e

Observation 2e1d3a8a-d15c-409d-bba9-fc01f777ec28 · inbound

Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models cites this paper.

Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models Massive Activations in Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T18:36:04.236333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:36:04.236333Z digest=sha256:ab5cd864f14afe7d1a3039cdc57371accbafe0717ce726cb4a2d665fc09a8f07

Observation a9597c0d-7cd2-468a-9765-fd31b87c8fbe · inbound

Characterization and Mitigation of Training Instabilities in Microscaling Formats cites this paper.

Characterization and Mitigation of Training Instabilities in Microscaling Formats Massive Activations in Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:47:52.279011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:47:52.279011Z digest=sha256:bee1dfb0655d51817f25273022ebe36a5eabfd5d503f7e2984d48cf6135775ec

Observation 328b74aa-c4eb-4829-b374-fc4130bcb1e3 · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Massive Activations in Large Language Models

Reference 130

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.342861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.342861Z digest=sha256:178c31565284c6f37aaf720a12b99dcb2d14db00805df61dd09cd340beae5d4c

Observation ac939149-f285-4c96-b1e9-8d3269bf9ef8 · inbound

Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers cites this paper.

Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers Massive Activations in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:49.076209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:49.076209Z digest=sha256:0c712aa66a082c52926aaf3bd0d9b5175fe0c349952816646925588c74603e33

Observation 127db34a-44aa-44c1-9ba1-042b06b07865 · inbound

Mitigating Spurious Correlations in Weakly Supervised Semantic Segmentation via Cross-architecture Consistency Regularization cites this paper.

Mitigating Spurious Correlations in Weakly Supervised Semantic Segmentation via Cross-architecture Consistency Regularization Massive Activations in Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:25.673273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:25.673273Z digest=sha256:c3eb6b488ef9619efe8d3890c34d265e63e856b874462364736a1f0b5da051e1

Observation a4a28053-910a-4330-9e3e-a3bb82c42a71 · inbound

Discriminating Distal Ischemic Stroke from Seizure-Induced Stroke Mimics Using Dynamic Susceptibility Contrast MRI cites this paper.

Discriminating Distal Ischemic Stroke from Seizure-Induced Stroke Mimics Using Dynamic Susceptibility Contrast MRI Massive Activations in Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:56.368022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:56.368022Z digest=sha256:314ab701f2cf74d7207a3a0f27edc56e196c2b3d5544618c3cc7d4177ad49650

Observation 0ef4d47a-fef7-4492-9b73-f80f95bab925 · inbound

CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models cites this paper.

CAT: Causal Attention Tuning For Injecting Fine-grained Causal Knowledge into Large Language Models Massive Activations in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:32:13.895867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:32:13.895867Z digest=sha256:95f45362db536065be94945a08f2ca83a5287451f60133d4a8c3228bd9b1c41e

Observation 74c8d797-2c9b-4f86-b5a7-7b1ba4d2cee5 · inbound

Prophecy: Inferring Formal Properties from Neuron Activations cites this paper.

Prophecy: Inferring Formal Properties from Neuron Activations Massive Activations in Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:31:24.705418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T13:31:04.425374Z digest=sha256:b791f82845ff7af6ec8b9587f349e4ed2b14f9b690b61ec4e7150f6fb5a84bab

Observation cba88781-4174-4f78-8112-203773de3080 · inbound

CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs cites this paper.

CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs Massive Activations in Large Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:14:00.183201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T04:13:13.231293Z digest=sha256:3cffed9e5fe3ceadef118886102d3ff0b0700342cec2efc0da1593ae775d8dbe

Observation 11853ee6-3d40-4068-ac60-7334cef886f2 · inbound

MiMo-V2-Flash Technical Report cites this paper.

MiMo-V2-Flash Technical Report Massive Activations in Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T11:33:32.568261Z digest=sha256:c0b2f59569018dc2b22de19f3c0346e4870a93a91eed8eeb137c2601ecc784ff

Observation 7ff386ff-d10b-413c-842a-21c3972864b8 · inbound

Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers cites this paper.

Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers Massive Activations in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T11:16:00.419063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:16:00.419063Z digest=sha256:5c10d40ee2e41d7c32dfc2508ceb55f787de79f4a52bf19795f7c00df6b43a06

Observation 3020cc14-71ed-4415-8184-1da2152ecdd9 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Massive Activations in Large Language Models

Reference 291

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:40:54.708987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:46dbaca6eefb4f54f29cb328ba3cfb8ee33f4acbb1f3d3019ff74637a6db202b

Observation ffc5e7d6-2ead-4d3b-9349-b6a1fc54ba35 · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Massive Activations in Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:98e960c6dc1f441915e29daa4ade0fab58cff7da9229c4dccd703f59616bf253

Observation e7f4d679-9ca7-453f-ba8c-a03d7c8ba61f · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge Massive Activations in Large Language Models

Reference 122

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:76e0837df4686f51c0c0cbc8fa8044bec7fe6128be2ec535aae634ef4352b347

Observation 62b6b7d8-2d88-4b30-a58a-3d7207d95271 · inbound

When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models cites this paper.

When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models Massive Activations in Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T23:20:51.899127Z digest=sha256:dfde64dbd0e0d0470a1935b97bf5807f2d082ac3566e7e6bcc56fb5951f2c5f8

Observation 6dee6136-6b0a-4d77-b30f-d58da2d1b15a · inbound

When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models cites this paper.

When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models Massive Activations in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T17:04:24.797847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:04:24.797847Z digest=sha256:147de9bc94a73b4d63efcc5427ce0236c23160fdd6a74430f73640ba28b3a37b

Observation c80a1876-0fdb-4220-b322-79369b779401 · inbound

Noise Steering for Controlled Text Generation: Improving Diversity and Reading-Level Fidelity in Arabic Educational Story Generation cites this paper.

Noise Steering for Controlled Text Generation: Improving Diversity and Reading-Level Fidelity in Arabic Educational Story Generation Massive Activations in Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T19:39:07.039348Z digest=sha256:a361c419b9b3f6e6ee04fd83c1b6ea42c701486c6e9da030642a1d8e0739697b

Observation f49cec62-35b2-4572-bea6-adc1f4d780ae · inbound

OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension cites this paper.

OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension Massive Activations in Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:15:59.622461Z digest=sha256:45bc361af2191d71c530627e42694a7dc74207b8312fc6e91a4db27184824c71

Observation b1dc4a5b-5ac9-4a65-81c5-c0d9c08aab5e · inbound

Graph-Guided Adaptive Channel Elimination for KV Cache Compression cites this paper.

Graph-Guided Adaptive Channel Elimination for KV Cache Compression Massive Activations in Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T07:06:08.786535Z digest=sha256:88377ca3063d1514ed773c0c585fab77ee1b6438141d4349583c2a82922a5529

Observation 83b3ef74-a5a2-466f-b0c2-4c4937fa1a99 · inbound

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization cites this paper.

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization Massive Activations in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:59:23.738476Z digest=sha256:3a169260e18c050c17d04d3796765019f4b596f5357fc7d80cc0c162501cbe4d

Observation 95e5c181-0bb2-476c-87e6-8d1b341ab6f0 · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs Massive Activations in Large Language Models

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:b767d915123d6e96896d7dd7a1b569f2e0a201854f3c3681b6ad61bc524a6660

Observation 2dd3e400-3304-40f7-b222-803c265efc7a · inbound

Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing cites this paper.

Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing Massive Activations in Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:09:43.879809Z digest=sha256:44b4e87654bb70b90d22c1610aea9b4c88f9f0d82f54a86da02d8d7f41a4f8ed

Observation 55732636-5021-41b6-9195-eab9969b97c2 · inbound

Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay cites this paper.

Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay Massive Activations in Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T14:24:54.946996Z digest=sha256:1875355132bc8c6ac48bba0ffe3fd077bed7a99c8cfd4baca74fc85181693196

Observation 9b70fc32-94e4-4282-b233-b2fae0127f2d · inbound

Taming Outlier Tokens in Diffusion Transformers cites this paper.

Taming Outlier Tokens in Diffusion Transformers Massive Activations in Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T17:20:16.402422Z digest=sha256:09531c4e629f7c5a4ab777be3fce6751d81b04ac99dce8142d4596ea02db4448

Observation d6009a94-3323-4e32-b7bd-c6b32bf871cc · inbound

HyperLens: Quantifying Cognitive Effort in LLMs with Fine-grained Confidence Trajectory cites this paper.

HyperLens: Quantifying Cognitive Effort in LLMs with Fine-grained Confidence Trajectory Massive Activations in Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T11:38:49.630171Z digest=sha256:202bc2d75c52a3b4ab468f15b1772ffaa1daae3a54fa0054d499810615299a4e

Observation 3cfaea57-1004-4290-98c1-8905cd8367e0 · inbound

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models cites this paper.

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Massive Activations in Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:29:20.796512Z digest=sha256:e34dd5bbe066bd8bdb7fde9cc2dcc1afa5c656802e6cf880a9128db1fecc4d35

Observation ae2b1b10-7ccd-43af-b80f-0655ff6f4e04 · inbound

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models cites this paper.

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models Massive Activations in Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T21:03:25.624300Z digest=sha256:034e66c4d66272e62ac41d631a9eabf6412a5a2f999c8e73ec4569e39c6870c3

Observation 922095c1-efa6-420a-a1c3-694242266ca4 · inbound

Attention Sinks in Diffusion Transformers: A Causal Analysis cites this paper.

Attention Sinks in Diffusion Transformers: A Causal Analysis Massive Activations in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:18:26.883899Z digest=sha256:90d7ebaf012bbb8fdf5e6fb4f164d76605b64796a6ffad33c19894699ee174c5

Observation cb3946b7-ea77-4c40-9bbb-e6d4d25819e1 · inbound

Attention Sinks in Diffusion Transformers: A Causal Analysis cites this paper.

Attention Sinks in Diffusion Transformers: A Causal Analysis Massive Activations in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T05:59:16.233848Z digest=sha256:41ce088b24eb5688422c2ef969348c3213a2273e0282bfabf7399c5892e99129

Observation ceba6819-4be6-47be-b841-f47b7fbf3bcf · inbound

Attention Sinks in Diffusion Transformers: A Causal Analysis cites this paper.

Attention Sinks in Diffusion Transformers: A Causal Analysis Massive Activations in Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:45:46.094318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T22:47:11.360530Z digest=sha256:aa11a918fdaca0e2823175adb5a5c3e7f8f8b1c9455bce5f08dfbc1fcccc3326

Observation ea74b8f5-9e0e-4191-8a50-34700be059a3 · inbound

Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination cites this paper.

Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination Massive Activations in Large Language Models

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:02:54.298502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T05:15:34.156717Z digest=sha256:aee00ddde77112ef5163ffd27df53db99284c3a4b1e1f9873b172b7eb681404c

Observation 3c9d3649-708d-45a2-b29f-a57bcd5c25e4 · inbound

Registers Matter for Pixel-Space Diffusion Transformers cites this paper.

Registers Matter for Pixel-Space Diffusion Transformers Massive Activations in Large Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:08:54.075715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T19:08:23.052621Z digest=sha256:f2ec1cb651fb56723cdf372f3d3e1d655ad038b390dc0be04a15441448516bb9

Observation 80abfa05-8b46-4a34-94b2-b50624b91b6d · inbound

Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise cites this paper.

Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise Massive Activations in Large Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:49:05.404901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T21:44:58.434589Z digest=sha256:e5ac26c8b49c3d82d9f023c67643be628f4e7bcf93c5b16272e727aeb820f0bb

Observation e08a1f29-2709-4f9f-bf97-79d48ac48c83 · inbound

A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions cites this paper.

A Two-Parameter Weibull Framework for Diagnosing Transformer Weight Distributions Massive Activations in Large Language Models

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:23:21.552672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T14:20:15.499338Z digest=sha256:35680e9a2bf110da8a683253feff1d9e66941d911c4f74dc14f88fcf899b1e7e

Observation f08fd51c-3415-4236-a430-80a793d5822d · inbound

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register cites this paper.

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register Massive Activations in Large Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:48:05.658978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T06:48:00.733211Z digest=sha256:759f8f2e60aecad1305d91dfdb3887d653bea28bc45525fe5c6d6d1f68adaa3c

Observation 394692bc-6410-4042-a7aa-acc161029fe6 · inbound

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond cites this paper.

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Massive Activations in Large Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:58:07.583961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T07:57:51.032025Z digest=sha256:fecb1b7744a9bcf5675cd9ab826aeba298dba1dca80185947734031ee66b09e6

Observation 44d160ba-f95a-4d6e-bfb4-cbcc6e847c31 · inbound

Steered Generation via Gradient-Based Optimization on Sparse Query Features cites this paper.

Steered Generation via Gradient-Based Optimization on Sparse Query Features Massive Activations in Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:36:40.308302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T05:31:29.510639Z digest=sha256:94898f5b921a242221e205355fab712e33ed4bfbe14b0f2f3418e10918822e14

Observation 27532da8-eaba-4e7b-b0c6-7be961a8d607 · inbound

A Simple Plug-in for Improving Eviction-Based KV Cache Compression cites this paper.

A Simple Plug-in for Improving Eviction-Based KV Cache Compression Massive Activations in Large Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-25T04:55:24.127620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T04:51:04.068354Z digest=sha256:1279ebf23ee9c9a8c61cb37b8fb54c99ec73ce071fdce8d23ff29049a0f0832c

Observation 5a9ecc30-18ea-4fd4-93e1-ef7e3d804773 · inbound

Multi-Gate Residuals cites this paper.

Multi-Gate Residuals Massive Activations in Large Language Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T04:50:20.372266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-25T04:48:06.164938Z digest=sha256:1f3617b7a5297ebe77792543d3cfb49d8bf57eb5a280dc72128a3969c45da367

Observation 09f478f2-335a-4ce6-80da-eb1f509e88f4 · inbound

YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models cites this paper.

YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models Massive Activations in Large Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:22:46.925024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T23:19:18.264462Z digest=sha256:b37743f9011bcceca57a388f2813b34942a9a4d33903d1aa4bcfb2ca63368a47

Observation 18d3ee7d-4d24-41b5-af02-b14fbeb4c824 · inbound

Rethinking the Role of Tensor Decompositions in Post-Training LLM Compression cites this paper.

Rethinking the Role of Tensor Decompositions in Post-Training LLM Compression Massive Activations in Large Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T01:46:26.869250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T11:31:25.851340Z digest=sha256:ba7632d0cac20817e220e6a86308bcd37c6525cc0f25b006e8b9611132b13c59

Observation 94ed6142-bd48-4a66-8ca7-1bb3d67694f2 · inbound

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models cites this paper.

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models Massive Activations in Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:56:27.444608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T11:27:09.779776Z digest=sha256:9f073fffbca717b4c36a88c281f2103b67cc1c9b73e697986c0ceed13950c478

Observation 6a09064c-2530-4d0f-ae0e-1305d741deac · inbound

Dominant-Layer ZO: A Single Layer Dominates Zeroth-Order Fine-Tuning of LLMs cites this paper.

Dominant-Layer ZO: A Single Layer Dominates Zeroth-Order Fine-Tuning of LLMs Massive Activations in Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:46.987673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T06:35:46.175065Z digest=sha256:b7e985c6304162c82a8d706e4d702e5bba39565b4841a14574fd72b144c4d7d5

Observation ab8c87e8-10b9-4d2b-a08b-368bcd703a83 · inbound

Dead Directions: Geometric Singular Learning cites this paper.

Dead Directions: Geometric Singular Learning Massive Activations in Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:36:55.209040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T03:20:09.365073Z digest=sha256:3953c16f5e9753372b7bdea45b2c72b191775927fc8d5c4f7d540d3c669d299f

Observation c3a4e679-9f32-4593-819f-9832b2a8718b · inbound

P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8 cites this paper.

P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8 Massive Activations in Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T05:56:40.586160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T07:55:23.257970Z digest=sha256:fdfdebbbef4757b57b9617482b23b5f44334ba50828cc288444ea2851bea3656

Observation 02f250f4-fe71-41c5-8299-34fea20461bb · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Massive Activations in Large Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:32:46.731937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:f23123092da18f5dcab6ab984a73807a4eb590bda72fbd75625ab8a9fa451d24

Observation 51a33f4e-b687-467c-8c08-3ba3eebdd8ca · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Massive Activations in Large Language Models

Reference 172

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:32:47.298617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:02d3067f0c8979e1832533b9247ab3b508a0d2bd80690c1d899777ab2de23b37

Observation 5d824933-ff7a-4041-842b-d0eaf2b5bec4 · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs Massive Activations in Large Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:57:32.370082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:b230fbf6d7821ea944c61c6ee04b826a1f6a5d79b86a4572b8a4d7d5511e68d7

Observation 172cd7cd-07e9-484b-8f53-863739c41580 · inbound

ICA Lens: Interpreting Language Models Without Training Another Dictionary cites this paper.

ICA Lens: Interpreting Language Models Without Training Another Dictionary Massive Activations in Large Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:17:49.067937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T10:21:58.878499Z digest=sha256:29814253e5020745bf7d920bb13c9b8d9fc6a061b8f10db63b272ed310bdd92e

Observation 2e7efc17-490e-4eca-a6db-881506a87601 · inbound

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models cites this paper.

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models Massive Activations in Large Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:28:04.157411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T09:35:24.118536Z digest=sha256:ada415a08fd5d7fe05e210289ef3918b8c26bbbf8a42ea474fd9c3acb1677a14

Observation 762dfa80-5f30-456e-a465-a0ac29fb917e · inbound

DynamicPTQ: Mitigating Activation Quantization Collapse via Residual-Stream Dynamics cites this paper.

DynamicPTQ: Mitigating Activation Quantization Collapse via Residual-Stream Dynamics Massive Activations in Large Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:57:47.899068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T10:36:33.893923Z digest=sha256:eb973eea99496a676401eb3039853275e25848891ac03e7b5d109b7f1e359938

Observation a4d8e077-dc61-44be-a07e-303fd8770aca · inbound

MiniMax Sparse Attention cites this paper.

MiniMax Sparse Attention Massive Activations in Large Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T15:28:34.115353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T06:29:46.190303Z digest=sha256:518381b66025d00a8b55751ad2e7f09f5db2eebf1202055707810f46f122c94c

Observation f5782f62-4696-4b0c-b7ce-bd0696b564f7 · inbound

Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications cites this paper.

Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications Massive Activations in Large Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:20.670950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T20:47:36.337189Z digest=sha256:f53f2631080a7a5776cda0aaed7e9dfa72d415cb37637026bc3373b81af45787

Observation d2115d05-ec7e-42a5-b466-24f950eb2ad8 · inbound

Algebraic Dead Directions in LayerNorm Transformers: A Forward-Pass-Only Diagnostic at LLM Scale cites this paper.

Algebraic Dead Directions in LayerNorm Transformers: A Forward-Pass-Only Diagnostic at LLM Scale Massive Activations in Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.257781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T21:14:11.815522Z digest=sha256:d71a4e76a3707b23f6209ba57dae0708d79311a5edaf5ea5b8b95f7839b4eb33

Observation 16a81b1e-96f9-4708-948f-ae046188af58 · inbound

Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test cites this paper.

Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test Massive Activations in Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:59:20.553890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T20:47:54.133087Z digest=sha256:d4358391e5c8d62a9a11d4f2d3b76874f5dd156e1a69bbf6b7676a04d36b16ef

Observation cb690365-a9ba-4e4a-bfc6-16a759ba42a4 · inbound

Demystifying Numerical Instability in LLM Inference: Achieving Reproducible Inference for Mission-Critical Tasks with HEAL cites this paper.

Demystifying Numerical Instability in LLM Inference: Achieving Reproducible Inference for Mission-Critical Tasks with HEAL Massive Activations in Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:59:38.000608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T14:58:55.680525Z digest=sha256:dc06fda6734876ce1dc8934bffce5deb4c2d4c1a4ae2a29502084e4ae303a795

Observation 46bd5bef-62cc-4abc-848b-9ef34ed5d7e1 · inbound

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference cites this paper.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference Massive Activations in Large Language Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T12:59:52.345822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T05:41:39.052865Z digest=sha256:3a09d422754febeaf28ae5ca9ea9e690515e606ae89d773c3d123c1cb2f81d1e

Observation b028e290-b7d9-4a7b-bcc9-987a888f41fd · inbound

Information-Regularized Attention for Visual-Centric Reasoning cites this paper.

Information-Regularized Attention for Visual-Centric Reasoning Massive Activations in Large Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:17:07.253514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T15:12:43.475802Z digest=sha256:9e1e4f2f6f79b6fd2cca8d2fa7a1e9edf60dcf52e37614219e20c048befbd842

Observation 4d21c1c0-3162-443d-9e09-7db8d820e31f · inbound

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation cites this paper.

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation Massive Activations in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T05:46:45.293902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:46:45.293902Z digest=sha256:95af551e3a8ecaca781ede67de7a60e94455447a65576e08a7a188e4985eb653

Observation e5ab7a8d-73d7-4e49-a1f5-600ae34915e1 · inbound

Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale cites this paper.

Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale Massive Activations in Large Language Models

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T18:27:31.317526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-10T18:22:48.495407Z digest=sha256:8e8b6641834663656cc84f6472fba10794bf719b4d4c68f2bbe6aa71d9979327

Observation 70a0c0dc-55c0-4f0d-b0c4-7fe95e173d1c · inbound

Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator cites this paper.

Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator Massive Activations in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T04:11:56.715045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:11:56.715045Z digest=sha256:9c4e5381827fc747aaff4e18c01bc85e3cdcf304c82c63b819b55289590c0646

Observation 1c206bbc-5ba5-4cdc-b4de-c344a95c3fc6 · inbound

Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization cites this paper.

Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization Massive Activations in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T14:15:46.074413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:15:46.074413Z digest=sha256:7d427a99f8a4dfd819d8306cc96572dfb72d88df99ff7060934995beb993e122

Observation dbaf008b-e25d-4a52-ae7d-8becdb250a27 · inbound

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering cites this paper.

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering Massive Activations in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T04:52:40.524880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:52:40.524880Z digest=sha256:837f97efe352cc8e8a506a8dff1f2734b12f23c6d134bee21170e49b0ae340b7

Observation d6794ba7-bd94-4812-94aa-56d7ae62c82a · inbound

Phasor Attention: Mean Root Square Normalization for Phase Manifold Preservation cites this paper.

Phasor Attention: Mean Root Square Normalization for Phase Manifold Preservation Massive Activations in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:39:38.536822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:39:38.536822Z digest=sha256:f4d298f2e6da17e3fb174216bae025fb8dcf1ff0fe4510710dd49baed4fb624a

Observation 4a443176-c429-41cb-85b5-64995baa358c · inbound

Estimating Rare Events in Language Models with Proper Evaluation cites this paper.

Estimating Rare Events in Language Models with Proper Evaluation Massive Activations in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T15:26:28.124965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:26:28.124965Z digest=sha256:8479d8a4ee239af37cd661e25edcaaea551b33efac9d8be073a5eb366aa77399

Observation f39f12d3-f021-4699-939d-3efa67c21e5e · inbound

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers cites this paper.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Massive Activations in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T13:25:57.390684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:25:57.390684Z digest=sha256:6b4f6a4298f051b9fa084db9e8b28161607c2a52f525f5ec9972ab28381e33f0

Observation 6459f884-0819-411f-a870-8fb057a89af0 · inbound

Reference Feature Atlases for Mechanistic Auditing of Language Models cites this paper.

Reference Feature Atlases for Mechanistic Auditing of Language Models Massive Activations in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:17.430243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:35:17.430243Z digest=sha256:db2f24001ca4b282d989aba2d815394badc7a3d3cc67b10cbcadd96a16992e8a

Observation 4374f278-588b-48ab-808a-386da25b46e2 · inbound

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization cites this paper.

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization Massive Activations in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T00:43:52.070083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:43:52.070083Z digest=sha256:294b50461227cf915d05f7cc524aca5ddfdfd90626b2805442cd4877cc4fe824

Observation f8ff366e-1770-45e8-ac9e-fe0d045b1de9 · inbound

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs cites this paper.

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs Massive Activations in Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:44.582319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:44.582319Z digest=sha256:3060539ee0cb0ca1f3b685294c0302fc5bdbc4f918ddb5d11ce6f3482291102e

Observation 8096263e-d7c3-4f93-8069-42b2d96aabc1 · inbound

Hidden Language Consistency Phenomena in Reasoning LLMs cites this paper.

Hidden Language Consistency Phenomena in Reasoning LLMs Massive Activations in Large Language Models

Reference 210

Resolution
unresolved
no resolver link, observed 2026-08-14T04:40:10.034226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:40:10.034226Z digest=sha256:83aee0d68e047691a1f72a68641e746daf0ec5caabff0624ec84e92315b8292c