Pith. sign in

Paper Citation Record · LEDGER

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

As of 13 August 2026, this Paper Citation Record lists 100 of 199 outbound references and 4 inbound Pith citation observations for arXiv:2604.10098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10098 v1

Coverage vector

measured 100 of 199 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:17:09.834609Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T19:08:23.052621Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T19:08:54.035553Z

Reference resolution

100 of 199 outbound references displayed

  • verified exact39
  • verified fuzzy60
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa24d1ed-4df9-4616-87d0-923a5702cda2 · outbound

This paper cites Attentionisallyouneed.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Attentionisallyouneed

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.177831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:3a1407b6684461e9ed55f437745187a165cae34fa2fdb89f94ff6f4e3e21c608

Observation ee950246-922b-4ce5-b1d6-72867d807dbf · outbound

This paper cites A survey of transformers.AI open, 3:111–132.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation A survey of transformers.AI open, 3:111–132

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.174727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:b79b1c8318506e765d2e5970f4652708bf57b76e733e76d6e5a445ebbb9fa82e

Observation cac07e63-6615-443a-843c-fcecb91c42ea · outbound

This paper cites A Survey of Large Language Models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation A Survey of Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:57.971328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:f3ea0716eacbfc5b602780ae7b0ea3b73eb19502b549ed452173dbf0a24cbdf3

Observation e55857f5-80ff-4e3d-b1c2-7de3f66bbb0d · outbound

This paper cites A survey on vision transformer.IEEE transactions on pattern analysis and machine intelligence, 45(1):87–110.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation A survey on vision transformer.IEEE transactions on pattern analysis and machine intelligence, 45(1):87–110

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.376498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:b309318b471389191f307fc15c37753e02a9b090cd940a6f52c98eb865005577

Observation c6df39ab-eede-4575-9d58-4364ab95c7f9 · outbound

This paper cites A survey on multimodal large language models.National Science Review, 11(12):nwae403.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation A survey on multimodal large language models.National Science Review, 11(12):nwae403

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.233501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:234c53cbf7f7836d32ac62f8b148d8a20e4a55138dc6802114547486f7f1ad0e

Observation cbfe3669-25c5-4ee4-8745-06a0566685d1 · outbound

This paper cites A survey of embodied ai: From simulators to research tasks.IEEE Transactions on Emerging Topics in Computational Intelligence, 6(2):230–244.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation A survey of embodied ai: From simulators to research tasks.IEEE Transactions on Emerging Topics in Computational Intelligence, 6(2):230–244

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.195014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:19195babeb2d111a60fd618366bf24bc592ee867d548e98c39571941fb950407

Observation b88c1003-3aee-431e-9f69-9ec7cf38f5ed · outbound

This paper cites Longcat-flash technical report.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Longcat-flash technical report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.303917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:5a9b17c62272818738996a3743a132251995dd10f0159d258088d550ec5a283d

Observation 4d2bcc5e-c999-409d-a858-248bc09478c9 · outbound

This paper cites Longcat-flash-omni technical report.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Longcat-flash-omni technical report

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.048016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:5c03bdeab2b18899382683190bc3ff345577f049efb87991f9ed00275e73bb19

Observation e76e36bd-6df1-439d-b30f-ab8df2a5736b · outbound

This paper cites Introducing LongCat-flash-thinking: A technical report.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Introducing LongCat-flash-thinking: A technical report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.275862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:0666d8d8bdefe940e6023ea1ddd26ab3620b158dfe7143b2909d904d883713f1

Observation 4020cc53-582e-4bc0-bb4a-2caf4c0fa977 · outbound

This paper cites Longcat-flash-thinking-2601 technical report.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Longcat-flash-thinking-2601 technical report

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.153487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:6c24a4404fcaaa109668bb51b95eb94eb7ad45745c40cf255a95aa563a567d30

Observation dcf4fd8c-f2b9-4198-bd7b-0bf4200a0326 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Vggt: Visual geometry grounded transformer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.167958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:dc7ed741a171d51626903924384d805bcfcb14226acff1d85823f0cd9a096bb0

Observation d8e7aaf8-2c98-42fe-9b6b-924dbb82e4e3 · outbound

This paper cites Xstreamvggt: Extremely memory-efficient streaming vision geometry grounded transformer with kv cache compression.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Xstreamvggt: Extremely memory-efficient streaming vision geometry grounded transformer with kv cache compression

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.055150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:f2934ff35405f2f6f16722a1e8f231b7a27a609851506a2a1e3148e627d77eea

Observation 32819257-1579-4fa8-89d7-f924e2c465ef · outbound

This paper cites MapAnything: Universal Feed-Forward Metric 3D Reconstruction.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation MapAnything: Universal Feed-Forward Metric 3D Reconstruction

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:06:15.374425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:38fa39bf4354d078e904a74a31cd7164840ce213e82cb46745a7309124d0ba86

Observation 32f393c0-5e39-4dad-b06e-12f81a13b2eb · outbound

This paper cites A survey on model compression for large language models.Transactions of the Association for Computational Linguistics, 12:1556–1577.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation A survey on model compression for large language models.Transactions of the Association for Computational Linguistics, 12:1556–1577

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.181587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:efd84b4548b2c627df2f4558275e1d8e4e976c7da79143a374a3dd66fa5d7239

Observation 168935be-15c8-480e-bfcb-529ffdd705f6 · outbound

This paper cites Efficient large language models: A survey.Transactions on Machine Learning Research.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Efficient large language models: A survey.Transactions on Machine Learning Research

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.191707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:94ef1d892ed03b5db58fcdeeed962e877916504de2126d2463b8f33544267ffa

Observation d84989f8-3c51-41bb-9b67-bf363b2c12e7 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Kvquant: Towards 10 million context length llm inference with kv cache quantization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.478668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:ff9ed1b683466519a9b15643a71d39eee0e385f0a32130f1dc4888e29a8b6457

Observation b8481f70-56b4-4759-a618-a19ecdabd51a · outbound

This paper cites Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.188305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:49211bb7608e5529973d86b0cf9345a22c0fcbf67f8512bf8b604a066f81150f

Observation aa58d049-a37f-4072-9b53-63927d93ef9c · outbound

This paper cites Efficient attention mechanisms for large language models: A survey.Visual Intelligence.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Efficient attention mechanisms for large language models: A survey.Visual Intelligence

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.576481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:3289c16822c2f285ec7300c61cf0e3dc9df9a267eaff55ce98ab1222f7286963

Observation e44bbf8d-916b-4222-9cb8-1119b1d06ad6 · outbound

This paper cites Speed Always Wins: A Survey on Efficient Architectures for Large Language Models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Speed Always Wins: A Survey on Efficient Architectures for Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.088861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:0ef1f44e805b1cb434d89e39c65302181ab1f4e3b1abda68570bca5f218fefe8

Observation f81ad7ad-7f82-475f-93e8-0f3a5f87a551 · outbound

This paper cites Gated delta networks: Improving mamba2 with delta rule.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Gated delta networks: Improving mamba2 with delta rule

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.360005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:4eb15c8c6b8cba9964e1935bbc08d0b749c28d6ddce969aed4c45ee5a908a84b

Observation de6dd072-f4d6-4212-b67c-d5e29b8340ac · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:11.451207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:72691d21f3d2e74f491f043f29b2d950476856475a97a9e09ece0005a3b1fa4b

Observation 39064cea-29aa-4a0c-916c-d3c497d5bbcb · outbound

This paper cites Titans: Learning to memorize at test time.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Titans: Learning to memorize at test time

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.335021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:f099e31200b48f4a78cc02b731769d99644ec45401bc1de653bc9c3d5c4321b9

Observation 1f507c2f-d205-4357-bc1c-f7dd9b5705f1 · outbound

This paper cites SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.161185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:ced8e118a7c7ee9371e893f5ff9aaa3eb2c375a2f0864d74d1586af18a297590

Observation 393f78f2-11e0-41ad-bb1d-5b6641c983c1 · outbound

This paper cites Efficientstreaminglanguage models with attention sinks.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Efficientstreaminglanguage models with attention sinks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.241116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:937c0e75d25cb30bc949ff5b2f0f11c7ba204a8c7ca77e10f9353349dae56954

Observation 1efefa74-c70c-405c-ad0e-5de673143071 · outbound

This paper cites When attention sink emerges in language models: An empirical view.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation When attention sink emerges in language models: An empirical view

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.257031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:22e4645d60f8a184523d3349e80bf6f60b5f8242637dab1c1fe9df6ea1b72478

Observation 6760d8d1-b4b8-4e22-83cd-d258739857b9 · outbound

This paper cites Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.252877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:698408007ec02108bcb955a0870358eed50341407274c4ee8fa78ac1ab7ddb0c

Observation db85d0cf-59e9-4f09-946c-4e1f39b1bb09 · outbound

This paper cites Why do llms attend to the first token? InSecond Conference on Language Modeling.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Why do llms attend to the first token? InSecond Conference on Language Modeling

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.214422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:8f014fe13eeddb75ddc54448b0a0784904bcb2fdafdfb259e87485e2f3a04c83

Observation a5a07e3f-2cfc-4b2c-8199-fe04345f26d2 · outbound

This paper cites Kvsink: Understanding and enhancing the preservation of attention sinks in kv cache quantization for llms.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Kvsink: Understanding and enhancing the preservation of attention sinks in kv cache quantization for llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.299064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:c9f56222b9f0f78c927a9985bbe0149da7094fd904cef3173e06fadc9dfeb34a

Observation 3df883b1-4497-4f1a-bca7-dfb63386f7d5 · outbound

This paper cites Quantizable transformers: Removing outliers by helping attention heads do nothing.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Quantizable transformers: Removing outliers by helping attention heads do nothing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.396459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:e0026ce3701a7db90685478b326e9dacdded9f7ea9f77af0c6047ab8936d1636

Observation a50651e7-c8f5-4354-8398-d541d342ac69 · outbound

This paper cites Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.172892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:b7d137ad17d091f6a1ff8bd9df8c0027bde54a633f960160dc2d384b8b59ed28

Observation 65866687-b36b-4dd0-8b12-8914262073ad · outbound

This paper cites Attention reallocation: Towards zero-cost and controllable hallucination mitigation of mllms.International Journal of Computer Vision, 134(1):22.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Attention reallocation: Towards zero-cost and controllable hallucination mitigation of mllms.International Journal of Computer Vision, 134(1):22

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.443368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:9043149ed9ab9ad9f34687876a1f9a376fd763110f846317bb2eaeb1e036ba60

Observation bfd8a969-39f3-4a61-a07e-c33628358d1a · outbound

This paper cites Vasparse: Towards efficient visual hallucination mitigation via visual-aware token sparsification.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Vasparse: Towards efficient visual hallucination mitigation via visual-aware token sparsification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.568874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:03abfb7a50ca7d8f34bd48a2d738034b212a3606224268af48e55dac9ab05b7b

Observation e2987534-74e4-409c-8891-857451b67892 · outbound

This paper cites Forget- ting to forget: Attention sink as a gateway for backdoor- ing llm unlearning.arXiv preprint arXiv:2510.17021.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Forget- ting to forget: Attention sink as a gateway for backdoor- ing llm unlearning.arXiv preprint arXiv:2510.17021

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.194894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:2c9a4a7e9eb9eaf0dd4681098317549af22f337d0863c87fa9d4d7acdefbfcb6

Observation f6a1d808-9c04-4d7d-87d2-01eb8a530128 · outbound

This paper cites Interpreting the repeated token phenomenon in large language models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Interpreting the repeated token phenomenon in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.338513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:f916ad060892ec08ec24dc971a29e9c31b6df5b4f8040b625f91151d672bec7d

Observation e045dc42-cd84-4e90-b98e-5196a0f5e5ec · outbound

This paper cites Leveraging registers in vision trans- formers for robust adaptation.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Leveraging registers in vision trans- formers for robust adaptation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.218447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:a0bff73584ff9e3c40070b4a50b5a8d7b30fefb3ec5c134b84a7c3faea9ca67e

Observation f834c925-b530-43f9-9a58-04361dd44fce · outbound

This paper cites Nosa: Native and offloadable sparse attention.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Nosa: Native and offloadable sparse attention

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.272183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:6e43a72122054bf7820affe6e33896da8026c30d8037d4fad488a8f022473e62

Observation e446f852-d717-4982-bde7-a17c29e5d233 · outbound

This paper cites OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:02:20.220338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:4b7386643e1dc82a0467e52915934533460f3f2d932dfb9b56a2f2099840be38

Observation 9f4a8eca-af8a-46d6-a13f-a952e0fd9773 · outbound

This paper cites SALS: Sparse attention in latent space for KV cache compression.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation SALS: Sparse attention in latent space for KV cache compression

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.341908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:2f0cb2273cf978f68018a957c870bbc6bde276ff78de0374077ad6cf5c205e61

Observation 2137741a-cb04-4054-9a6c-7edd591f158b · outbound

This paper cites OjaKV: Context-Aware Online Low-Rank KV Cache Compression.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation OjaKV: Context-Aware Online Low-Rank KV Cache Compression

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:57.862495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:506af1be4708d23190887ea6d84ad8fdea0e2824e056d7f7f48005f2e8b622a1

Observation f93cbd5e-34d1-4dee-9543-fbd0992eb31e · outbound

This paper cites Mitigating attention sinks and massive activations in audio-visual speech recognition with llms.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Mitigating attention sinks and massive activations in audio-visual speech recognition with llms

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.141700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:ce47fc0670af8b8abc868412cbd9b452120f5d17f70cfa3bde45a9128e68a34a

Observation 771e016f-74eb-4e44-8c0f-fd8cd0748642 · outbound

This paper cites Attention sinks and compression valleys in llms are two sides of the same coin.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Attention sinks and compression valleys in llms are two sides of the same coin

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.486196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:66c30ed8f386a8bb174e986fbce124221695784debc13ddf2a90282f316eb5d6

Observation 98a8b51d-41ea-4fb7-972b-0e6d7e477954 · outbound

This paper cites Outlier-safe pre-training for robust 4-bit quantization of large language models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Outlier-safe pre-training for robust 4-bit quantization of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.496430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:17e50eb7d0965076015a965995b253e246503cc844cb5f084585e8f130e6e433

Observation 768890e9-e056-4db0-bb79-bcb2b7f7d8b9 · outbound

This paper cites Unveiling super experts in mixture-of-experts large language models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Unveiling super experts in mixture-of-experts large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.506535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:d018eb7b3734a41451076fbee87e674070146cbe45020fd87db6b68cbd543b04

Observation ef20d4e6-ce66-462f-ae9d-c9befa908093 · outbound

This paper cites Value-state gated attention for mitigating extreme-token phenomena in transformers.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Value-state gated attention for mitigating extreme-token phenomena in transformers

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.839886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:26b1258a18615462459e9af08a5648a5e685dcb092a8a68add964fc44937ce2f

Observation 71fe4d71-65e8-45b8-a10f-4c853a3c44d1 · outbound

This paper cites Research and latest advancements.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Research and latest advancements

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.200975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:77a9e866ca75fd3b5f1ff8e1eb207c1378c533bb2b9f0f8442d49a16d147f174

Observation d68d2ff5-7e28-4169-a9ba-e57d6582fa6b · outbound

This paper cites What are you sinking? a geometric approach on attention sink.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation What are you sinking? a geometric approach on attention sink

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.554852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:036566ccd9214e7bf17c1cfa627c4ea5fee2666ce702d2dc7eb1c26f1ee6cddc

Observation 114ccecc-d602-430e-ac76-ba5dc9bf27a2 · outbound

This paper cites CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-03T02:05:41.714027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:215bf88f187ff5435de875125860c6a27db52655628c6d11a7b1a9b7f528e197

Observation 1bdedb93-1215-4d49-b290-1109a8048d31 · outbound

This paper cites Does roberta perform better than bert in continual learning: An attention sink perspective.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Does roberta perform better than bert in continual learning: An attention sink perspective

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.409523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:c91ba80eef44f224686561a0e10f783fee35f92bfd93540b473760ae67e301e9

Observation ea2d9ab0-7a6b-4407-a97e-11bea1f4a85a · outbound

This paper cites Outlier dimensions that disrupttransformersaredrivenbyfrequency.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Outlier dimensions that disrupttransformersaredrivenbyfrequency

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.447264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:9e02b26902c938b0a45e0a2757a213fed38c7c3e62a2a010b57815abebecf40b

Observation 2f179ad1-0e5e-4064-96d6-24c8ead91245 · outbound

This paper cites Understanding and overcoming the challenges of efficient transformer quantization.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Understanding and overcoming the challenges of efficient transformer quantization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.380628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:7b8b2e2526f212964fb1b5003b18a152c9520a2f1d37d200dd71a9eb0e6422c2

Observation 246e340d-2548-400b-8953-3ecb2b5c7c8c · outbound

This paper cites Bert busters: Outlier dimensions that disrupt transformers.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Bert busters: Outlier dimensions that disrupt transformers

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.413020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:cd94f134ecfb59466a1fcbeb93fa800ab8a64cc6fd44ae0916b6e20bf2610656

Observation 5a3fc787-8962-4e36-b383-f53c3f750ef3 · outbound

This paper cites Positional artefacts propagate through masked language model embeddings.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Positional artefacts propagate through masked language model embeddings

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.557840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:c0f9b8882b662292e470f2c1220c4feab18df4a572b6c414de2a5d945ed4d3c4

Observation cf02788f-8034-429b-b9f5-4aa2a28ca8d7 · outbound

This paper cites What does bert look at? an analysis of bert’s attention.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation What does bert look at? an analysis of bert’s attention

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.403131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:d7cd462f694d1f88b69db63ea8c4a6d68be573c83666984dc27d76d83ba4211b

Observation 960f5649-5ff8-4d3e-8efa-a60693af1ec3 · outbound

This paper cites MiMo-V2-Flash Technical Report.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation MiMo-V2-Flash Technical Report

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:33:32.947936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:ae4b7a8ab8f941dbceca1562f4dce21a7f909aaef7a140351cd8db0bcaec4e4c

Observation 2db3180d-4542-4554-bdf8-9c6aea642ede · outbound

This paper cites Attention needs to focus: A unified perspective on attention allocation.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Attention needs to focus: A unified perspective on attention allocation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.925228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:1509335064b07bd532457d015b22afeb5e4061e2d7ce8b96b46c6689fdfb7f5e

Observation be95aad9-6650-4018-be11-c998d820fb7e · outbound

This paper cites On the existence and behaviour of secondary attention sinks.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation On the existence and behaviour of secondary attention sinks

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.834936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:56f36faaa60f28be776b2a3f82fde0b525952ed3ab6bffc2616605273b119384

Observation 1854e6bb-ec19-4971-b475-4d91bd360291 · outbound

This paper cites Sliding window attention adaptation.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Sliding window attention adaptation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.899852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:11fbf9c4f8a839519e856ca7ed608c889ea365b4794846009f475247eee93684

Observation f3c2dcbf-fbba-4479-8225-4d5f44fb2955 · outbound

This paper cites Dope: Denoising rotary position embedding.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Dope: Denoising rotary position embedding

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.932730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:9120c8eb38a9ea7a139865fd63622026f9f21710bba5094341d71922339ac43c

Observation f37dca4a-b180-431c-8a22-8f35b8cf15ee · outbound

This paper cites Tweo: Transformers without extreme outliers enables fp8 training and quantization for dummies.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Tweo: Transformers without extreme outliers enables fp8 training and quantization for dummies

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.888361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:066a9e8ae00086d2d0b7c6ad9b03dec214ace6903575fe90880e9fb6b7b81147

Observation d31aa59b-3ce7-434b-a2db-31062d808ae7 · outbound

This paper cites Lost in the middle: An emergent property from information retrieval demands in LLMs.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Lost in the middle: An emergent property from information retrieval demands in LLMs

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.983080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:f3d297d5dcbc5e2098aa8896bbd96d55dd0c8eec6922d8d2dce113738ff84940

Observation 3f1037da-5a52-4862-bdac-73ab5e75d78f · outbound

This paper cites Hybrid Architectures for Language Models: Systematic Analysis and Design Insights.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Hybrid Architectures for Language Models: Systematic Analysis and Design Insights

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.030354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:8b17ec555f3e42c4da861fc1fe403ea9afd3f50a619b9b7ea33e406b3622635f

Observation 50321c96-3347-433c-9926-0d01d8d28f19 · outbound

This paper cites CacheClip: Accelerating RAG with Effective KV Cache Reuse.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation CacheClip: Accelerating RAG with Effective KV Cache Reuse

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:50.604177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:c895d224306c0104dfbf30e7c96a0ca355034efae1325bb53a15ccf720c73c81

Observation 25c2da19-77a3-45dc-b3c5-5851a371073f · outbound

This paper cites Artificial hippocampus networks for efficient long-context modeling.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Artificial hippocampus networks for efficient long-context modeling

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.940007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:914936a2589ac6d6387aaf7b7a9cbf70da95dda0b7b4f5fec1579d382eb9570c

Observation 207ef988-1d6f-4bac-86f3-3fd01dc26dba · outbound

This paper cites vattention: Verified sparse attention via sampling.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation vattention: Verified sparse attention via sampling

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.237214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:4acb2153cf8e8d67da03210d62326891bc2b2beecca826af27088433876497f3

Observation 54bf7f7a-fb52-4402-b64b-bc7c801bcdac · outbound

This paper cites All for one: Llms solve mental math at the last token with information transferred from other tokens.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation All for one: Llms solve mental math at the last token with information transferred from other tokens

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.269537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:1807137ffda150043a9afc265b3a5db556106b4feee035f6da0667194b029a97

Observation 03d321b5-84f7-41f9-898f-29e2f3a152f9 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation gpt-oss-120b & gpt-oss-20b Model Card

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.295555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:d42b9c55f8fb6ff1c4dc13249f940a0740b476301d8857d4671cf81464ff0a5d

Observation ff7cee7d-36e9-4391-b84a-5505e1d91cb2 · outbound

This paper cites Integral transformer: Denoising attention, not too much not too little.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Integral transformer: Denoising attention, not too much not too little

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.302360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:4137c80e280fa26fb296cde33aec1d95df330efa4cd94e5f0585ebed600ae94a

Observation 336dfe7d-52d5-45d8-bbdb-ef3468de23b6 · outbound

This paper cites H2eal: Hybrid-bonding architecture with hybrid sparse attention for efficient long-context llm inference.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation H2eal: Hybrid-bonding architecture with hybrid sparse attention for efficient long-context llm inference

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.164002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:24b5bdf0f7b216dd14fbf58dcd735bb6e735900bb5abed03af021aed7d37166f

Observation 5a3aa25d-3f69-4882-8886-3b6d2f0c7957 · outbound

This paper cites Accelerating Prefilling via Decoding-time Contribution Sparsity.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Accelerating Prefilling via Decoding-time Contribution Sparsity

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.042049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:89e1dc79c6ba51047137a41a18a6e8e104cadb002c63403b6bd8cc44810db063

Observation 9ec1466d-6bf2-48ed-932c-5caa6f95b8d0 · outbound

This paper cites Orthorank: Token selection via sink token orthogonality for efficient llm inference.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Orthorank: Token selection via sink token orthogonality for efficient llm inference

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.320335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:41e465fca9abedf32007ef5a2ff2d51e2121df1fc31198860f878b35800a8f4a

Observation e2bde12f-8c3e-4e22-9bb3-ff06e7bdbfcd · outbound

This paper cites Earn: Efficient inference acceleration for llm-based generative recommendation by register tokens.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Earn: Efficient inference acceleration for llm-based generative recommendation by register tokens

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.349150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:67699dfbd1d9c576d72e2449a88b77b47f0a90a2097f4ff9b9a9d469c44be35c

Observation b9c372f8-f66e-4582-b2d8-0b483a5920f6 · outbound

This paper cites DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.881229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:205cd1664e681484345c662e3cd0893514b09fc6c306eb6358929637e48f6c90

Observation 795242bf-3d0f-49a8-96ad-9f827583f903 · outbound

This paper cites Two Heads Are Better than One: Simulating Large Transformers with Small Ones.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Two Heads Are Better than One: Simulating Large Transformers with Small Ones

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.232597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:6bfc91a29bd27abd3bb94e210a2b6df1130005487f27338d9893f717e79dfbee

Observation adc68189-911a-4ea2-bb71-b3f964fbeaf9 · outbound

This paper cites Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.988048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:1bbdc4c428b5218950690058e90fa73d7666f83539606415b65b1a751b697aa0

Observation 1cf09c47-39bd-4361-80bd-044d0ef29ddf · outbound

This paper cites Zerotuning: Unlocking the initial token’s power to enhance large language models without training.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Zerotuning: Unlocking the initial token’s power to enhance large language models without training

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.561092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:05ce676a7abab47fb782eadf52d755890249c29eb667fb95501abfb5f47aaecc

Observation 1095e0a4-1e1c-49cf-a782-17eca38f0219 · outbound

This paper cites Delta attention: Fast and accurate sparse attention inference by delta correction.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Delta attention: Fast and accurate sparse attention inference by delta correction

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.009864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:1d1cd2e59a656ff2d13efd4edc757488c9b0f82b8be5dc6258c172feb06f42d9

Observation 1639b654-9aab-48e4-bd4f-b3999b366418 · outbound

This paper cites Softpick: No Attention Sink, No Massive Activations with Rectified Softmax.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Softpick: No Attention Sink, No Massive Activations with Rectified Softmax

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:05:58.105206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:375b78a51dba73c05eb4e6e1424dd44200537f9cff829a35c5a7b4d929e72bfb

Observation 3ee72983-7d7e-4f35-8819-1d5a02246a15 · outbound

This paper cites Keydiff: Key similarity-based kv cache eviction for long-context llm inference in resource-constrained environments.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Keydiff: Key similarity-based kv cache eviction for long-context llm inference in resource-constrained environments

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.591502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:5edbe061eff976b7a6967aa42f7ab59fb6b9a01c6153bdc2b4a7f02d9957ac10

Observation b2b788cb-8714-4281-9d76-4ce1c7ff662d · outbound

This paper cites Edgeinfinite: A memory-efficient infinite-context transformer for edge devices.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Edgeinfinite: A memory-efficient infinite-context transformer for edge devices

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.539087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:6abeb96731fdbf16f208f4497b09cafb93bcfd4357f5a44a33acfed9077b4609

Observation da0721a4-2131-4411-a749-dce0a4205400 · outbound

This paper cites Efficient many-shot in-context learning with dynamic block-sparse attention.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Efficient many-shot in-context learning with dynamic block-sparse attention

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.516167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:b2b655560980c5e4825366c382c87618d0549a59488f03a82a1ea5d8d2802392

Observation 84ca17e2-b00a-40d1-adc7-abaa8a0af1db · outbound

This paper cites On the emergence of position bias in transformers.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation On the emergence of position bias in transformers

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.529286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:28d9667b35f1150cb5b2680ad453f07b06616e42cf0808c09dda3457898b8ad3

Observation 7f68eefa-308c-4eee-8c79-407d050c2206 · outbound

This paper cites Systematic outliers in large language models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Systematic outliers in large language models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.541810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:6527a1f5559310765fde47699457eafb3ea72cd6716a76051cbc3604d200a0bd

Observation 504b86fc-5c12-47e3-a090-8a0a6792d74f · outbound

This paper cites Sliding Window Attention Training for Efficient Large Language Models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Sliding Window Attention Training for Efficient Large Language Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.944778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:67465d8ee007244420a237c8ce29ee8a28ddfea20a1e6ab079eec8ae7ca50172

Observation 7897ecfb-2b3e-4b52-bc2a-5b13b7ea0df4 · outbound

This paper cites Rotatekv: Accurate and robust 2-bit kv cache quantization for llms via outlier-aware adaptive rotations.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Rotatekv: Accurate and robust 2-bit kv cache quantization for llms via outlier-aware adaptive rotations

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.588022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:4e9d13e4ba74af119ee82409f10c39f16bf1c42b82a567e3c8a87ee90e66ecd4

Observation fb8c00e9-eaeb-46e1-915f-b24ca2ef0ae7 · outbound

This paper cites Weight-based analysis of detokenization in language models: Understanding the first stage of inference without inference.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Weight-based analysis of detokenization in language models: Understanding the first stage of inference without inference

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.583256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:976bbca2627245e265847fc67aa570da63b801e1b619c5db31fadff34a50faa9

Observation 84eace44-4527-4779-9b34-87224001a7b8 · outbound

This paper cites Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.130976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:94876ea1d65719c8f0ccfc7f94961679ae2267b060456729b43be95057e55d5e

Observation 603475af-914f-474d-a411-ae9f945955aa · outbound

This paper cites LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:05:58.252809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:ca2d2b954f0e01953b655be137cd1e968989f346758b6475b919d75b417019d2

Observation 91ddf077-214f-423d-ad7b-374bb4a8388d · outbound

This paper cites Unigist: Towards general and hardware-aligned sequence-level long context compression.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Unigist: Towards general and hardware-aligned sequence-level long context compression

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.489633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:8f738acf7f1081fadddbc7a20bc8c985095218802499348f324183aa2923263b

Observation 27c70ef9-2217-497b-ad11-3b63e1763b40 · outbound

This paper cites Cache me if you must: Adaptive key-value quantization for large language models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Cache me if you must: Adaptive key-value quantization for large language models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.532382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:1c6313c5455779e429b25ec98548b568b2a2e58773f14102422ebfb5331f37c5

Observation 89d228a3-631e-47a7-842f-9e04d8bc6c14 · outbound

This paper cites Position bias mitigates position bias: Mitigate position bias through inter-position knowledge distillation.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Position bias mitigates position bias: Mitigate position bias through inter-position knowledge distillation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.466482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:fb3fc85b9d2546fab709ea7e866ba41f708c14753249925a77f01791029a5500

Observation 3bbb1062-47a9-45c6-8b6c-5a19fd9ed013 · outbound

This paper cites Attention sinks: A’catch, tag, release’mechanism for embeddings.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Attention sinks: A’catch, tag, release’mechanism for embeddings

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.352899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:48e7e0524478075084b105d3b219f6e1fdad8ff11a167bf454b509329111986d

Observation 0ccd0280-dc05-4452-90db-33226c650173 · outbound

This paper cites The singular anchor: First token dominance in large language model attention sinks.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation The singular anchor: First token dominance in large language model attention sinks

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.264328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:9a245b2802da6e6cd0f338fb86cd2bb3c0470d33149ed11fbf5c0b25f4786cb6

Observation 16d92e27-cecf-43db-ac4e-ec692a42a7c8 · outbound

This paper cites Attention entropy is a key factor: An analysis of parallel context encoding with full-attention-based pre-trained language models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Attention entropy is a key factor: An analysis of parallel context encoding with full-attention-based pre-trained language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.287632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:b7e3ba08d720237f176cf00b4d865636007b80cecf2c8064b1d6da9697203a3e

Observation 91e055e8-98f2-4bc1-b6cf-dbd9e1794167 · outbound

This paper cites Sgd-kv: Summarization guided kv cache compression.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Sgd-kv: Summarization guided kv cache compression

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.597540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:5f0ed4b21e6cc2a98cdfe9b9706ce750a1d3a687c7c2021adc95a23a0d206ad0

Observation e16d1cbe-8c4b-4ce2-ae7b-270cd756d960 · outbound

This paper cites Initial-key cache: An efficient kv cache strategy focusing on initial and key tokens for llms.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Initial-key cache: An efficient kv cache strategy focusing on initial and key tokens for llms

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.482687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:c7fb50074b85308986324e917a515182bad28593ab40227d1db048377481595b

Observation e660d061-3644-48d4-9a6c-d5993d303167 · outbound

This paper cites Entropy-guided kv caching for efficient llm inference.Mathematics, 13(15):2366.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Entropy-guided kv caching for efficient llm inference.Mathematics, 13(15):2366

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.434982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:451867b11fe9f50dc3f131e6bf9238f85f90e6f9807ce6804cd4eca3f526d631

Observation c5a3008c-c18e-4efb-b018-f2ed58490db3 · outbound

This paper cites Subkv: Quantizing long context kv cache for sub-billion parameter language models on edge devices.Software: Practice and Experience.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Subkv: Quantizing long context kv cache for sub-billion parameter language models on edge devices.Software: Practice and Experience

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.535584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:ff3d48e7dde11a5de85f835be8f08471201c27400bcd951ae3784bf143cf9450

Observation 4e853cd0-58bf-4c20-9f95-edc29fb87b1e · outbound

This paper cites Massive activations in large language models.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Massive activations in large language models

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.503056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:4ab1caab655dccf45d49dce2c57c52ffb527212c111ad73a5fff123d62f41e8f

Observation f6c75b00-d2a4-49f4-995c-699f964361e9 · outbound

This paper cites Omnisparse: Training-aware fine-grained sparse attention for long-video mllms.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Omnisparse: Training-aware fine-grained sparse attention for long-video mllms

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.996556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:9bf05cbbbaf116b1216e8576c04852802edd8ec58c4ea617c6e70065ef0b917b

Observation 8268f3fa-58e1-4791-b138-81a3cabacf43 · outbound

This paper cites Bitmar: Low-bit multimodal fusion with episodic memory for edge devices.

Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Bitmar: Low-bit multimodal fusion with episodic memory for edge devices

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:04:57.331581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:17:09.834609Z digest=sha256:70f90367ab8fcac64e74ba5165eb9f8e1c613097f9ceac4a9351fd92e94ccd3d

Pith citing papers

Observation 4fdc6361-eb0e-4b00-a440-7d96b8680378 · inbound

When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression cites this paper.

When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:51:42.864258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T01:46:59.423533Z digest=sha256:a0690fc8d870e201b7ed131de6bb2cbb8df3cf29038083783d7980e8d910d3e9

Observation 8922d645-bb58-49af-94e9-c1af7a027842 · inbound

Priming: Hybrid State Space Models From Pre-trained Transformers cites this paper.

Priming: Hybrid State Space Models From Pre-trained Transformers Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.357221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:b75c040d63a481130ade7bace6e4e12f68ee87fcad0ba4309d5e2c1ac870fd16

Observation 4f600ed0-32bd-4188-b9ed-d20e479b24bd · inbound

Registers Matter for Pixel-Space Diffusion Transformers cites this paper.

Registers Matter for Pixel-Space Diffusion Transformers Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:08:54.036916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T19:08:23.052621Z digest=sha256:e80494579b7ed07e158eb07554a9b24dcd74c892ae3affefeb188a03c3ee670f

Observation 4c0a1a53-3f16-4983-80fb-9df02a31f889 · inbound

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond cites this paper.

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:58:07.543129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T07:57:51.032025Z digest=sha256:e2f5f7f868b8a187787c594dfab196deb58aeaf398117cda47f8aa04d5250895