Pith. sign in

Paper Citation Record · LEDGER

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

As of 11 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 12 inbound Pith citation observations for arXiv:2506.08889.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08889 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:06:32.755775Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:11:48.900029Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:36:59.523914Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 603767e6-da74-41e3-9625-00bbc598aeab · outbound

This paper cites URLhttps://github.com/tile-ai/tilelang.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning URLhttps://github.com/tile-ai/tilelang

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.702936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.546420Z digest=sha256:655ca7ae79a7ffc8502c71dc3559862789b754407f88e2d18f4ce5e1df96101b

Observation 3fca4ecb-e378-4f01-908d-ef2b18764eae · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.Proceedings of Machine Learning and Systems, 6:114–127, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.Proceedings of Machine Learning and Systems, 6:114–127, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.644684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.550117Z digest=sha256:23708fef664d076a46fe543288000f1651cd80faeeced5a7b501dc7d8fb5497b

Observation 7f5078a8-1d1e-4198-979f-7049585bdb4c · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.553125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.553125Z digest=sha256:85f4eed11f94390f6cef1a4b82a5bf59a4863cab0a704461792f6cfd60417cdc

Observation 283fb409-0322-4c71-9453-dd78c3e8b1cb · outbound

This paper cites xLSTM: Extended Long Short-Term Memory.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning xLSTM: Extended Long Short-Term Memory

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.556320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.556320Z digest=sha256:4f48d2f58d0a93c1730e6759473ca02b139f7f44c8f490b3a82ac28617e96c10

Observation 09743327-ca1d-4c10-b8f3-6af12b746a81 · outbound

This paper cites RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.559683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.559683Z digest=sha256:b745157384b2a8fcfb040cbaaa94734b97c401b686d3228b12ca4e033d52be61

Observation 6000f946-9a71-4274-9b21-340fc650a4f6 · outbound

This paper cites Longformer: The Long-Document Transformer.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Longformer: The Long-Document Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.563190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.563190Z digest=sha256:696042768bfc56b77f14567f9f3d44d33093cf87992781adb8c216267d34a540

Observation 0181f93e-8425-4ed1-a1f6-06a70775a97f · outbound

This paper cites Reducing transformer key-value cache size with cross-layer attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Reducing transformer key-value cache size with cross-layer attention

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.609507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.566599Z digest=sha256:a69981632b2b91906ba21d513315b132b38fe44f399a5addeec38e7048dde840

Observation 4815cf23-38f5-4b5a-ba4e-31639b59f255 · outbound

This paper cites R-kv: Redundancy-aware kv cache compression for training-free reasoning models acceleration.arXiv preprint arXiv:2505.24133, 2025.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning R-kv: Redundancy-aware kv cache compression for training-free reasoning models acceleration.arXiv preprint arXiv:2505.24133, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.569593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.569593Z digest=sha256:a6061cacd9bc5b43e67e9e38488dd1f4f642b2a555849f3fa023f10127020f43

Observation 8f496848-d855-4471-8142-2c892c297b6b · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.572234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.572234Z digest=sha256:57cd72fdde81b312b736b0629b47f4547bd4028279db6d387621a3cb5a6bc610

Observation 44a057aa-bae1-493a-bc57-03ec9dab8d18 · outbound

This paper cites RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.575162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.575162Z digest=sha256:983be678932e3ca221a93e9ac97f0cbbcad699feb0fbef83ec96b22d5ddeb9cb

Observation bfb725fa-9e3e-45ef-80a5-30ddc1b6c7da · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.577962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.577962Z digest=sha256:69c43e95731a9ecb413fd96b134ff10095f537a99285dfb10ffc5628e8d77e26

Observation 2bf6d40a-4219-422e-b3fc-64ba19b1d892 · outbound

This paper cites PipeThreader: Software-defined pipelining for efficient dnn execution.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning PipeThreader: Software-defined pipelining for efficient dnn execution

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.570272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.580743Z digest=sha256:9ea479cd1e48cc982163d8ad4a41ae3904255f28ace34b43d81f5e32346581f4

Observation 3abc9a08-dfa0-4a4e-b361-d1f14374e254 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Generating Long Sequences with Sparse Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.583149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.583149Z digest=sha256:15d6f7a9de2c968be6f33a5396fd079218bfd1886521aad295ffe5f1e90eab5e

Observation 8673676c-8c53-4f7c-a57a-84f2dca653bf · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.585847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.585847Z digest=sha256:91a48562590de099796b588d67e370f813bbaa3cf08129be55b91393533664a8

Observation 97b6ae02-56b1-4b19-9cf1-daf8bc52f13f · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.588405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.588405Z digest=sha256:4db88f51bcecae9fa16c36bfe281ba64d2594854c0b7abe1ab8fbcb1889abd2d

Observation 5d536d28-56ee-48fd-98f7-a858af3525ea · outbound

This paper cites Hymba: A Hybrid-head Architecture for Small Language Models.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hymba: A Hybrid-head Architecture for Small Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.591124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.591124Z digest=sha256:43cb4c624d1ebd15110f8d18b14f550736f91177b230d632b5759274c721d8f9

Observation e4e97183-6acd-47af-9892-7ea5a68e7b7b · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.594106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.594106Z digest=sha256:7205857328a5fca9b85ad2366bf2508aeb3462b6f0a7ad7f6e51d2034f13a21f

Observation ec087002-3424-487c-802c-c6f20a8d8573 · outbound

This paper cites Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.596786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.596786Z digest=sha256:30ac0777921259701b844aa66784d45a0b2df39cc91fed4bfa287b0c083be1a2

Observation 59d96b03-d758-480d-8dea-c50af189e0f7 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.599244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.599244Z digest=sha256:717888f15793854f0d355ef0360da5f23727425ffa6883455ea86a5b3146b54a

Observation 6574c471-194b-4c87-9736-87f3e24d4f67 · outbound

This paper cites Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.602023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.602023Z digest=sha256:cf3c6a3e9319592044e1e9d2bd291925d7b65c33b970f7e4f67da627def3f916

Observation f3700472-5d13-4c40-a724-ad7110675c01 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Better & Faster Large Language Models via Multi-token Prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.604608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.604608Z digest=sha256:66f6fe597b4fe867ff008537b50e79002f2839f89f254021c68bbfd32e9d646b

Observation 3d1573ba-6df3-4109-8a22-a7c0cb30fbfb · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.607376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.607376Z digest=sha256:f105d9c2e2504e5cbd3ebd391c7bf5ce304bb4189a02eee6f33e9bc4a285b14d

Observation 9fafa75e-6103-4783-8a4f-1d254ddb0115 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.610150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.610150Z digest=sha256:acbdb0a5c0f3b0513161c571e3f09931d9f84a87af30931bec29d6fba9742387

Observation f224880f-aba3-45d4-b743-5c6e89313233 · outbound

This paper cites Omnikv: Dynamic context selection for efficient long-context llms.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Omnikv: Dynamic context selection for efficient long-context llms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.538895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.612606Z digest=sha256:5e60b1ed04f9117c15704fe9175cc5d0db6c7b1b7552118f538dc37507f4f748

Observation 26d5b8c7-b6ca-4bdb-b8bb-248d491e1837 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Measuring Massive Multitask Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.615350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.615350Z digest=sha256:7459de0e2b6c791464b93b605b7d2f02d6afa7416c9996078194b1e8f394340d

Observation a125b793-69f1-4931-b3c5-c781c13f3bc4 · outbound

This paper cites Squeezed attention: Accelerating long context length llm inference.arXiv preprint arXiv:2411.09688, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Squeezed attention: Accelerating long context length llm inference.arXiv preprint arXiv:2411.09688, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.618100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.618100Z digest=sha256:96a84e5abe929856796eeed674fa758c8fb8f419adedd661b7de8b73c4c496b8

Observation e30bed75-7a91-4ce3-afde-a4304d5cc93a · outbound

This paper cites RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.620481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.620481Z digest=sha256:aebb71256c3def1434433b6a78a278538377f9209620ec75b750669841c51835

Observation f630c6ae-6e46-4b5c-bb03-cf706b18b457 · outbound

This paper cites OpenAI o1 System Card.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning OpenAI o1 System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.622907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.622907Z digest=sha256:74dfd42b015b8d4063d326b484c650e5fa865856595ebb8abcb84d5a1cc41c3a

Observation 4999cafd-4f82-4a86-8e13-90bd51e734a5 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.625484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.625484Z digest=sha256:46881d22e76ab90f744c53c0e627d60aa9d30dc4c48b0177ec70b6baa72ecb60

Observation e755f89d-3c32-44d1-89a3-42874f1b507c · outbound

This paper cites Kullback-leibler divergence.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Kullback-leibler divergence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.518944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.628140Z digest=sha256:228e903fa925b8bb58b602d1948a163b395781a2cd9b34a5e712efbb4eea9491

Observation 6ebb34e9-7d68-4134-a99f-d5fbb211f785 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.630511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.630511Z digest=sha256:5b99ec677e408130a1cd58fd8bd136b588a546419a9fdcd1a1434f9f2d4114ec

Observation bcbff325-2586-45a4-86c8-cc1488914b75 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gonzalez, Hao Zhang, and Ion Stoica

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.632901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.632901Z digest=sha256:de0aad432c6ce4a5d584802a1b4d3978921e73dd09ed3c82e921140a1fbaa46d

Observation 94ff5f04-9d57-4434-b600-a547ef219e7c · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.635500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.635500Z digest=sha256:f69f6b75b0f2f1dbaf31805b1fa3a6713ea42d6ac0ab125bfcd6f5440b9d9c2a

Observation 7879778f-56b8-4abc-885c-8df83eb59889 · outbound

This paper cites Fast inference from transformers via speculative decoding.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Fast inference from transformers via speculative decoding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.637982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.637982Z digest=sha256:11592dbb60560a39ca4ccdf9fe36ee6ca138507972b98b933a36aec241a45a6c

Observation 67473948-fa99-448e-bea7-ef575a77dd23 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.640407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.640407Z digest=sha256:0cc6ff1e176b811d99b349163ab9177f3d8d76f644d7d24a6fba2d0b4f4d3145

Observation 1d7ae517-d409-4357-b5c4-6d3b18cdd479 · outbound

This paper cites SCBench: A KV Cache-Centric Analysis of Long-Context Methods.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.642988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.642988Z digest=sha256:8f63a400f3cbdb8de20b592b1f396b2bad84ae14823f77f0f6d7794fb35339ed

Observation 03f08edf-0805-4bd4-9d62-5ae7f5e8cf61 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.645647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.645647Z digest=sha256:00ff11e0ea54e865aebb2a6d4d3772f0c26213213bcba03a7b54f3ee989f5b7f

Observation f3838f42-b8e1-41e0-8dfb-6e1323801559 · outbound

This paper cites Twilight: Adaptive attention sparsity with hierarchical top- p pruning.arXiv preprint arXiv:2502.02770, 2025.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Twilight: Adaptive attention sparsity with hierarchical top- p pruning.arXiv preprint arXiv:2502.02770, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.648079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.648079Z digest=sha256:a689d2fd3e0c1be4693df3d0e749d0b1b203172a7634bc0aeefaa40a70e211ff

Observation babc598e-3810-4f5a-bff9-881b61898fe8 · outbound

This paper cites Adaptive Computation Pruning for the Forgetting Transformer.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Adaptive Computation Pruning for the Forgetting Transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.650507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.650507Z digest=sha256:cd321a7003fe8fbb74be995c89acafcb3bf865c0fc458d9e06eb0d952d5470c4

Observation 0ff6959f-9958-4507-a9ec-1f287dd68ea6 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.652878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.652878Z digest=sha256:1b691702e846183fc987375527894039ee52f2f21a6e8f1b3f6a88d590b4f806

Observation c219a013-8686-4648-8283-db95ac9584b7 · outbound

This paper cites DeepSeek-V3 Technical Report.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-V3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.655660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.655660Z digest=sha256:c085134362b6b15d52f3a16580f64a743def26882a940e9976116f2a4012fbe0

Observation 2608c643-fffb-422b-877e-691614308d6c · outbound

This paper cites RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.658257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.658257Z digest=sha256:987eb388b91e7893ec5299936be6a252a541ec7d55140166534bcd427e9f7790

Observation 3530e94f-8a37-4537-b556-73ba454253da · outbound

This paper cites Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.660824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.660824Z digest=sha256:0855adba60a8870794e0cb53022370aa9d05ce966ce36da970e986ec188313ac

Observation 5e68eb24-e125-495b-beec-fbcb87c254a1 · outbound

This paper cites an unresolved cited work.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.663343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.663343Z digest=sha256:eb7f1dde43f9037565a5eeb8c14185c056d37c0881b2adf5de833b715448cb8c

Observation 293d79b4-1d15-40dc-a4f6-305781561c89 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.665648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.665648Z digest=sha256:42342f9ba2dbef54b16fa0cbd090c89f94c1a03b1fe5d75091eb99bf5b07b12b

Observation 7a362d9b-f617-41b7-bbf1-6fb73bd032fe · outbound

This paper cites Inference-time sparse attention with asymmetric indexing.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Inference-time sparse attention with asymmetric indexing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.671170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.671170Z digest=sha256:5c46f02c5684808b674e22ddfa3f41622a47849f12a90df39e7abf13b7009a0d

Observation d667d25b-3933-414d-a5b6-143357e4e96c · outbound

This paper cites Aime problems and solutions.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Aime problems and solutions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.472215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.673524Z digest=sha256:5783cfd09a9fbe12995842b6e7d4636ee2aa8427e60849499c6bf5c81a7563a6

Observation c01b0e2c-6de0-4f45-ad94-953af6858da4 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RWKV: Reinventing RNNs for the Transformer Era

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.676223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.676223Z digest=sha256:1f28ff109e0476ea27658e1b1f72af3bdecd4c9c166615764e54b9eea42cdc3e

Observation 0cf432c8-7479-46db-b6a4-144757bb9505 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gpqa: A graduate-level google-proof q&a benchmark

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.678846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.678846Z digest=sha256:5c5c1538be9cfe4e391c2e049fb53befdb0c8333a09487266f8cf63f7092fac5

Observation d49d1105-3f23-496a-83b3-036244b4a626 · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.681295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.681295Z digest=sha256:410c8be8cee152e47a5904a2442a12beb8b962e1c046f763049d37e3a856f324

Observation 8b4bec04-2885-409e-a38b-b7d3854429c1 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Fast Transformer Decoding: One Write-Head is All You Need

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.683732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.683732Z digest=sha256:43e168d82ff602905f54b0b9206bda1fdc85dca4ab4680a026be628b325140b0

Observation d984e696-1486-4856-95d9-d4a857901aa0 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.686218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.686218Z digest=sha256:d43eb5aec8040cebca49d3204895277c6459dba16dd2a64f2354a196b9c1f065

Observation e4e05d7f-6813-44c6-813c-7b8878ce31b7 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Retentive Network: A Successor to Transformer for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.688555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.688555Z digest=sha256:6432ba14a310cf71e3e152bf421df18b642f0dc26b21a688efb1349d4b8bb186

Observation fd83167e-1222-400d-8695-d27c9f5de1de · outbound

This paper cites You only cache once: Decoder-decoder architectures for language models.Advances in Neural Information Processing Systems, 37:7339–7361, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning You only cache once: Decoder-decoder architectures for language models.Advances in Neural Information Processing Systems, 37:7339–7361, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.691080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.691080Z digest=sha256:c49099aca1c5c3b993a93d6688a4870e658640ea471fc974cae6bf37c0fa52a8

Observation 973f0b67-f390-49b2-b41c-433cfb715649 · outbound

This paper cites Rectified Sparse Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Rectified Sparse Attention

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.693476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.693476Z digest=sha256:d1008a645876704c4d763238dc0083c691eeeeaf2027eb9bdb5ff495a2eb357d

Observation e7e5477f-7825-4fd7-9640-a0efa6e1f46c · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.696425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.696425Z digest=sha256:c3523a44cc2c2ada690a74b9845bb7f498b1ef3e45cdbc577265f5cb146b4419

Observation a67b2b37-dab9-4ff1-a390-54ec09637381 · outbound

This paper cites Minicpm4: Ultra-efficient llms on end devices.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Minicpm4: Ultra-efficient llms on end devices

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.432664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.699380Z digest=sha256:9c682d885910663f0f2043528ea0004a04655addd1a46812839531f62b473aa9

Observation cd3ca030-3cf0-40aa-a4f0-706877f84057 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Triton: an intermediate language and compiler for tiled neural network computations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.701827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.701827Z digest=sha256:0ab736877195ba5ad7c44252b0854c1c67a45e4acf5107a18aba5eb149561098

Observation 6ac5c415-4962-4918-ba04-6d5b0b62b380 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.411160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.704664Z digest=sha256:49672652186d0ef482182d5e57d77b24efde910fab768b561f1205ce11c39bbe

Observation a4ba9d5f-4361-4aed-88d1-e2f654aeaaf8 · outbound

This paper cites Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.401791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.707584Z digest=sha256:b105815f8f3b222dcfab6bb849332093a4c0ec0857387b3239bdb35a581ae616

Observation e51bd4e9-f1ae-497b-b21f-cb827e7bdd73 · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.710225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.710225Z digest=sha256:83a7022ec40bc4d02832713c19ed571ede653e4001de8548fc67ad86c19c5b06

Observation 0da04754-dc55-459a-abb4-c029a17bb797 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Efficient Streaming Language Models with Attention Sinks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.712682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.712682Z digest=sha256:c2fa19fae43d931c65a61215e6e016ceff46abcb339749776eca17ebabab1dc7

Observation af72d851-2eef-4c38-93dc-e2ca446dc95f · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.715316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.715316Z digest=sha256:b39cc521bf92349231216320df881fe2c1d4885fcaa120d88c97d3f75a755dd3

Observation c1f77f0b-62cc-436f-a5a8-9fb99f5859c8 · outbound

This paper cites XAttention: Block Sparse Attention with Antidiagonal Scoring.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning XAttention: Block Sparse Attention with Antidiagonal Scoring

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.718359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.718359Z digest=sha256:28d2718b110d4709837ba370314d16062cb8f2ed6a4fc7bf72b3dbe8240167e0

Observation c2b5e48b-48d0-4ef7-8fc6-80bd77976df3 · outbound

This paper cites Qwen3 Technical Report.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Qwen3 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.721183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.721183Z digest=sha256:4e9c7a0c76a2c9ef6028c4b4d632972284950b57bcf0bf9f6f37df3739220865

Observation e6a73f34-695e-4872-ae7d-4196c0ddbddb · outbound

This paper cites LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.724124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.724124Z digest=sha256:77e2c1b969e71384d32d138c8aab94d471d30c564a41c8479f75c6416aecee84

Observation bab96539-ebe4-47de-8355-181e7446dd8d · outbound

This paper cites Post-Training Sparse Attention with Double Sparsity.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Post-Training Sparse Attention with Double Sparsity

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.726866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.726866Z digest=sha256:7309b36fa46a7ca39bf2704e2928d138dfbee3ba13c6d7ceeeb1df7e858a187b

Observation 95dfe03b-fd92-4e07-8d3f-2a2c169094e6 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.729535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.729535Z digest=sha256:9671af6fe5fd715aec63fee27316df636b33f5cbbd88ea18a0713224b34784c4

Observation ba368b4c-83f6-4f2c-9b86-76d1891b91af · outbound

This paper cites Parallelizing Linear Transformers with the Delta Rule over Sequence Length.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.732274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.732274Z digest=sha256:3189d5064febe9226b8129874447f0be3eff48c00e89de69285125b8eacc43b6

Observation 92fa3dfc-1f8d-4cf3-90db-dd119862ed8b · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.735020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.735020Z digest=sha256:f658238b6c3dcf999fde18f1f9b817e2262ae9821a245b9301e01ae60eb65913

Observation 292cce17-8c9a-4178-887e-7497c8c511e0 · outbound

This paper cites Hardware-Efficient Attention for Fast Decoding.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hardware-Efficient Attention for Fast Decoding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.737715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.737715Z digest=sha256:c8972568a50f7ba51ba1704dc9cb9189bb284732a2cd66e9c0dca0c234b85228

Observation 563df4f2-cca6-45f9-a9b7-96a0005e3186 · outbound

This paper cites Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297, 2020.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297, 2020

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.740450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.740450Z digest=sha256:cdb1cf628e88debda3e52faddca7fd9faa087cafc6ee16a75ecf33e2f39574ca

Observation ce549a7a-91ea-40a8-b8bd-ec2b2a44d192 · outbound

This paper cites In-context KV-Cache Eviction for LLMs via Attention-Gate.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning In-context KV-Cache Eviction for LLMs via Attention-Gate

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.742963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.742963Z digest=sha256:1ad004e66ed9a14929cae4cd19a3010468fae75d246d84d7f73085d7bc75356b

Observation 506d2d3c-2bcb-4aea-9efb-b6071374512a · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.745585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.745585Z digest=sha256:d8c67dc2e72dcf80ccd99222bbe7aca56712aa71f8e62c6936e34ed2bd5219d1

Observation f109a2ac-2869-4b5e-9f1b-2099b952ac7d · outbound

This paper cites Spargeattn: Accurate sparse attention accelerating any model inference.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Spargeattn: Accurate sparse attention accelerating any model inference

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.387280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.748548Z digest=sha256:560190525eaa77da34bfba591b5780500e8f28c81f6326cc0245ff475b8ba138

Observation ab92dd5a-4c2b-42ca-936f-3deb329e8c07 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.379076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.751054Z digest=sha256:47e798e6a4aa1f4b2e14492212ac343f60f495966019bbb2a1a5a8ca4076ec54

Observation 0e4c6dc1-84ad-4eee-897e-4fe0150e1e66 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in Neural Information Processing Systems, 37:62557–62583, 2024.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Sglang: Efficient execution of structured language model programs.Advances in Neural Information Processing Systems, 37:62557–62583, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.371151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.753405Z digest=sha256:502b75112aa0df72e6cf328a2821b9480d5751b4dc55280fc0797f51ded4d0b3

Observation 53b8b144-c7aa-40e1-9c49-15a0541be264 · outbound

This paper cites ROLLER: Fast and efficient tensor compilation for deep learning.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning ROLLER: Fast and efficient tensor compilation for deep learning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:06:33.362723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T05:06:32.755775Z digest=sha256:30067a0bb8df68f23b030879ffce844dc5531a908d74f3153e79c4eac87488dd

Pith citing papers

Observation 9efce4c6-ac26-4d67-b107-bec22e6a9b2e · inbound

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning cites this paper.

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:26:02.860786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T07:25:56.953876Z digest=sha256:fef09452bccde1b82ae66459d58988d5dc7f9d3d345d1822c28cecfd3d3c8dbf

Observation b4a83254-06b0-4481-9dd6-7cc3af0c7330 · inbound

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding cites this paper.

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.608942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T22:20:53.856657Z digest=sha256:aaaefc515431221826c0f24b959148fe28cc88c0d4f9278304abd57106681066

Observation 1e34d13a-9f28-459a-b6b1-af2ff9ee8997 · inbound

Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference cites this paper.

Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:28:29.860253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T00:25:49.807277Z digest=sha256:1b450321585514f2152f02652fedcc3b6fd2cb1b130db4d1418603b5cba96176

Observation 3772f3b3-2b63-4da8-a594-fef652386ee2 · inbound

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning cites this paper.

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:10.553722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T11:17:43.769244Z digest=sha256:4b66b44137cdbfeed283bf72ff1ea2d4d7115b07700df66ab9dab2b5fe76c849

Observation 7e737445-b0fd-423c-8c2c-b47f79990846 · inbound

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving cites this paper.

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:01:25.896526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T13:15:21.201950Z digest=sha256:77399745d54a8babbf14c6a75f10eb3cb98ec2ce5c6c89c47d1505c9de462ddd

Observation ea5e97a0-7941-4c73-84d0-4223c643f029 · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:54.140296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:9da5592af711d0c1f9431ac7f0e7ca899503d6e29f8fa82fa8e709cf69328817

Observation 4e375f8d-a619-4f26-8b6a-d8c8893f666e · inbound

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction cites this paper.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:26.703884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:ada4df7bfd0bbc291a8f0b600c2a1bea83afe6e9524e2f4bde5d59b30d53be4e

Observation 015213c6-84eb-4fbb-bb4f-1b04d4a97998 · inbound

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention cites this paper.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:53:13.586873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:b8169f1b7e3a2198b32e98ea9bc057343e576b11571e2ad59f703b48894fc4b6

Observation 25b2bac7-6670-42dc-a6cd-e008de0a2d35 · inbound

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing cites this paper.

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.525347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T01:06:04.896501Z digest=sha256:37abbbf43bc42c2b6ebb0778d00321bd4be1ef84054f82000de4b910f149e185

Observation 5ee18d25-c704-4b04-9543-64c33b112150 · inbound

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention cites this paper.

PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T11:04:04.584216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:04:04.584216Z digest=sha256:4642d34c9176daa277c96ed05194c0073e79de7de82e93c868a838e94458ce57

Observation dac1432e-871b-4332-b81a-15b0cca4a735 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:53.318889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:53.318889Z digest=sha256:94fa604e71193b9bac3423b7353f7c333731b32c3c6a6a796cdf9713dfe75d7b

Observation 5ba11f5f-1fb7-44d2-be6d-de0d04b20416 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SeerAttention-R: Sparse Attention Adaptation for Long Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:11:48.900029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:11:48.900029Z digest=sha256:69a8a959dfc39c412c2f7c367cb0b0e29014bbb6344e914ddd273a1cb0b566bd