Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:06:32.755775Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 12 inbound Pith citation observations for arXiv:2506.08889.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:06:32.755775Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:11:48.900029Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T13:36:59.523914Z
78 of 78 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 603767e6-da74-41e3-9625-00bbc598aeab · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning URLhttps://github.com/tile-ai/tilelang
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3fca4ecb-e378-4f01-908d-ef2b18764eae · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Keyformer: Kv cache reduction through key tokens selection for efficient generative inference.Proceedings of Machine Learning and Systems, 6:114–127, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f5078a8-1d1e-4198-979f-7049585bdb4c · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283fb409-0322-4c71-9453-dd78c3e8b1cb · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning xLSTM: Extended Long Short-Term Memory
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09743327-ca1d-4c10-b8f3-6af12b746a81 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6000f946-9a71-4274-9b21-340fc650a4f6 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Longformer: The Long-Document Transformer
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0181f93e-8425-4ed1-a1f6-06a70775a97f · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Reducing transformer key-value cache size with cross-layer attention
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4815cf23-38f5-4b5a-ba4e-31639b59f255 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning R-kv: Redundancy-aware kv cache compression for training-free reasoning models acceleration.arXiv preprint arXiv:2505.24133, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f496848-d855-4471-8142-2c892c297b6b · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a057aa-bae1-493a-bc57-03ec9dab8d18 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb725fa-9e3e-45ef-80a5-30ddc1b6c7da · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MagicPIG: LSH Sampling for Efficient LLM Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf6d40a-4219-422e-b3fc-64ba19b1d892 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning PipeThreader: Software-defined pipelining for efficient dnn execution
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3abc9a08-dfa0-4a4e-b361-d1f14374e254 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Generating Long Sequences with Sparse Transformers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8673676c-8c53-4f7c-a57a-84f2dca653bf · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b6ae02-56b1-4b19-9cf1-daf8bc52f13f · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d536d28-56ee-48fd-98f7-a858af3525ea · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hymba: A Hybrid-head Architecture for Small Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e97183-6acd-47af-9892-7ea5a68e7b7b · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec087002-3424-487c-802c-c6f20a8d8573 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Moa: Mixture of sparse attention for automatic large language model compression.arXiv preprint arXiv:2406.14909, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d96b03-d758-480d-8dea-c50af189e0f7 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6574c471-194b-4c87-9736-87f3e24d4f67 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3700472-5d13-4c40-a724-ad7110675c01 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Better & Faster Large Language Models via Multi-token Prediction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1573ba-6df3-4109-8a22-a7c0cb30fbfb · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fafa75e-6103-4783-8a4f-1d254ddb0115 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f224880f-aba3-45d4-b743-5c6e89313233 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Omnikv: Dynamic context selection for efficient long-context llms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 26d5b8c7-b6ca-4bdb-b8bb-248d491e1837 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Measuring Massive Multitask Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a125b793-69f1-4931-b3c5-c781c13f3bc4 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Squeezed attention: Accelerating long context length llm inference.arXiv preprint arXiv:2411.09688, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e30bed75-7a91-4ce3-afde-a4304d5cc93a · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RaaS: Reasoning-Aware Attention Sparsity for Efficient LLM Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f630c6ae-6e46-4b5c-bb03-cf706b18b457 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning OpenAI o1 System Card
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4999cafd-4f82-4a86-8e13-90bd51e734a5 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e755f89d-3c32-44d1-89a3-42874f1b507c · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Kullback-leibler divergence
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ebb34e9-7d68-4134-a99f-d5fbb211f785 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Transformers are rnns: Fast autoregressive transformers with linear attention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcbff325-2586-45a4-86c8-cc1488914b75 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gonzalez, Hao Zhang, and Ion Stoica
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94ff5f04-9d57-4434-b600-a547ef219e7c · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7879778f-56b8-4abc-885c-8df83eb59889 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Fast inference from transformers via speculative decoding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67473948-fa99-448e-bea7-ef575a77dd23 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d7ae517-d409-4357-b5c4-6d3b18cdd479 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning SCBench: A KV Cache-Centric Analysis of Long-Context Methods
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f08edf-0805-4bd4-9d62-5ae7f5e8cf61 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Snapkv: Llm knows what you are looking for before generation.Advances in Neural Information Processing Systems, 37:22947–22970, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3838f42-b8e1-41e0-8dfb-6e1323801559 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Twilight: Adaptive attention sparsity with hierarchical top- p pruning.arXiv preprint arXiv:2502.02770, 2025
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation babc598e-3810-4f5a-bff9-881b61898fe8 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Adaptive Computation Pruning for the Forgetting Transformer
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ff6959f-9958-4507-a9ec-1f287dd68ea6 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c219a013-8686-4648-8283-db95ac9584b7 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DeepSeek-V3 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2608c643-fffb-422b-877e-691614308d6c · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3530e94f-8a37-4537-b556-73ba454253da · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e68eb24-e125-495b-beec-fbcb87c254a1 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 293d79b4-1d15-40dc-a4f6-305781561c89 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a362d9b-f617-41b7-bbf1-6fb73bd032fe · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Inference-time sparse attention with asymmetric indexing
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d667d25b-3933-414d-a5b6-143357e4e96c · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Aime problems and solutions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c01b0e2c-6de0-4f45-ad94-953af6858da4 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning RWKV: Reinventing RNNs for the Transformer Era
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf432c8-7479-46db-b6a4-144757bb9505 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gpqa: A graduate-level google-proof q&a benchmark
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49d1105-3f23-496a-83b3-036244b4a626 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Flashattention-3: Fast and accurate attention with asynchrony and low-precision.Advances in Neural Information Processing Systems, 37:68658–68685, 2024
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b4bec04-2885-409e-a38b-b7d3854429c1 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Fast Transformer Decoding: One Write-Head is All You Need
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d984e696-1486-4856-95d9-d4a857901aa0 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e05d7f-6813-44c6-813c-7b8878ce31b7 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Retentive Network: A Successor to Transformer for Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd83167e-1222-400d-8695-d27c9f5de1de · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning You only cache once: Decoder-decoder architectures for language models.Advances in Neural Information Processing Systems, 37:7339–7361, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 973f0b67-f390-49b2-b41c-433cfb715649 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Rectified Sparse Attention
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e5477f-7825-4fd7-9640-a0efa6e1f46c · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67b2b37-dab9-4ff1-a390-54ec09637381 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Minicpm4: Ultra-efficient llms on end devices
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd3ca030-3cf0-40aa-a4f0-706877f84057 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Triton: an intermediate language and compiler for tiled neural network computations
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac5c415-4962-4918-ba04-6d5b0b62b380 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a4ba9d5f-4361-4aed-88d1-e2f654aeaaf8 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e51bd4e9-f1ae-497b-b21f-cb827e7bdd73 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da04754-dc55-459a-abb4-c029a17bb797 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Efficient Streaming Language Models with Attention Sinks
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af72d851-2eef-4c38-93dc-e2ca446dc95f · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f77f0b-62cc-436f-a5a8-9fb99f5859c8 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning XAttention: Block Sparse Attention with Antidiagonal Scoring
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b5e48b-48d0-4ef7-8fc6-80bd77976df3 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Qwen3 Technical Report
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a73f34-695e-4872-ae7d-4196c0ddbddb · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab96539-ebe4-47de-8355-181e7446dd8d · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Post-Training Sparse Attention with Double Sparsity
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95dfe03b-fd92-4e07-8d3f-2a2c169094e6 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Gated Linear Attention Transformers with Hardware-Efficient Training
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba368b4c-83f6-4f2c-9b86-76d1891b91af · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Parallelizing Linear Transformers with the Delta Rule over Sequence Length
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92fa3dfc-1f8d-4cf3-90db-dd119862ed8b · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292cce17-8c9a-4178-887e-7497c8c511e0 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Hardware-Efficient Attention for Fast Decoding
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 563df4f2-cca6-45f9-a9b7-96a0005e3186 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33: 17283–17297, 2020
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce549a7a-91ea-40a8-b8bd-ec2b2a44d192 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning In-context KV-Cache Eviction for LLMs via Attention-Gate
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506d2d3c-2bcb-4aea-9efb-b6071374512a · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning PQCache: Product Quantization-based KVCache for Long Context LLM Inference
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f109a2ac-2869-4b5e-9f1b-2099b952ac7d · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Spargeattn: Accurate sparse attention accelerating any model inference
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ab92dd5a-4c2b-42ca-936f-3deb329e8c07 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e4c6dc1-84ad-4eee-897e-4fe0150e1e66 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning Sglang: Efficient execution of structured language model programs.Advances in Neural Information Processing Systems, 37:62557–62583, 2024
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53b8b144-c7aa-40e1-9c49-15a0541be264 · outbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning ROLLER: Fast and efficient tensor compilation for deep learning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9efce4c6-ac26-4d67-b107-bec22e6a9b2e · inbound
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b4a83254-06b0-4481-9dd6-7cc3af0c7330 · inbound
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e34d13a-9f28-459a-b6b1-af2ff9ee8997 · inbound
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3772f3b3-2b63-4da8-a594-fef652386ee2 · inbound
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7e737445-b0fd-423c-8c2c-b47f79990846 · inbound
Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ea5e97a0-7941-4c73-84d0-4223c643f029 · inbound
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4e375f8d-a619-4f26-8b6a-d8c8893f666e · inbound
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 015213c6-84eb-4fbb-bb4f-1b04d4a97998 · inbound
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 25b2bac7-6670-42dc-a6cd-e008de0a2d35 · inbound
You Only Index Once: Cross-Layer Sparse Attention with Shared Routing SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5ee18d25-c704-4b04-9543-64c33b112150 · inbound
PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dac1432e-871b-4332-b81a-15b0cca4a735 · inbound
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba11f5f-1fb7-44d2-be6d-de0d04b20416 · inbound
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.