Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:10.193536Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2505.24179.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:10.193536Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f1181232-b07a-4c73-a2fe-3033e696e651 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Booksum: A collection of datasets for long-form narrative summarization,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b13ab18e-98c6-42c3-beed-4206e077c3af · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformer Based Implementation for Automatic Book Summarization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 512c5be2-fb3e-47e2-8561-02f765bf403d · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Booookscore: A systematic exploration of book-length summarization in the era of LLMs,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d871b43-4e3d-458f-adad-b036b31dd6f2 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Peek across: Improving multi-document modeling via cross-document question-answering,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0492a883-5869-4d00-b33a-23eaf3cf4da0 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Quality: Question answering with long input texts, yes!,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51bd3625-2748-411d-8673-c8a17d439cb9 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Eli5: Long form question answering,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11e139f6-855e-441e-9eb1-dee98a1dcd55 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Teaching code llms to use autocompletion tools in repository-level code generation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d394e65c-6b84-4b07-85e6-84e86266e87e · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Rlcoder: Reinforcement learning for repository-level code completion,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd7e2f28-2281-477c-ae9e-67a4033939f2 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling The llama 3 herd of models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dbd470f-d605-4360-a5f6-439816f638b5 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Qwen2.5-1M Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7024eba2-1cc9-4e41-8937-4623bc89b760 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Gemma 3 technical report,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6c144f1-0568-490a-8e08-926250a7c254 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Deepseek-v3 technical report,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ebd4d0b-4c23-4611-a066-be719b0961b9 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Attention is all you need,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d31f66-3863-43e5-8cba-31d97e3a4ff1 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Challenges in deploying long-context transformers: A theoretical peak performance analysis,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc30b61e-5109-45e4-afa2-82763aa26a9e · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b8de900-60eb-4bab-8ac4-71d127049db3 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1fbed45-88d9-46ac-9bde-e090760fa7d3 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Generating Long Sequences with Sparse Transformers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fb3dcee-4063-43f4-9f7a-9f4a119ed900 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Big bird: Transformers for longer sequences,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1077f3fa-2255-4a52-aecf-c71753c526b4 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Longformer: The Long-Document Transformer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd0ba8b8-d185-4029-98eb-6479edf9cd5d · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Efficient streaming language models with attention sinks,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c1f97f4-31be-47bf-a7e4-89aa14373d90 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 880c0f54-d757-48f2-b888-07afa42268bf · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23c86693-7505-4dfc-ab59-2ab4c6e8a005 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flexprefill: A context-aware sparse attention mechanism for ef- ficient long-sequence inference,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c91011ad-266a-48c1-a33f-3149672e4140 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Spargeattn: Accurate sparse attention accelerating any model inference,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c552fb85-094a-49c9-ab38-0dc97356c943 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling A training-free sub-quadratic cost transformer model serving framework with hierarchically pruned attention,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d02b2b5-2c14-4ecb-9d69-11e44e8a1782 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling When attention sink emerges in language models: An empirical view,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 777fd1ba-d746-48d0-8058-35432def5b16 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling H2o: Heavy-hitter oracle for efficient generative inference of large language models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b2391d7-3bc3-4e6a-9457-1abb8c98786d · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Snapkv: Llm knows what you are looking for before generation,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0d3eeb6-e310-4965-99c8-b85aa2d41096 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling PQCache: Product Quantization-based KVCache for Long Context LLM Inference
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68f5277a-7d78-4afb-813e-57c80128c8d0 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d099f893-9734-41ba-994f-43a4a313878a · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Qwen2.5 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fbdeb21-5945-4ade-a74b-dfe82e553374 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Characterizing Prompt Compression Methods for Long Context Inference
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e17897b-9f3b-4af4-bd78-0b15be0fe53d · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb1ec33-d153-40f9-9093-18fd3a040d60 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8677bee6-8834-4439-a1f5-446277b5db5b · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Compressing Context to Enhance Inference Efficiency of Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ad1def-52a4-41d4-9065-0738cec69d22 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c9d2ff-0379-4a51-80bc-4a1411f3cd37 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Moa: Mix- ture of sparse attention for automatic large language model compression,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9985a76-0204-4317-95d5-3abf1cd66568 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 783534c0-5b4b-49b8-bff0-f2b48d7627f8 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1b6144-7239-4bab-a852-bd4d7f9f8a3e · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52639c41-fccf-42ba-a00f-b3047216356f · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122136bc-932d-4d23-a356-0c9231bfa3bc · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling RWKV: Reinventing RNNs for the Transformer Era
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ffea3ab-fd1e-43dd-ac41-cd4805e84c8e · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Gated Linear Attention Transformers with Hardware-Efficient Training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072a6031-85e0-4157-9026-1e82c3954bfd · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Mamba: Linear-time sequence modeling with selective state spaces,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 656dcb64-0540-4ede-b338-6e62a40d085d · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformers are ssms: generalized models and efficient algorithms through structured state space duality,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8094898b-a09f-4271-ac6e-1e47a32bf967 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ede19b-d24e-4f37-9da5-4551d0f4b7eb · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8b8f53b-8397-443e-8325-0edc21cec549 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling MagicPIG: LSH sampling for efficient LLM generation,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 421f902d-d766-48e5-9a7c-159250df986e · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e50e07-b88e-4883-b6a4-fc715debffd4 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 815689dc-7de3-4398-8a7d-9d441cf63646 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc3a49e-e94f-44ed-96cd-72c9db4abe33 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling A Simple and Effective $L_2$ Norm-Based Strategy for KV Cache Compression
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7771ceee-8c80-42c4-ab8e-5e1c995482d3 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Cam: Cache merging for memory-efficient llms inference,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc45d404-06d6-4b1f-b8bb-d6a0af867bdd · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling SubGen: Token Generation in Sublinear Time and Memory
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d3e88fc-7553-4943-bd0e-8208da72cbd5 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flashattention: Fast and memory-efficient exact attention with io-awareness,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc81e05c-8df2-4ac4-9c0f-1c127d93a221 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a37b6a5-5f58-4758-9f2b-8da82c070851 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Flashattention-3: Fast and accurate attention with asynchrony and low-precision,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d15aa6d2-c9ba-4662-8b4f-9464d76ad745 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b210d8f7-8e27-4561-b883-0e2d0f162c3c · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Sageattention2 technical report: Accurate 4 bit attention for plug-and-play inference acceleration,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3fe215-278a-4b2a-a516-3d74a2c82088 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a3992ac-f256-4aa1-87be-50971fe84b26 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Triton: an intermediate language and compiler for tiled neural network computations,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b2475b7e-d19d-470d-bab2-1864008f8952 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Transformers: State-of-the-art natural language processing,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7f2c074-e5c4-4711-9955-6e96db117fca · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44475d5-4d5b-4b6a-aa3b-fe28cae2bbc4 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b2b6213-5e68-4e18-8bb8-1ed454617470 · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling LongBench: A bilingual, multitask benchmark for long context understanding,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01963edd-4a23-4919-b6cd-d2123191a12a · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling ∞bench: Extending long context evaluation beyond 100k tokens,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a29cb238-8b36-448d-a202-e4a809b09dea · outbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Needle in a haystack - pressure testing llms,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.