Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:41.059939Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 2 inbound Pith citation observations for arXiv:2506.15704.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:41.059939Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T07:03:08.617257Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T07:04:20.888331Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7e9b751b-e021-4306-9155-4dc11094e21f · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6905c1fa-97b4-4f8c-8379-e543693552b3 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4818a542-f672-4751-9d18-0e618417980e · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Loki then and now: the trickster against civilization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fbc23119-1404-43fb-bcbe-a736885ebc2c · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MagicPIG: LSH Sampling for Efficient LLM Generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b64cb96-4555-4e9f-96da-9b6ea86f7653 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c429e67-9fbe-4570-8558-ab2d3cd7d6ab · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Human-like episodic memory for infinite context llms
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c97e9d-4fe5-4fba-ba4d-0a63006dc951 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550258a3-3477-4a84-96ab-1b4fb5775f14 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd36822-23d0-4a69-a797-88e0c5ff4a97 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f11bb7e-adc0-419b-9367-506515bf257a · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e46c474-80d9-4b1b-8ce7-4d9e4401fd52 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34eb7be-896c-4b89-b6ba-31cd8f00ef2a · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Compute Or Load KV Cache? Why Not Both?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1bfc8f9-aeab-4816-b3b3-d2d996cf9159 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Efficient memory management for large language model serving with pagedattention
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 661a023e-a340-4058-832c-8704d5177c46 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache management
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eada5741-34d4-4ef6-a9c4-bf9fa63b8c1a · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Snapkv: Llm knows what you are looking for before generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16d82528-4536-421b-acd9-37cd4177d06c · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding DeepSeek-V3 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 977bba42-95d7-49de-a137-96176e3d0470 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e878039f-0681-4e7b-9be6-f5d943c4b78e · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc97e210-8720-4018-962e-cd4d98de5a87 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e514596-25c8-486d-96ff-73b94b3258e6 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 952fef15-d388-47c7-aa55-f96d7c1b1060 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Code Llama: Open Foundation Models for Code
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc38a9bf-8c56-45a5-abae-122c51efc83e · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Flexgen: High-throughput generative inference of large language models with a single gpu
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9545cbf4-34d8-46f4-9560-9cacafa53065 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f60f49-951a-40a4-a26d-2e8c88f2a8bf · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed9b6f2-dc40-4894-92bb-2aa4d92a1ff7 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding LLaMA: Open and Efficient Foundation Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a23d7f7-75e2-45b2-9981-415acf380d00 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3bc751-1820-4855-b643-3b526a4061b9 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Infllm: Unveiling the intrinsic capacity of llms for under- standing extremely long sequences with training-free memory.arXiv e-prints, pages arXiv–2402, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa115202-2d34-45d5-9d2a-faa4234d69e4 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d557360f-bcb7-4c09-9807-c01bc5effd5e · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Efficient Streaming Language Models with Attention Sinks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23115589-32d9-4e42-b217-8997271304a3 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Qwen2.5-1M Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5437d5b9-1181-4359-adaf-152295e111b3 · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding Orca: A distributed serving system for {Transformer-Based} generative models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5774c3-cbdc-4a02-aab5-6c6eaea9326c · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding PQCache: Product Quantization-based KVCache for Long Context LLM Inference
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0057c0b2-26c2-4e71-9083-60f8c3ffe1db · outbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adc68189-911a-4ea2-bb71-b3f964fbeaf9 · inbound
Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b692aa6d-3478-4390-a79c-b57fa3cd51c3 · inbound
Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.