Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:45:23.455170Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2608.07009.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:45:23.455170Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 08450613-eb6a-43d6-aae1-f173b5525e8e · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b827acd-2743-44e3-9406-16faa4e31499 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management IndexCache: Accelerating sparse attention via cross-layer index reuse, 2026
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18eecb4a-35bc-4e38-9ce3-169bcea1ace4 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63cddaf1-c477-4db0-9da6-49064ca8accf · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Peters, and Arman Cohan
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4fe85e3-d7ab-442d-8289-be80dff3b30e · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ArkVale: Efficient generative LLM inference with recallable key-value eviction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0dd0d590-d780-4eb8-922c-7edb139f8458 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ESS: An offload- centric latent-cache management architecture for DeepSeek-V3.2-Exp, 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 206a681e-7d02-48f3-a697-3ce2dc7cc6f7 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management MagicPIG: LSH sampling for efficient LLM generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bf2644b7-59bd-4cd7-9575-b6aee5ef0486 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FlashAttention-2: Faster attention with better parallelism and work partitioning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 443f7eff-8486-4d41-9538-da640552c794 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 950f44f3-a46f-4438-a616-2d5359620b94 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V3.2: Efficient reasoning & agentic AI
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 51644309-8908-4f0b-8327-d9b009988433 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e82dd7-d097-4c32-a51c-b445c0fcbad2 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V4: Towards highly efficient million-token context intelligence, 2026
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6639af0c-5fa4-4eca-b8e8-d67202168ffe · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Cost-efficient large language model serving for multi-turn conversations with CachedAttention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db37646f-9abf-42aa-b2a3-918176736f6b · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management GLM-5: from Vibe Coding to Agentic Engineering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b101c3db-c662-4e7b-aa80-9a0f26e7f45c · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 617066a6-f89a-486d-afbc-84989ba4f922 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management NEO: Saving GPU memory crisis with CPU offloading for online LLM inference
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 95365699-59c1-4ad6-8de1-4cceda3b1a08 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Reformer: The efficient transformer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b6669bd-e03c-467f-99d5-3191d56b2bbe · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Gonzalez, Hao Zhang, and Ion Stoica
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3219acc-acf7-418a-98fd-1dfa903ad32d · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management InfiniGen: Efficient generative inference of large language models with dynamic KV cache management
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 037b387f-a27f-4a2e-85ad-31ae52571801 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management SnapKV: LLM knows what you are looking for before generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a899c11b-847a-401d-ab04-7ac5315e7d33 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ECHO: Efficient KV cache offloading with lossless prefetching for serving native sparse attention LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 18955ac0-7ab2-4126-9599-a8b5f22a4dc1 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management KIVI: A tuning-free asymmetric 2bit quantization for KV cache
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e5d2f803-6b22-43b3-9283-ec291a443307 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management NVIDIA GH200 Grace Hopper superchip
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1b2b2c5b-985b-4edd-acd9-cb1afae4eca1 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Splitwise: Efficient generative LLM inference using phase splitting
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2c8afc0a-6b0f-4cb4-bfab-05e645e472c4 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Mooncake: Trading more storage for less computation—a KVCache-centric architecture for serving LLM chatbot
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 47246dca-dccc-402b-9779-3fcad698a246 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Qwen3-30B-A3B-Thinking-2507
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 39509634-85f7-41c8-a6c7-7517926305b4 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Bench serving guide
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4b0eecd4-41e1-4e3b-8ef8-a7534b2d602c · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management HiSparse: Hierarchical sparse attention
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d841648c-33c0-4963-9330-7a92637bf412 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FlexGen: High-throughput generative inference of large language models with a single GPU
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cbefa9c6-2a90-43ae-bb7a-7eedeebdf0c6 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ShadowKV: KV cache in shadows for high-throughput long-context LLM inference
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2ff478fc-ffba-4d28-be6f-4d8d73867a95 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Quest: Query-aware sparsity for efficient long-context LLM inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 610d0934-c4f7-4171-906e-006158d27888 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Linformer: Self-Attention with Linear Complexity
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f6430f-d3a9-4593-8931-bccccdfc123d · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Efficient streaming language models with attention sinks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ef9abf-fc94-4a59-9ae7-35d9ac0249d9 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management SGLang HiCache: Fast hierarchical KV caching with your fa- vorite storage backends
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 229d2781-e7f0-422d-9803-08c083ce68a3 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management HiSparse: Turbocharging sparse attention with hierarchical memory
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf222180-b25d-4802-9352-9aabc53fcfa9 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Strata: Hierarchical Context Caching for Long Context Language Model Serving
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4644d1c-4603-434f-9620-4e0d0f09fdcb · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Qwen3 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ddc875-30a7-4cce-966b-3257618e3929 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Native sparse attention: Hardware-aligned and natively trainable sparse attention
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6683b0-4b82-4e1d-8838-331b0e82d336 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management GLM-5.2: Built for long-horizon tasks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 126ae745-301c-44e3-8cb1-1c20ab255f32 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management PQCache: Product quantization-based KVCache for long context LLM inference.Proceedings of the ACM on Management of Data, 3(3):201:1–201:30, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbdb14c9-42e4-49cf-beff-ad7264c493c0 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management H2O: Heavy-hitter oracle for efficient generative inference of large language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4e21e5-f645-416c-a7cb-0cdc26cc8fa7 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Gonzalez, Clark W
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d6f3142d-bdb0-409f-9231-eb31c460ab82 · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7ea78599-617f-4bfd-82cd-0322b5c4281c · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Longformer: The Long-Document Transformer
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccbcf82a-b759-4eb3-8e41-62a5b8a1444d · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d260293f-3b99-47e9-909d-a75a55b1efae · outbound
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Accessed 2026-05-04
Reference 2026
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
No inbound Pith citation observations are available.