Pith. sign in

Paper Citation Record · LEDGER

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management

As of 23 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2608.07009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07009 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:45:23.455170Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08450613-eb6a-43d6-aae1-f173b5525e8e · outbound

This paper cites LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management LongBench v2: Towards deeper understanding and reasoning on realistic long-context multitasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.291458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.291458Z digest=sha256:11f80ec49d753895ffc65667659b1d9aae9b98f5f64488dd02d9e50fb83bdb6d

Observation 1b827acd-2743-44e3-9406-16faa4e31499 · outbound

This paper cites IndexCache: Accelerating sparse attention via cross-layer index reuse, 2026.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management IndexCache: Accelerating sparse attention via cross-layer index reuse, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.295722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.295722Z digest=sha256:4feee27a5b0b85307b12d6613fe41c82efa844a2c9cff447dae2c2d4974d04eb

Observation 18eecb4a-35bc-4e38-9ce3-169bcea1ace4 · outbound

This paper cites an unresolved cited work.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.299171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.299171Z digest=sha256:2e7e3338cf6fa93fe221ded333a7586bbd6e110867df9d585fd7d1f2531ccbfc

Observation 63cddaf1-c477-4db0-9da6-49064ca8accf · outbound

This paper cites Peters, and Arman Cohan.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Peters, and Arman Cohan

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.302928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.302928Z digest=sha256:0dbfd0e3dfe837b4539a5c25eb5702a22df9729337adc36b5cf1a2c96e91c65c

Observation f4fe85e3-d7ab-442d-8289-be80dff3b30e · outbound

This paper cites ArkVale: Efficient generative LLM inference with recallable key-value eviction.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ArkVale: Efficient generative LLM inference with recallable key-value eviction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.178841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.310628Z digest=sha256:c2384fec03e412586556e874e317d0d8ec806e8287c877f9f12b510efbb06b43

Observation 0dd0d590-d780-4eb8-922c-7edb139f8458 · outbound

This paper cites ESS: An offload- centric latent-cache management architecture for DeepSeek-V3.2-Exp, 2025.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ESS: An offload- centric latent-cache management architecture for DeepSeek-V3.2-Exp, 2025

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-10T16:45:23.810177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.314548Z digest=sha256:4815f76d6d14e9e02bc840331f07da078200818519b1e2809c21d3aad0c310c1

Observation 206a681e-7d02-48f3-a697-3ce2dc7cc6f7 · outbound

This paper cites MagicPIG: LSH sampling for efficient LLM generation.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management MagicPIG: LSH sampling for efficient LLM generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.167750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.318523Z digest=sha256:ba4dc92b1605b2de63c965bafea027c7aa0d3130b49c099ce0e5558ee9d95328

Observation bf2644b7-59bd-4cd7-9575-b6aee5ef0486 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.156172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.321833Z digest=sha256:2763466ef13e036e2335a068c92a3dfbb083b399a314e2fe747a97e63356885f

Observation 443f7eff-8486-4d41-9538-da640552c794 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.145252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.325223Z digest=sha256:51431d9b767632a2eb1fef8dcaa98926908562fd6fe99c600c2efffab52d895d

Observation 950f44f3-a46f-4438-a616-2d5359620b94 · outbound

This paper cites DeepSeek-V3.2: Efficient reasoning & agentic AI.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V3.2: Efficient reasoning & agentic AI

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.132858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.328699Z digest=sha256:05cd1d2e745e162b43a3e2ccb9e25b47243e7ed24437754233cf03dc206ea897

Observation 51644309-8908-4f0b-8327-d9b009988433 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.332283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.332283Z digest=sha256:c4d314784184def4a01856891b3265a3ca2db0f5c65e4e9120ef612ffce2c15e

Observation 87e82dd7-d097-4c32-a51c-b445c0fcbad2 · outbound

This paper cites DeepSeek-V4: Towards highly efficient million-token context intelligence, 2026.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DeepSeek-V4: Towards highly efficient million-token context intelligence, 2026

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.335996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.335996Z digest=sha256:a92fb100cdda28ea2c8a13f7da5526410c2443e0f69109d01fd8606cbf48d474

Observation 6639af0c-5fa4-4eca-b8e8-d67202168ffe · outbound

This paper cites Cost-efficient large language model serving for multi-turn conversations with CachedAttention.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Cost-efficient large language model serving for multi-turn conversations with CachedAttention

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.120914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.339727Z digest=sha256:96e3890563757ddf6b4b612c987ff0d2c6489cf3319d9abba86d832102b6da99

Observation db37646f-9abf-42aa-b2a3-918176736f6b · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management GLM-5: from Vibe Coding to Agentic Engineering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.343078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.343078Z digest=sha256:a85842c867f07cfb5efaec2a03775d48862a7357886035588d3ddd3768de831f

Observation b101c3db-c662-4e7b-aa80-9a0f26e7f45c · outbound

This paper cites FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.346750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.346750Z digest=sha256:a897d86e45a1fdfef6fc876f4ca4654dfee364ee8dc77f578b15223a76717fa1

Observation 617066a6-f89a-486d-afbc-84989ba4f922 · outbound

This paper cites NEO: Saving GPU memory crisis with CPU offloading for online LLM inference.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management NEO: Saving GPU memory crisis with CPU offloading for online LLM inference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.109280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.350702Z digest=sha256:64678672f4b6f06f85c64b63e67a13766a582d68b03be7a9e32ed149f993258e

Observation 95365699-59c1-4ad6-8de1-4cceda3b1a08 · outbound

This paper cites Reformer: The efficient transformer.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Reformer: The efficient transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.354354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.354354Z digest=sha256:d3f1c38db1161924af46a2a0f4f118e6ae58c493c16a66be7dd6def2605f8deb

Observation 7b6669bd-e03c-467f-99d5-3191d56b2bbe · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.358099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.358099Z digest=sha256:a527b4cbf24ecf104401bf3491f74b6c1d713047b4bc2501a1ddf4662dc47785

Observation c3219acc-acf7-418a-98fd-1dfa903ad32d · outbound

This paper cites InfiniGen: Efficient generative inference of large language models with dynamic KV cache management.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management InfiniGen: Efficient generative inference of large language models with dynamic KV cache management

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.091321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.361753Z digest=sha256:3fca6b4d94498d156eb96e59f3d6dd01c6c1a97a55c227b7e3348df76ba245c0

Observation 037b387f-a27f-4a2e-85ad-31ae52571801 · outbound

This paper cites SnapKV: LLM knows what you are looking for before generation.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management SnapKV: LLM knows what you are looking for before generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.079545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.365570Z digest=sha256:5219ba49bf4578d113138e1495a5e108af1f46e204354d1269667185ece4f871

Observation a899c11b-847a-401d-ab04-7ac5315e7d33 · outbound

This paper cites ECHO: Efficient KV cache offloading with lossless prefetching for serving native sparse attention LLMs.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ECHO: Efficient KV cache offloading with lossless prefetching for serving native sparse attention LLMs

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.068201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.368982Z digest=sha256:08ea52736e6aeef904e2eba4d952c0ef56e9da6523cbd4aef85c9832a5eb6cdf

Observation 18955ac0-7ab2-4126-9599-a8b5f22a4dc1 · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.057458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.372519Z digest=sha256:472cd247e0c1f31c4740632460732a993b96015d8c10d8268bc8369566e88168

Observation e5d2f803-6b22-43b3-9283-ec291a443307 · outbound

This paper cites NVIDIA GH200 Grace Hopper superchip.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management NVIDIA GH200 Grace Hopper superchip

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.046620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.375879Z digest=sha256:4ac37168e9cb7ba79d26d130f6269536623d100b23aa023ff689b9500f05904c

Observation 1b2b2c5b-985b-4edd-acd9-cb1afae4eca1 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Splitwise: Efficient generative LLM inference using phase splitting

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.034020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.379612Z digest=sha256:40cecf82dbf0dc7362f573802c2a7bf87ac30698c3eae759d11edb723d725e5f

Observation 2c8afc0a-6b0f-4cb4-bfab-05e645e472c4 · outbound

This paper cites Mooncake: Trading more storage for less computation—a KVCache-centric architecture for serving LLM chatbot.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Mooncake: Trading more storage for less computation—a KVCache-centric architecture for serving LLM chatbot

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.023147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.387043Z digest=sha256:094eada038070061b763113cfa2a71d68669a9059e3f990de4e1a30dad11a74d

Observation 47246dca-dccc-402b-9779-3fcad698a246 · outbound

This paper cites Qwen3-30B-A3B-Thinking-2507.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Qwen3-30B-A3B-Thinking-2507

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:24.011712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.390523Z digest=sha256:6667a695354d02b3c60d3e3dbf4d96420ee5847e56147a5e9342df33edbc58e2

Observation 39509634-85f7-41c8-a6c7-7517926305b4 · outbound

This paper cites Bench serving guide.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Bench serving guide

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.999816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.393999Z digest=sha256:6d3ee7855a517917aec741e89890eb49c832333b8eb5619a442b8cfda2b7cba2

Observation 4b0eecd4-41e1-4e3b-8ef8-a7534b2d602c · outbound

This paper cites HiSparse: Hierarchical sparse attention.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management HiSparse: Hierarchical sparse attention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.987113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.397212Z digest=sha256:b11452af4340c339e1772f67c5506bb92e30919fe63fa1f166f390572eb6d7f5

Observation d841648c-33c0-4963-9330-7a92637bf412 · outbound

This paper cites FlexGen: High-throughput generative inference of large language models with a single GPU.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management FlexGen: High-throughput generative inference of large language models with a single GPU

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.964365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.403797Z digest=sha256:9f25fb483ccff609f076a75c56d0c2286365ed60b1b6816ebbb6c2d9b1d45f86

Observation cbefa9c6-2a90-43ae-bb7a-7eedeebdf0c6 · outbound

This paper cites ShadowKV: KV cache in shadows for high-throughput long-context LLM inference.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management ShadowKV: KV cache in shadows for high-throughput long-context LLM inference

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.952778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.406916Z digest=sha256:65ef1a6532da87dd76bc9a3961c174295c5517e6303f54d90b1cada7c669fdf2

Observation 2ff478fc-ffba-4d28-be6f-4d8d73867a95 · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context LLM inference.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Quest: Query-aware sparsity for efficient long-context LLM inference

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.940899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.410330Z digest=sha256:f3e2e785c2dacabd718271a28c57e34cbe8925cffc2f08cb5b206dfc6ca25c9f

Observation 610d0934-c4f7-4171-906e-006158d27888 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Linformer: Self-Attention with Linear Complexity

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.414556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.414556Z digest=sha256:35837f318d17ae59640f2d0b3d2759189f7caae4af5a40748b3b93a0c84d51a9

Observation b2f6430f-d3a9-4593-8931-bccccdfc123d · outbound

This paper cites Efficient streaming language models with attention sinks.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Efficient streaming language models with attention sinks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.419073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.419073Z digest=sha256:9493437140a05163d5cf9024c749130aca4585facaae22e19c9df17ec8527c34

Observation 73ef9abf-fc94-4a59-9ae7-35d9ac0249d9 · outbound

This paper cites SGLang HiCache: Fast hierarchical KV caching with your fa- vorite storage backends.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management SGLang HiCache: Fast hierarchical KV caching with your fa- vorite storage backends

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.924134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.423045Z digest=sha256:f1d4405ca248dbdf31f46735802a9bc5f2973a6dcba0f0c235857e36f9d4f684

Observation 229d2781-e7f0-422d-9803-08c083ce68a3 · outbound

This paper cites HiSparse: Turbocharging sparse attention with hierarchical memory.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management HiSparse: Turbocharging sparse attention with hierarchical memory

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.910968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.426667Z digest=sha256:afca9d6193f83402ab24916aa6312955dc6bda24c831c4c175f48f53bf0d5886

Observation cf222180-b25d-4802-9352-9aabc53fcfa9 · outbound

This paper cites Strata: Hierarchical Context Caching for Long Context Language Model Serving.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Strata: Hierarchical Context Caching for Long Context Language Model Serving

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.429982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.429982Z digest=sha256:4a1439e24e9e8a71e157c7d2d09f51bb3b3da0faccccf52cbd915bac7f282658

Observation d4644d1c-4603-434f-9620-4e0d0f09fdcb · outbound

This paper cites Qwen3 Technical Report.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Qwen3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.433719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.433719Z digest=sha256:08f9296313abf8df6c0be1e2476767392dc4270b05a94c3e081036cf01631474

Observation 78ddc875-30a7-4cce-966b-3257618e3929 · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.437624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.437624Z digest=sha256:3c8ff6aa529ad56167cd84e69fea1014412994af3393321aedb2483965c1dbc0

Observation 2c6683b0-4b82-4e1d-8838-331b0e82d336 · outbound

This paper cites GLM-5.2: Built for long-horizon tasks.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management GLM-5.2: Built for long-horizon tasks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.899561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.441276Z digest=sha256:6665925a04360e9480b3cff3237ba2178986b035a071e1cb7875d498ef66ba5d

Observation 126ae745-301c-44e3-8cb1-1c20ab255f32 · outbound

This paper cites PQCache: Product quantization-based KVCache for long context LLM inference.Proceedings of the ACM on Management of Data, 3(3):201:1–201:30, 2025.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management PQCache: Product quantization-based KVCache for long context LLM inference.Proceedings of the ACM on Management of Data, 3(3):201:1–201:30, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.444854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.444854Z digest=sha256:4e8ff41790d1a5a015500c615177d716dd05dd28c6ac56dbf9f5ab984c220243

Observation dbdb14c9-42e4-49cf-beff-ad7264c493c0 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.448711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.448711Z digest=sha256:c9d46e8df0c45644f795daae3f6333f7a31568cd55e057cda94dbce9d38e2194

Observation 1b4e21e5-f645-416c-a7cb-0cdc26cc8fa7 · outbound

This paper cites Gonzalez, Clark W.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Gonzalez, Clark W

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.882352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.451960Z digest=sha256:c6b01b46c5d2a3a1c217c6115f0da8f07d525ef1288931b266589f2197368088

Observation d6f3142d-bdb0-409f-9231-eb31c460ab82 · outbound

This paper cites DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.870987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.455170Z digest=sha256:789db123e30d650951a75e99c2049cbf81e092729ab2609a374a0baf7f55d587

Observation 7ea78599-617f-4bfd-82cd-0322b5c4281c · outbound

This paper cites Longformer: The Long-Document Transformer.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Longformer: The Long-Document Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.306275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.306275Z digest=sha256:c7ca7505f7c953a81751630d5ad23ed3e04b956cdb83a3afe92b4e5b6741ccc2

Observation ccbcf82a-b759-4eb3-8e41-62a5b8a1444d · outbound

This paper cites an unresolved cited work.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T16:45:23.383477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:45:23.383477Z digest=sha256:ad6a6d236d70417f5a48b7d75dbba7833fe2fa7bfcc5e479ebddb2fba780b379

Observation d260293f-3b99-47e9-909d-a75a55b1efae · outbound

This paper cites Accessed 2026-05-04.

HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management Accessed 2026-05-04

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:45:23.975469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T16:45:23.400534Z digest=sha256:f139fe5e8d36a7cd4ec19deba7e4ba5f6ffac8b802cd7c7b65c5022badf0a3f7

Pith citing papers

No inbound Pith citation observations are available.