Pith. sign in

Paper Citation Record · LEDGER

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2510.24606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.24606 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:44:18.354744Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7ec3d41-8d54-4bb0-8924-a73383bd884b · outbound

This paper cites The Compressor-Retriever Architecture for Language Model OS.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference The Compressor-Retriever Architecture for Language Model OS

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:15.795227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:15.795227Z digest=sha256:19615c8c6d3f4a861c2f9ae42c8ef92553187e7c46dfc8f19ed22c51aea289f2

Observation 2e6cc209-8069-4c30-a792-be0f76ef7649 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Gemma 2: Improving Open Language Models at a Practical Size

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:16.434996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:16.434996Z digest=sha256:9638dbf2d15e32649f6031e6cd7a308716e6ff77b32121820f0721ca58a3cf33

Observation 1873a0f2-cba9-45ce-9fb3-a3311a189cf7 · outbound

This paper cites Gemma 3 Technical Report.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Gemma 3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:16.564775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:16.564775Z digest=sha256:e07be20e823f22f23521f6366abd020685470ca77ae4f58365bc2cd92de7df51

Observation 4f1735dd-4f94-42a9-9304-23b6ea7908f2 · outbound

This paper cites ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:16.794770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:16.794770Z digest=sha256:600052fe5083079625ac642a9b49bc95be6c4f157e6741f106a233f4a47cc26f

Observation 095e062d-1ca3-47bf-a3b7-d314d4b98893 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:17.063957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:17.063957Z digest=sha256:75a8c9bfebe4315f156f7f0753cad023fbc298818785ad9c1c9a1bc9a7d03f6a

Observation 136bda79-b2af-44c4-a312-226f3b42fda4 · outbound

This paper cites an unresolved cited work.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:17.215338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:17.215338Z digest=sha256:91e4354f8d4e766438e47f725841641ac2cf82984dd17bf54d65309e110f7e2f

Observation 1b78366d-7034-4a81-b543-ad98329af1f6 · outbound

This paper cites On Gemma2-2b-it, retaining the top 1k tokens per layer, DHSA matches dense attention while substantially outperforming block sparse attention.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference On Gemma2-2b-it, retaining the top 1k tokens per layer, DHSA matches dense attention while substantially outperforming block sparse attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:17.535583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:17.535583Z digest=sha256:f346d8d50443882e6e19d69756e308cafc7285d8a4153478fc819396493b9126

Observation 1d36bdba-36ce-4b52-99c2-b9026f914c84 · outbound

This paper cites Sliding-window attention uses a budget of 2048, block sparse attention 512 and KV compression (streaming LLM, h2o, pyramidKV).

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Sliding-window attention uses a budget of 2048, block sparse attention 512 and KV compression (streaming LLM, h2o, pyramidKV)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:17.684744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:17.684744Z digest=sha256:74c3ceef6ad033676b792bcce87beda461a762311ae6df64e4eb01e456e7f918

Observation 54711909-d2fc-45da-af4c-e45b342aa9e8 · outbound

This paper cites Sliding- window attention uses a budget of 2048, block sparse attention 512 and KV compression (streaming LLM, h2o, pyramidKV).

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Sliding- window attention uses a budget of 2048, block sparse attention 512 and KV compression (streaming LLM, h2o, pyramidKV)

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:17.964749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:17.964749Z digest=sha256:181174a33f36d05b231d7930c25926bbe3118b6d91c17ae82e7a1b4c6a525983

Observation 3fe1254e-71a7-41bc-a966-74fbff2ef4e1 · outbound

This paper cites Finally, we note several potential influencing factors in KV compression methods (Fig.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Finally, we note several potential influencing factors in KV compression methods (Fig

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:18.183182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:18.183182Z digest=sha256:ceb4942050c4aa060f4508b4e0749e161b575b8eb1a68e6da8878043432adb30

Observation 87125184-9e45-400c-8394-5f47199d66a5 · outbound

This paper cites 0 2000 4000 6000 8000 10000 12000 14000 16000 Max KV Cache Capacity 3.2 3.4 3.6 3.8 4.0 4.2Time (s) Latency v.s.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference 0 2000 4000 6000 8000 10000 12000 14000 16000 Max KV Cache Capacity 3.2 3.4 3.6 3.8 4.0 4.2Time (s) Latency v.s

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:18.354744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:18.354744Z digest=sha256:7f5b483d601665cb305df86a79ae5c5dd3b87df9251175bdb9cde424f15d0eb1

Observation a0d3ba07-b0ee-4528-9f43-3cea5950c520 · outbound

This paper cites On the other hand, for all decode-stage methods, prefill time increases with context length since KV cache compression only applies during decoding.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference On the other hand, for all decode-stage methods, prefill time increases with context length since KV cache compression only applies during decoding

Reference 128

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:17.604884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:17.604884Z digest=sha256:0baeebf805b45bb2de8654b9893749adb79deea2108bbc9e526908f2dcb2ab5b

Observation 9c7d66e6-b2cb-45c8-a476-0885ed71c286 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Efficient Streaming Language Models with Attention Sinks

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:16.904840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:16.904840Z digest=sha256:93f4fcc79d8c2f396f2e6e81f1ab2ec39ffebe61041147eb89f2dbd805afd434

Observation 1a44c33b-1804-44a1-9a42-835bc06d0035 · outbound

This paper cites Longformer: The Long-Document Transformer.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Longformer: The Long-Document Transformer

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:15.991143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:15.991143Z digest=sha256:4f27c7c2c36a7c7d89494e62334592870b52381868567c18ad4165b3df7d049b

Observation 9578a1aa-a2c1-4bc3-97ce-afc38ff1ad6c · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:16.274893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:16.274893Z digest=sha256:429732d24eb300edcb19acec0a1febff24b8f64f18594d9c3e2e6c908fe2f06d

Observation d1f67058-5aeb-4794-9502-cf08d945e7e4 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:16.124951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:16.124951Z digest=sha256:0fd4c57f421e6ec8912ed007fad79c48b6c99e9b14265f6cdb5dc467990e2e4e

Observation 1c2aa70a-518a-4ec3-b830-f9a654911648 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T07:44:16.653934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:44:16.653934Z digest=sha256:0b8dd0a0696cc4c809b4b4ab07c276e6ecb40fd6e698b6415b71156c9574d9f1

Pith citing papers

No inbound Pith citation observations are available.