Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:44:18.354744Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2510.24606.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:44:18.354744Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a7ec3d41-8d54-4bb0-8924-a73383bd884b · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference The Compressor-Retriever Architecture for Language Model OS
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e6cc209-8069-4c30-a792-be0f76ef7649 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Gemma 2: Improving Open Language Models at a Practical Size
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1873a0f2-cba9-45ce-9fb3-a3311a189cf7 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Gemma 3 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f1735dd-4f94-42a9-9304-23b6ea7908f2 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095e062d-1ca3-47bf-a3b7-d314d4b98893 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 136bda79-b2af-44c4-a312-226f3b42fda4 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b78366d-7034-4a81-b543-ad98329af1f6 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference On Gemma2-2b-it, retaining the top 1k tokens per layer, DHSA matches dense attention while substantially outperforming block sparse attention
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d36bdba-36ce-4b52-99c2-b9026f914c84 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Sliding-window attention uses a budget of 2048, block sparse attention 512 and KV compression (streaming LLM, h2o, pyramidKV)
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54711909-d2fc-45da-af4c-e45b342aa9e8 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Sliding- window attention uses a budget of 2048, block sparse attention 512 and KV compression (streaming LLM, h2o, pyramidKV)
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe1254e-71a7-41bc-a966-74fbff2ef4e1 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Finally, we note several potential influencing factors in KV compression methods (Fig
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87125184-9e45-400c-8394-5f47199d66a5 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference 0 2000 4000 6000 8000 10000 12000 14000 16000 Max KV Cache Capacity 3.2 3.4 3.6 3.8 4.0 4.2Time (s) Latency v.s
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0d3ba07-b0ee-4528-9f43-3cea5950c520 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference On the other hand, for all decode-stage methods, prefill time increases with context length since KV cache compression only applies during decoding
Reference 128
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c7d66e6-b2cb-45c8-a476-0885ed71c286 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Efficient Streaming Language Models with Attention Sinks
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a44c33b-1804-44a1-9a42-835bc06d0035 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference Longformer: The Long-Document Transformer
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9578a1aa-a2c1-4bc3-97ce-afc38ff1ad6c · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1f67058-5aeb-4794-9502-cf08d945e7e4 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2aa70a-518a-4ec3-b830-f9a654911648 · outbound
Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.