Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2407.02490.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T11:54:14.367710Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 9b2ac96e-76a3-4183-bfa8-0a575b874f14 · inbound
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e637f98-6e48-42a6-8a2c-3a55a8834366 · inbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a6c4347b-3505-4498-8c4a-ecc4928d63e9 · inbound
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c4f212c6-33a8-4b7f-830c-206048453844 · inbound
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 681f0673-0b72-43a4-abd7-1ab0cb30250d · inbound
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7277d24d-9352-4285-820c-0cb3f6496af7 · inbound
MoBA: Mixture of Block Attention for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65217528-7120-4a7a-b774-60cfb657162c · inbound
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ca52ff46-90bc-4a78-ae48-08d8c6da7500 · inbound
Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87954259-ef6d-4d24-a4b2-7bf3539ab4e2 · inbound
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 691b9a4c-d2e9-45d6-876a-a9f44abdf1c1 · inbound
Prism: Spectral-Aware Block-Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f6ced1-9abc-41e9-a942-c1eb01f4f7dd · inbound
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e8fe1080-28b1-4054-957e-2b2ed372cf7f · inbound
Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b77b6de-5270-476c-8eec-3f8775bcb0e7 · inbound
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 371997cc-fc2a-45a3-a6fc-09deef0a4244 · inbound
StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3bcb5b8b-a10f-4efd-8c81-1b156e449bba · inbound
Long Context Pre-Training with Lighthouse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e946a6a-3acd-40b1-887a-3f10d559de34 · inbound
Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44cbbdd6-45e9-46b0-93d7-d27c7a5d104e · inbound
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 82c0a072-b8d5-43c0-a43e-a8a86834bd94 · inbound
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f543456c-1afd-4396-ba3f-1c23272a88e3 · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40b9c12d-7e6f-48da-b1f5-6d1606fbe75d · inbound
How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 220dc4c9-14fb-475a-915e-5bbb705f6507 · inbound
SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4a10a0d-e597-4083-a80a-1d11a8bd2bbf · inbound
From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d46ba56a-1e4d-4629-a0b0-f65436b9c3a7 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91bc4a43-88c4-4dc1-a536-759e80b6eaf9 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 132
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46920f48-52ae-436b-9d5a-203aa039cf3a · inbound
NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 517cc6a6-ee9e-4000-bfeb-1fec7201f700 · inbound
Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ea4c474-9722-4fe5-a252-a84a522b42ae · inbound
Hierarchical Global Attention (HGA) MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0fcff984-9533-4404-b38f-594b5faea618 · inbound
Uncertainty-gated selection for block-sparse attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45063632-881f-418d-8182-17cf68216e57 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f387e1f-e50b-47f9-88e7-abfa0bf8fff3 · inbound
KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.