Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 60 inbound Pith citation observations for arXiv:2407.02490.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:45:20.504684Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 9b2ac96e-76a3-4183-bfa8-0a575b874f14 · inbound
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e637f98-6e48-42a6-8a2c-3a55a8834366 · inbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a6c4347b-3505-4498-8c4a-ecc4928d63e9 · inbound
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 726bed92-0016-47fd-bf64-5646b9ccce34 · inbound
Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 117ea9d5-a022-4e3d-b118-60286fc7f7ba · inbound
Boosting Long-Context Management via Query-Guided Activation Refilling MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd7d52b-6ac0-4ad4-8121-e7826a92f1f3 · inbound
A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf4cfec-b40e-49b2-aef5-aeff1eecc7fd · inbound
Bootstrap Your Own Context Length MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70bb0c2-05b7-47f3-a30e-f3e8c6faca67 · inbound
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbe6fb6-d8c0-41bd-b5d2-ce0472f62f3e · inbound
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2582b16-804f-4781-8624-a8375140aac3 · inbound
Scaling Inference-Efficient Language Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4f212c6-33a8-4b7f-830c-206048453844 · inbound
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cd041ad0-0ee2-45fc-9301-b0f8a83ec79d · inbound
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e538726b-8ab0-46d7-9e42-01e58a067ce0 · inbound
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI) MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81687733-b76f-4519-b52d-389e12613747 · inbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ca09d6-542e-4079-a12d-02af750dd0e8 · inbound
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 869c4170-945e-4fed-824b-5686e7ea904b · inbound
APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d83b80b3-601d-441d-b7cf-6f2b72df8a91 · inbound
LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ed5391c-3d37-4a5c-be11-bce0d9996892 · inbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681f0673-0b72-43a4-abd7-1ab0cb30250d · inbound
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7277d24d-9352-4285-820c-0cb3f6496af7 · inbound
MoBA: Mixture of Block Attention for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9b1d3248-b234-4373-9449-4043757d043a · inbound
LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1033f936-1941-4a9e-9745-7e5a7660e1f8 · inbound
Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa330f3-898e-4217-8508-1b36ea25a20b · inbound
R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeec097a-a6a7-4685-b0a1-af1c8d334661 · inbound
Reward Model Generalization for Compute-Aware Test-Time Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c98ba6d-c323-4975-9849-0c133478a1d6 · inbound
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4999cafd-4f82-4a86-8e13-90bd51e734a5 · inbound
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f11bb7e-adc0-419b-9367-506515bf257a · inbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35b17de-6ccc-4761-af96-722dbe43503c · inbound
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15d543b0-28aa-4420-97a8-ca925f1e9c4a · inbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b745b4-46bb-49aa-9d1a-daeb32a4bd09 · inbound
Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adcc9033-8d91-4d3d-87c7-087cc89f694c · inbound
Dynamic Sparse Causal-Attention Temporal Networks for Interpretable Causality Discovery in Multivariate Time Series MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df5ba07c-3d6d-494f-b1f5-228867137817 · inbound
Synergy: End-to-end Concept Model MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e6b9a98-e7c4-4f03-8bc3-bd9a83e5d865 · inbound
NABLA: Neighborhood Adaptive Block-Level Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada1c4d4-38ac-4c43-a5f7-507ee13b7a3d · inbound
DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65217528-7120-4a7a-b774-60cfb657162c · inbound
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ca52ff46-90bc-4a78-ae48-08d8c6da7500 · inbound
Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87954259-ef6d-4d24-a4b2-7bf3539ab4e2 · inbound
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 691b9a4c-d2e9-45d6-876a-a9f44abdf1c1 · inbound
Prism: Spectral-Aware Block-Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f6ced1-9abc-41e9-a942-c1eb01f4f7dd · inbound
PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e8fe1080-28b1-4054-957e-2b2ed372cf7f · inbound
Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b77b6de-5270-476c-8eec-3f8775bcb0e7 · inbound
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 371997cc-fc2a-45a3-a6fc-09deef0a4244 · inbound
StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3bcb5b8b-a10f-4efd-8c81-1b156e449bba · inbound
Long Context Pre-Training with Lighthouse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e946a6a-3acd-40b1-887a-3f10d559de34 · inbound
Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 44cbbdd6-45e9-46b0-93d7-d27c7a5d104e · inbound
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 82c0a072-b8d5-43c0-a43e-a8a86834bd94 · inbound
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f543456c-1afd-4396-ba3f-1c23272a88e3 · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 40b9c12d-7e6f-48da-b1f5-6d1606fbe75d · inbound
How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 220dc4c9-14fb-475a-915e-5bbb705f6507 · inbound
SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e4a10a0d-e597-4083-a80a-1d11a8bd2bbf · inbound
From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d46ba56a-1e4d-4629-a0b0-f65436b9c3a7 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 91bc4a43-88c4-4dc1-a536-759e80b6eaf9 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 132
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46920f48-52ae-436b-9d5a-203aa039cf3a · inbound
NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 517cc6a6-ee9e-4000-bfeb-1fec7201f700 · inbound
Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ea4c474-9722-4fe5-a252-a84a522b42ae · inbound
Hierarchical Global Attention (HGA) MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0fcff984-9533-4404-b38f-594b5faea618 · inbound
Uncertainty-gated selection for block-sparse attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45063632-881f-418d-8182-17cf68216e57 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8f387e1f-e50b-47f9-88e7-abfa0bf8fff3 · inbound
KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a485c77-41fc-46fd-a732-9357902f67c8 · inbound
SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac14ab7f-d113-4e76-9b76-816b3b767422 · inbound
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.