Pith. sign in

Paper Citation Record · LEDGER

MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2407.02490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.02490 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:54:14.367710Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9b2ac96e-76a3-4183-bfa8-0a575b874f14 · inbound

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling cites this paper.

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:58:29.171156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T09:58:29.057357Z digest=sha256:f23c4f944090679b07ed27e10034c56924c6b52cd08d8d2985cfd0c27712e9a3

Observation 8e637f98-6e48-42a6-8a2c-3a55a8834366 · inbound

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval cites this paper.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:01.955613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:60f0a3ab24bb21f8b8920a3aa58eedbf5241424fceb3c926d881d308b5a71d1c

Observation a6c4347b-3505-4498-8c4a-ecc4928d63e9 · inbound

DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads cites this paper.

DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:49:16.795888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T11:49:16.654836Z digest=sha256:c08f9a36198ae9c255336b69ed1bf843af9599616859e8011afc060770db5f6e

Observation c4f212c6-33a8-4b7f-830c-206048453844 · inbound

FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration cites this paper.

FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:02:30.356815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T03:59:58.634512Z digest=sha256:9e191d69fc3e0d245c456c32c6127549663ff7558f1a827e4d0ec7ef5a053186

Observation 681f0673-0b72-43a4-abd7-1ab0cb30250d · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:46:30.161102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:ec284e84f5736bfa2be27a29c7615d1105ce2631b1f735705089207fb3b72632

Observation 7277d24d-9352-4285-820c-0cb3f6496af7 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.253626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:4518dfaacf241da96c0af3a2669b3ad58daced721ccefadc5bbd92be3af9aa0f

Observation 65217528-7120-4a7a-b774-60cfb657162c · inbound

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference cites this paper.

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:52.082431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T22:03:10.316005Z digest=sha256:7a19083f45e2633c36b2a735ea78aa2f70dd5ddc8b2dc4d452c0283bf62b82fc

Observation ca52ff46-90bc-4a78-ae48-08d8c6da7500 · inbound

Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety cites this paper.

Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:54:14.367710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:54:14.367710Z digest=sha256:b610cfc55c163a7eec3cc899d955fb749e82d7be2b3d2aee8d0be602c9ab12c1

Observation 87954259-ef6d-4d24-a4b2-7bf3539ab4e2 · inbound

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs cites this paper.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:30:55.155139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:61e4829a040fd877617f488eb5bd9046c2879110d2eacfab44d32507952095a5

Observation 691b9a4c-d2e9-45d6-876a-a9f44abdf1c1 · inbound

Prism: Spectral-Aware Block-Sparse Attention cites this paper.

Prism: Spectral-Aware Block-Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:22:53.923474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:22:53.923474Z digest=sha256:0774be0190c3abf3c5726101a3694926f7c74116e3eb4540f58a36709b236e3b

Observation 83f6ced1-9abc-41e9-a942-c1eb01f4f7dd · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.744912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:387bd3aa9e7cc44453d70ee53c0687fbf5e78975a1a97a1b758f392017c0278a

Observation e8fe1080-28b1-4054-957e-2b2ed372cf7f · inbound

Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions cites this paper.

Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:19:49.419805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:19:49.419805Z digest=sha256:cfcbfc434289bd07efc8df5343cced7d8d29a7bd6ef8ad06c5d7d7d7154df47f

Observation 5b77b6de-5270-476c-8eec-3f8775bcb0e7 · inbound

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention cites this paper.

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:43:00.751161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:40:28.854082Z digest=sha256:8ca283e79ff07b9568fedaab87f156d94a1e0b389d00c17495d9d4b74a0e5238

Observation 371997cc-fc2a-45a3-a6fc-09deef0a4244 · inbound

StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k cites this paper.

StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:15:37.731448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T18:44:42.456111Z digest=sha256:8799d12d996383b07a61bf4f5daa9b67c1909c2f3beee9a86cb4f89d7a14b2ec

Observation 3bcb5b8b-a10f-4efd-8c81-1b156e449bba · inbound

Long Context Pre-Training with Lighthouse Attention cites this paper.

Long Context Pre-Training with Lighthouse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:09.397118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:10:38.610613Z digest=sha256:7a54a0b9309a8149e727989021ba598a6a9385572cb0b4bc95b08bcd0424781e

Observation 6e946a6a-3acd-40b1-887a-3f10d559de34 · inbound

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache cites this paper.

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:15:50.429219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:15:35.871863Z digest=sha256:6781c27c11fb7fe610d1580e587ced6b8d2ab964650a91e0022fb1a44eae01ac

Observation 44cbbdd6-45e9-46b0-93d7-d27c7a5d104e · inbound

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention cites this paper.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.547437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:aef78a4575545d7983d6512204ea1fc52f43377e5911b23f259a4c1e83f15d3f

Observation 82c0a072-b8d5-43c0-a43e-a8a86834bd94 · inbound

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference cites this paper.

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:03:59.863178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:03:07.685253Z digest=sha256:c78edd0ea7bde2e3ec97bd3019a048ee44470f5d654c8e8c5ca60876a18cd34b

Observation f543456c-1afd-4396-ba3f-1c23272a88e3 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.749346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:354358a198a6d7564d0b678dc60642ad875511b295b1109195a27ef7318700f1

Observation 40b9c12d-7e6f-48da-b1f5-6d1606fbe75d · inbound

How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models cites this paper.

How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:09.159889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T22:52:13.337419Z digest=sha256:b9599942786b87898ae8bb43a43d6afb125896275256e6804d5689e5bee0798f

Observation 220dc4c9-14fb-475a-915e-5bbb705f6507 · inbound

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance cites this paper.

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:31.268022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:24:31.109508Z digest=sha256:68140a4dd438f247eacefa218b72fc43d68bd2d3c81c79d3f64534bee0cc8111

Observation e4a10a0d-e597-4083-a80a-1d11a8bd2bbf · inbound

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs cites this paper.

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:01:07.726301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T16:57:02.709967Z digest=sha256:203629831649cec08d22733b502806ba83a71533be25ff669e750c653c0cd852

Observation d46ba56a-1e4d-4629-a0b0-f65436b9c3a7 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:09:36.976767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:2f4d9aaf37a6b6a2b0b60d53d93b6a09883382c61808e7c54df385ec854da34f

Observation 91bc4a43-88c4-4dc1-a536-759e80b6eaf9 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:13.460701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:13.460701Z digest=sha256:512c1eb0d0078b322f569e57d5cc016f34b460a6439d5bff717d53076dd80593

Observation 46920f48-52ae-436b-9d5a-203aa039cf3a · inbound

NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation cites this paper.

NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.719634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T05:00:15.336703Z digest=sha256:2d64683b42526cd4b911c416338dc90c0f1f124e4782b38ac3868a65dfd34e31

Observation 517cc6a6-ee9e-4000-bfeb-1fec7201f700 · inbound

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM cites this paper.

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:24:21.385249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T07:19:48.530272Z digest=sha256:7048b9dfb0c753283e74897cdfc896bb76d55941ec07124c4325e02af3458709

Observation 8ea4c474-9722-4fe5-a252-a84a522b42ae · inbound

Hierarchical Global Attention (HGA) cites this paper.

Hierarchical Global Attention (HGA) MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:55:28.774424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T06:54:58.841989Z digest=sha256:dfc10f49cb68912442030ee6e1534ae8abacf2c0293e86b0cd25ad2f33585aa6

Observation 0fcff984-9533-4404-b38f-594b5faea618 · inbound

Uncertainty-gated selection for block-sparse attention cites this paper.

Uncertainty-gated selection for block-sparse attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T22:15:14.580916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T22:15:14.580916Z digest=sha256:f509800a62f0e9828592073f1c9aff3c1eac60cef8094f48c6acbee1d4dc3244

Observation 45063632-881f-418d-8182-17cf68216e57 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.231104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:37a5fbe3e0386da2db0d47125a104abbc1e1f0a7b239552bc38b7ec41cf3f440

Observation 8f387e1f-e50b-47f9-88e7-abfa0bf8fff3 · inbound

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems cites this paper.

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T19:52:40.866676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:52:40.866676Z digest=sha256:b0fc1099cb7cfe5a35f1f4bce5d56b5d83a4fcea8d65d95f5380f47572a55965