Pith. sign in

Paper Citation Record · LEDGER

MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 60 inbound Pith citation observations for arXiv:2407.02490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.02490 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 60 of 60 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:45:20.504684Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9b2ac96e-76a3-4183-bfa8-0a575b874f14 · inbound

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling cites this paper.

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:58:29.171156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T09:58:29.057357Z digest=sha256:bc806c01c7a3c657be6aa5647f418193f231eb50e0bf043f00ba95f7b1a3f754

Observation 8e637f98-6e48-42a6-8a2c-3a55a8834366 · inbound

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval cites this paper.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:01.955613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:f7336a97ce545270a6097ed0d02fa9bb2363afc2b880d2dc9b5eda1e059fefc4

Observation a6c4347b-3505-4498-8c4a-ecc4928d63e9 · inbound

DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads cites this paper.

DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:49:16.795888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T11:49:16.654836Z digest=sha256:6347e21a2f1e5689fe12493933ba7ddcf7a39a80aa800057f77b31e68423a49f

Observation 726bed92-0016-47fd-bf64-5646b9ccce34 · inbound

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity cites this paper.

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:02.028093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:46:02.028093Z digest=sha256:1134c901f47dbbc898b4e8a7d20467fce87cadf5aaed4318717a94b958aeafa1

Observation 117ea9d5-a022-4e3d-b118-60286fc7f7ba · inbound

Boosting Long-Context Management via Query-Guided Activation Refilling cites this paper.

Boosting Long-Context Management via Query-Guided Activation Refilling MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:07:30.856237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:07:30.856237Z digest=sha256:8c201f2eccebde94e07681a239b2bfbf9b16fd3661182f9e331b6b1d70acad82

Observation 6bd7d52b-6ac0-4ad4-8121-e7826a92f1f3 · inbound

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression cites this paper.

A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context Compression MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:32:29.429234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:32:29.429234Z digest=sha256:4e224dfe584258a4d2f80d0452ec5dde11d0ef5fc430646ff41bf0e2d93907c4

Observation 9cf4cfec-b40e-49b2-aef5-aeff1eecc7fd · inbound

Bootstrap Your Own Context Length cites this paper.

Bootstrap Your Own Context Length MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:28:38.710848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:28:38.710848Z digest=sha256:b4b3499c56cf39782b49220b10ae338501a06bb9e91f9c9ac4e6f709722a24ab

Observation c70bb0c2-05b7-47f3-a30e-f3e8c6faca67 · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:25.024283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:25.024283Z digest=sha256:2ecbbab5440788fd31d3ba3f70c098b0bd1c6e138f4aaf8604c359b6e7ffecac

Observation 1fbe6fb6-d8c0-41bd-b5d2-ce0472f62f3e · inbound

AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference cites this paper.

AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:51.272353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:51.272353Z digest=sha256:eff8e68a3d3926cc123e5737ae3333c2697e62798b0cf3ba88a65616fa32cd6e

Observation d2582b16-804f-4781-8624-a8375140aac3 · inbound

Scaling Inference-Efficient Language Models cites this paper.

Scaling Inference-Efficient Language Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T00:43:28.945003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:43:28.945003Z digest=sha256:f454577d06032b5dc438dc4a3af55a022e5a0c429723ed7deb14f92eecb5975a

Observation c4f212c6-33a8-4b7f-830c-206048453844 · inbound

FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration cites this paper.

FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:02:30.356815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-23T03:59:58.634512Z digest=sha256:50e2ca39963ffcd137b925dd69c7da6295fef2eb66e9d8baba362a25e6ed3200

Observation cd041ad0-0ee2-45fc-9301-b0f8a83ec79d · inbound

Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity cites this paper.

Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:35:10.472320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:35:10.472320Z digest=sha256:ee91df116ef58486ce8ba832c4c3b10283933d0f9e2d9d8945f292051e4b743a

Observation e538726b-8ab0-46d7-9e42-01e58a067ce0 · inbound

A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI) cites this paper.

A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI) MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:38:33.581299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:38:33.581299Z digest=sha256:4c3af787e3db615f1c63fa35af5031c6b2ff0f7a65de7cbc726263ecbb08a913

Observation 81687733-b76f-4519-b52d-389e12613747 · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.689930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.689930Z digest=sha256:b2c54815272dae6fe7dbc13a6090b20280e7f767d8a34a726cf343f9ef06e18a

Observation 33ca09d6-542e-4079-a12d-02af750dd0e8 · inbound

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective cites this paper.

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T00:45:56.872982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:45:56.872982Z digest=sha256:f962b487bf47f41e7fbf1fee5057670779401075befd5c4455ddd56c6c92c24c

Observation 869c4170-945e-4fed-824b-5686e7ea904b · inbound

APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding cites this paper.

APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T19:28:43.455099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:28:43.455099Z digest=sha256:3f8fdff73d872534c1962b37b5afa3a3c65b3c6d0440c257296a86ef28768101

Observation d83b80b3-601d-441d-b7cf-6f2b72df8a91 · inbound

LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMs cites this paper.

LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T16:40:27.078878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:40:27.078878Z digest=sha256:d960aeb0a818f389c48731f69630dec2daef422699811ed3e21b7b40e0b6baef

Observation 6ed5391c-3d37-4a5c-be11-bce0d9996892 · inbound

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU cites this paper.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.848294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.848294Z digest=sha256:5a2747601e8860379485f1af521fea4f3ed55b37ca9ae4a4cc1865c9446ee297

Observation 681f0673-0b72-43a4-abd7-1ab0cb30250d · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:46:30.161102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:b04df2a6d3f6ffd626da88a4589835be3816e35e78d23a51e4224eac5835fed0

Observation 7277d24d-9352-4285-820c-0cb3f6496af7 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.253626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f0e3aeeb4aeba15c9fda80ae8ff23ad6f307252adbf57ca10432c55792d95922

Observation 9b1d3248-b234-4373-9449-4043757d043a · inbound

LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams cites this paper.

LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:45:20.504684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:45:20.504684Z digest=sha256:5a8482ecc41905582f0f3bf071b082314d2c7f346eb39ed8b17192e8aae6f2b4

Observation 1033f936-1941-4a9e-9745-7e5a7660e1f8 · inbound

Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation cites this paper.

Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.454100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.454100Z digest=sha256:56e30bea2fe6050a772a2b2f05707e990d80448580f28e462de5e3fa5af317a6

Observation 8aa330f3-898e-4217-8508-1b36ea25a20b · inbound

R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference cites this paper.

R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:46.626948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:46.626948Z digest=sha256:22e3a1e5c0bba6ab5402debb6fbef243d0df79a9c4b012daac1c094e57d3b8ea

Observation aeec097a-a6a7-4685-b0a1-af1c8d334661 · inbound

Reward Model Generalization for Compute-Aware Test-Time Reasoning cites this paper.

Reward Model Generalization for Compute-Aware Test-Time Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:27.030634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:27.030634Z digest=sha256:26328c18c3443cdf684fddc32c854b0f9e3a4fb0f758305d40b1f92e115fe82d

Observation 7c98ba6d-c323-4975-9849-0c133478a1d6 · inbound

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity cites this paper.

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:53.290095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:53.290095Z digest=sha256:3a896e62c6ae406dfef91aa6909a5e71548cead6c1a5cdc3641dfe9ca24564ec

Observation 4999cafd-4f82-4a86-8e13-90bd51e734a5 · inbound

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning cites this paper.

SeerAttention-R: Sparse Attention Adaptation for Long Reasoning MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:32.625484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:32.625484Z digest=sha256:3c0d4e65f1a4e4c5352b9d3cb5fa309c37a56fbade9d612541c25402345a5c16

Observation 2f11bb7e-adc0-419b-9367-506515bf257a · inbound

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding cites this paper.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.911487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.911487Z digest=sha256:4f2b80d6b876566c5c5a10811f95901e47739f25410be7bff2b0525a16739885

Observation c35b17de-6ccc-4761-af96-722dbe43503c · inbound

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation cites this paper.

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:45.021883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:45.021883Z digest=sha256:a02ec1989944a13bdf811d75dcf26079c9f3f821c0676d1dcd08155bc47ab2fa

Observation 15d543b0-28aa-4420-97a8-ca925f1e9c4a · inbound

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification cites this paper.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.794754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.794754Z digest=sha256:fc0b72df333ad68cd765b29a406ab885d820287cdfe77b02573fdee529b0ed50

Observation 23b745b4-46bb-49aa-9d1a-daeb32a4bd09 · inbound

Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers cites this paper.

Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:51.617612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:51.617612Z digest=sha256:7bf146b8bcf81613b885fe4350c93ee9d245ba87eb18b7ef749e41d848c532db

Observation adcc9033-8d91-4d3d-87c7-087cc89f694c · inbound

Dynamic Sparse Causal-Attention Temporal Networks for Interpretable Causality Discovery in Multivariate Time Series cites this paper.

Dynamic Sparse Causal-Attention Temporal Networks for Interpretable Causality Discovery in Multivariate Time Series MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.953281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.953281Z digest=sha256:5b7341d4c5e8e688137ee159adc8552e4a108fc0d5c115cde014bf9933f7a6fc

Observation df5ba07c-3d6d-494f-b1f5-228867137817 · inbound

Synergy: End-to-end Concept Model cites this paper.

Synergy: End-to-end Concept Model MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:34.035954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:34.035954Z digest=sha256:bf394c6f6bf84c645d74ea23f0e15b091de411699175578a13a1950f55bc5d12

Observation 2e6b9a98-e7c4-4f03-8bc3-bd9a83e5d865 · inbound

NABLA: Neighborhood Adaptive Block-Level Attention cites this paper.

NABLA: Neighborhood Adaptive Block-Level Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:45.327714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:45.327714Z digest=sha256:cd0ba5e2a842f3ef9202b44be4f39bbecfbfd2921eba976087cb07e966a985e4

Observation ada1c4d4-38ac-4c43-a5f7-507ee13b7a3d · inbound

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference cites this paper.

DeltaLLM: A Training-Free Framework Exploiting Temporal Sparsity for Efficient Edge LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:44.377690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:17:44.377690Z digest=sha256:20c5ccf38e7124ba3b8e1609f116719e0758046f62429da44d442cb9a93aeeef

Observation 65217528-7120-4a7a-b774-60cfb657162c · inbound

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference cites this paper.

ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:06:52.082431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T22:03:10.316005Z digest=sha256:10b3ce511a862d9592807ca3e87d8962a554be33a1b115577884e8c5c404a15b

Observation ca52ff46-90bc-4a78-ae48-08d8c6da7500 · inbound

Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety cites this paper.

Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:54:14.367710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:54:14.367710Z digest=sha256:2f8da8c9a77d33be757ed9e0b10bbe2412011c9a0a57e7482ace952ea6a347ef

Observation 87954259-ef6d-4d24-a4b2-7bf3539ab4e2 · inbound

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs cites this paper.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:30:55.155139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:673aab4a496f0407a6fbb7313da300f28794ac5c94deeba06f40513804abeb33

Observation 691b9a4c-d2e9-45d6-876a-a9f44abdf1c1 · inbound

Prism: Spectral-Aware Block-Sparse Attention cites this paper.

Prism: Spectral-Aware Block-Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:22:53.923474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:22:53.923474Z digest=sha256:cc896bf47cf646d46536dbf9d783828e5e27c9b9a2510cd87e694d6d30308e44

Observation 83f6ced1-9abc-41e9-a942-c1eb01f4f7dd · inbound

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems cites this paper.

PrefixWall: Mitigating Prefix Caching Side Channels in Shared LLM Systems MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:15:06.744912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T12:14:09.509302Z digest=sha256:12a1e9b6c6ebe778b22081f44e0134f6fe5dbf926d30a9abdc8f6ce5f0a99d7f

Observation e8fe1080-28b1-4054-957e-2b2ed372cf7f · inbound

Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions cites this paper.

Three non-Hermitian random matrix universality classes of complex edge statistics: Spacing ratios and distributions MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:19:49.419805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:19:49.419805Z digest=sha256:273c4dbb6fc396d14f99ef45cd79be3ae1032a0a5f65e0eeb49b53ac78eb75a7

Observation 5b77b6de-5270-476c-8eec-3f8775bcb0e7 · inbound

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention cites this paper.

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:43:00.751161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:40:28.854082Z digest=sha256:821378f91a502fc124c3101f5569b08a4ef50c8c0c56e776c2f6383e982a2840

Observation 371997cc-fc2a-45a3-a6fc-09deef0a4244 · inbound

StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k cites this paper.

StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:15:37.731448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T18:44:42.456111Z digest=sha256:1d69e3d25f7362067d4d6e258467a9e6ac997ff4ef5c3f97ef03026c66bc494b

Observation 3bcb5b8b-a10f-4efd-8c81-1b156e449bba · inbound

Long Context Pre-Training with Lighthouse Attention cites this paper.

Long Context Pre-Training with Lighthouse Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:09.397118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T10:10:38.610613Z digest=sha256:dbe1c021c9968b7084adfbdd1a09e557b5735b05a575105e2389ae1a61dfc4f3

Observation 6e946a6a-3acd-40b1-887a-3f10d559de34 · inbound

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache cites this paper.

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:15:50.429219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:15:35.871863Z digest=sha256:3d4da508357a365341bdb8a043e8fc6360ae81f6895793799ac6e0818342c573

Observation 44cbbdd6-45e9-46b0-93d7-d27c7a5d104e · inbound

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention cites this paper.

DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.547437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:50:12.926232Z digest=sha256:39dc02189f931ebfbaf63c22df42430365b9705fc3ba992bb0512f056b5c86b5

Observation 82c0a072-b8d5-43c0-a43e-a8a86834bd94 · inbound

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference cites this paper.

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:03:59.863178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:03:07.685253Z digest=sha256:d4dcf3178f3dd56b88a7b8fbfbfe0ae53ae15329a86267765e5dadb908f5ce90

Observation f543456c-1afd-4396-ba3f-1c23272a88e3 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.749346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:0cb82694e963d896a665f78fec82071fffaa58d1ba44aa4dcba637b54abaeed3

Observation 40b9c12d-7e6f-48da-b1f5-6d1606fbe75d · inbound

How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models cites this paper.

How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:09.159889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:52:13.337419Z digest=sha256:009dfaed52307b88ad1bbdadc75380a3c54518fb01f5b5d78b1f13875916069e

Observation 220dc4c9-14fb-475a-915e-5bbb705f6507 · inbound

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance cites this paper.

SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:31.268022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T16:24:31.109508Z digest=sha256:7a17b1b5bfc3d4417c0316afc6b47672280bb6a3025ad26a540c515a93a78596

Observation e4a10a0d-e597-4083-a80a-1d11a8bd2bbf · inbound

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs cites this paper.

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:01:07.726301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T16:57:02.709967Z digest=sha256:7d924c7bc0738a479bd632c2bdcca8e5ee060451eaedc645305710b119cfd006

Observation d46ba56a-1e4d-4629-a0b0-f65436b9c3a7 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:09:36.976767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:ed7ece37074eaaa6bf3565fa4c6580949df999c87cb5dc4e6465cad0d71a6efe

Observation 91bc4a43-88c4-4dc1-a536-759e80b6eaf9 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:13.460701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:13.460701Z digest=sha256:135e26844bf68fa90672cfe401ad021b9476e8f4c7debc9e9c864a7ae1fd3308

Observation 46920f48-52ae-436b-9d5a-203aa039cf3a · inbound

NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation cites this paper.

NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.719634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T05:00:15.336703Z digest=sha256:1e33f883d63eda43270262f6d6d5b428611c5fac00a88995e542cc2426eb8ecd

Observation 517cc6a6-ee9e-4000-bfeb-1fec7201f700 · inbound

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM cites this paper.

Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:24:21.385249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T07:19:48.530272Z digest=sha256:7ff924584bdb5bf8e162e84c434686a54d5f67df3680e9842ceee6d1ca182ebe

Observation 8ea4c474-9722-4fe5-a252-a84a522b42ae · inbound

Hierarchical Global Attention (HGA) cites this paper.

Hierarchical Global Attention (HGA) MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:55:28.774424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T06:54:58.841989Z digest=sha256:9dbf85c80cb92372450f5654b97dd908f60621658470182627009910e7ba282b

Observation 0fcff984-9533-4404-b38f-594b5faea618 · inbound

Uncertainty-gated selection for block-sparse attention cites this paper.

Uncertainty-gated selection for block-sparse attention MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T22:15:14.580916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T22:15:14.580916Z digest=sha256:deeb4d23af677bf71677808209c0a79287e25e3c1c34daea159080740cecae31

Observation 45063632-881f-418d-8182-17cf68216e57 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.231104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:f47b05667d77213e81d4dae4da3d39a6a8a9d4567381785c8bed2a6158e6aa40

Observation 8f387e1f-e50b-47f9-88e7-abfa0bf8fff3 · inbound

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems cites this paper.

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T19:52:40.866676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:52:40.866676Z digest=sha256:453a340ed1dc0cb9e94a0b8db7d093f205a82ed63e00e10e2945dacbc41cc73d

Observation 6a485c77-41fc-46fd-a732-9357902f67c8 · inbound

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference cites this paper.

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:03.978556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:50:03.978556Z digest=sha256:f9745b3b088ecccc36c1475dc5ac37b5e2f8eceb731b015ba1f64316aa6aff8f

Observation ac14ab7f-d113-4e76-9b76-816b3b767422 · inbound

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching cites this paper.

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:32:37.605595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:32:37.605595Z digest=sha256:e743bf2fc5b64b57a21af456d847111cbb92e685a588856fef7a7c6d66799f1d