Pith. sign in

Paper Citation Record · LEDGER

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

As of 5 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 30 inbound Pith citation observations for arXiv:2409.10516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.10516 v3

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T08:12:01.798459Z

measured 130 of 130 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:02.786626Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact12
  • verified fuzzy69
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 73377b4f-3a13-438c-a642-8cbd085b289c · outbound

This paper cites Scaling Learning Algorithms Towards.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Scaling Learning Algorithms Towards

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.302095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:ec10fdff034ac29e532f60e8780a963e31f74375f5bfa90cfb9e60d81ac5907a

Observation 1a089186-eab7-4b7b-92f1-5729d0385891 · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval and Osindero, Simon and Teh, Yee Whye , journal =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.311323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:210c75dd6b6d71dd31167b9efd522be032bb695bb3b9b8738fb3df168db12c3c

Observation a5a5c367-551b-4164-bb4a-d682582842c0 · outbound

This paper cites 2016 , publisher=.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval 2016 , publisher=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.315146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:d4449c75d09770752c27c5bf105fcfcc56353b26ea600e3762a71663b95c8196

Observation 4770016a-b290-44c3-ac4f-3bf5ed777ecc · outbound

This paper cites IEEE Transactions on Computers , volume=.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval IEEE Transactions on Computers , volume=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.318420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:90bbe7f96373095a0641130fb78408f9264bacb665ba7507e2448f0cd9d1e8ff

Observation 1eb601b0-52b1-4549-be1d-bd86d0ea179c · outbound

This paper cites IEEE Transactions on Big Data , volume=.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval IEEE Transactions on Big Data , volume=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.321872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:a121c959b3bb309ca676525fba47d0cc0118cb08b27c8ff65c4a151cb9fb9938

Observation fb833c48-76a8-4886-b567-fd6b592e4807 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:45:36.571697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:29c3a94a2feb4e9bf1e7767768967254aadb8c8bdaeeea7606b3173ad0828745

Observation 8b466e0c-38b5-4a88-ae52-218beb04d046 · outbound

This paper cites OOD-DiskANN: Efficient and Scalable Graph.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval OOD-DiskANN: Efficient and Scalable Graph

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.327034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:cdd761f3ae9362d3eb6d8783517ba25f802357641e804759d2f789c5007453b0

Observation 6c4bf14a-0a2f-4ae7-81df-42c4df7d4c1c · outbound

This paper cites CoRR , volume =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval CoRR , volume =

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.330821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:d9c7f38d39a0af3f52542a33a0fbbd57290a8c2f5c83a781ab9ea26b10fc5aff

Observation feae6a00-4d95-46ab-b168-e16390f6b7c5 · outbound

This paper cites Gomez and Lukasz Kaiser and Illia Polosukhin , bibsource =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gomez and Lukasz Kaiser and Illia Polosukhin , bibsource =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.334994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:b654fae2f2d09c28388a5a8ed603beb36b0d9b44fb8ae7743376e247fcfc8866

Observation a3d8e2c6-0682-49a6-ad39-3f9f9f04647a · outbound

This paper cites A weighted nearest neighbor algorithm for learning with symbolic features , volume =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval A weighted nearest neighbor algorithm for learning with symbolic features , volume =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.338991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:1e6ebc6ba287117495e59d8685cf17763cc40bd6110dcd94d583a0e9bfdf38a2

Observation 68bef60b-9c91-4749-b434-cda636492075 · outbound

This paper cites seed , volume=.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval seed , volume=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.342638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:e90372394b9846fba679a0ae7d0472809e0d54f6020462da09819b63582e5ccd

Observation eaabb11d-b223-4045-a11b-876d5acb14a1 · outbound

This paper cites Query by image and video content: The QBIC system , volume =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Query by image and video content: The QBIC system , volume =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.346430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:35f723621d4274178596d199cc7f062a8388bcd2b28db2fe8cfd0d522ec8abe4

Observation 52c5614c-469d-4713-9da8-b377c1b7e5c2 · outbound

This paper cites PQCache: Product Quantization-based KVCache for Long Context LLM Inference , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval PQCache: Product Quantization-based KVCache for Long Context LLM Inference , url =

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.350936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:e2a8859a7067ad8851a5b61a9f1db342f4fada104f1dd2277e7645b7efbcd172

Observation 681e026c-d1d3-4d37-8881-d80bf084cafb · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.354825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:1a7fb3bdde40e1b5a1bd4f7d0571ed152c6caf6a3d3e2cb87687ca5ed3930ba3

Observation 98d952c5-9c26-47bb-b2e8-1ae74c794f5f · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.358099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:94a97943a68bf683d18f10be4e63a685a92da6b942dd2ab782f5b9960904c1b0

Observation 480d7388-ff99-4d8f-b3f9-5249d13b4eb5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T08:12:01.998159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:f8a3a496fb7410819837ba566d5e8c2d215cfdba3d683f41e2b869e4f546e59a

Observation fc65432e-90ea-4a2b-a055-9665cb4d7c1b · outbound

This paper cites The Twelfth International Conference on Learning Representations.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval The Twelfth International Conference on Learning Representations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.361802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:67ffcc79ce85550e9a996b4b9b20311c27bfd57c6065025f442a04dc9d141839

Observation fc2b2041-fb93-4cef-bb14-1326f0f33abd · outbound

This paper cites LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:01.900718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:6fc8dd949d8421885371f3408b40c00b264bf2a03248a4f5024840ba53be88eb

Observation 0c948570-1561-4632-8e8f-813da7fdd902 · outbound

This paper cites Sean , doi =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Sean , doi =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.365074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:c3b235e4887665f06461f71ebe72b66370fdb752218dd901777db0dbbb782d77

Observation 8136a523-356d-4264-8f3b-2d40edc397da · outbound

This paper cites Reformer: The Efficient Transformer , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Reformer: The Efficient Transformer , url =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.368907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:67c7d21ccd3c0964739ea0b33f38992f6645de7da8af618c872c2d541be1ab53

Observation 532345c6-e323-439d-a52c-6b6404c76e8d · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.372011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:2cd4de19657c2f9c46b26ecf4a5be5f70812910615e25f730cdb173e973b66ac

Observation 05401a3e-f480-4eeb-8c23-f3bc50dc22a1 · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.374988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:d2dd6be2f8adbfa697f8262450939b88facb6153dc5675a0c01a2ef46159d153

Observation 5f4d7360-0853-4764-bd73-9e654bfeb271 · outbound

This paper cites FlexGen: high-throughput generative inference of large language models with a single GPU , year =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval FlexGen: high-throughput generative inference of large language models with a single GPU , year =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.380262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:1392b2859167fae21374e1d4d822396ba4bf5ac0ecf2cd5886df1e03201293e1

Observation 2f5e2b5c-0adf-4a7a-bbc6-8c324166bae3 · outbound

This paper cites Model Tells You What to Discard: Adaptive.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Model Tells You What to Discard: Adaptive

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.383476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:1d97e507b0a665e6e9b39814f437300b39b70037f94b991d248bb288de063dd7

Observation a5cc8c76-d3a5-444a-a66a-316f6e25d7cf · outbound

This paper cites RingAttention with Blockwise Transformers for Near-Infinite Context , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval RingAttention with Blockwise Transformers for Near-Infinite Context , url =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.387295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:9ee8196284b7e9f5b948f5d84055646b97633a3c85708bc4970f609abb9c8098

Observation 10213dac-214f-4e5d-a057-f7a61a86d039 · outbound

This paper cites and Ermon, Stefano and Rudra, Atri and R.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval and Ermon, Stefano and Rudra, Atri and R

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.391085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:91e1e082c3f9ce91d2ac00087a00396ebf21a401881eed34483c3c6feb736fc4

Observation b945c80e-9091-46e9-abf7-51f6f0641b8d · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.394285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:9b5c67a80c814b74831efd9576399dc3488f2cba391741cbb8fc5b90c5fe49e3

Observation eff34419-a984-4160-ba45-8e5537f4592b · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting , year =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Splitwise: Efficient generative llm inference using phase splitting , year =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.022609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:5b7e8f83230f80994cc36587848105171ae7e7d00852dced1d9a3c3002f9293f

Observation 46767789-bc80-4e7b-a090-066f3c90e449 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks , year =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient Streaming Language Models with Attention Sinks , year =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.027529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:e11cb9c43d1872cf8a9a00bfaa796ce1fa2855abfb1fb0a1288cfefd04c10ddb

Observation 3d72fdbf-9b5a-45ee-9715-17f6b33d85c2 · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.031968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:7ea7bc8083a21fb6da389743475767e670d3fe907e88068add1af03c657b36fd

Observation d0f3b649-679c-47af-89a5-3bc5c3d58726 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Snapkv: Llm knows what you are looking for before generation , url =

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.036543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:fe50a588e9f0422cc9055e87d2967282420ed78e905971959d713f66081f57fa

Observation f4e29f8d-0352-48bb-aae7-583e9f7f56d2 · outbound

This paper cites MagicPiG: sParse Inference enGine for LLM , year =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MagicPiG: sParse Inference enGine for LLM , year =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.041571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:bc5f1065763bb939f5d0fc751fd66dae6bf90e2a9e74fb3a401248742e018d26

Observation 2541b9f8-7978-4b7d-88f8-70bd248259d1 · outbound

This paper cites Longformer: The long-document transformer , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Longformer: The long-document transformer , url =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.046074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:661b02f273b96c1b27e00281c6d8593de928d685a253f20b3f8ca031a0484a96

Observation 3d4ef521-4f59-4f76-8295-7fede33ec93c · outbound

This paper cites InfLLM: Unveiling the Intrinsic Capacity of LLMs for Understanding Extremely Long Sequences with Training-Free Memory , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval InfLLM: Unveiling the Intrinsic Capacity of LLMs for Understanding Extremely Long Sequences with Training-Free Memory , url =

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.050719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:ae67e2e07b1d045d677fa45c464e602fa0006ce0f7196309bd826c6ca5877c85

Observation 34979964-5b4d-44eb-91cc-e8261a15978a · outbound

This paper cites The Faiss library.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval The Faiss library

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T08:12:01.914175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:01fab85cce0fc9e36a822165e76ffc0087e4740fa9a176f5776c99c19f5b2f2a

Observation af1fe93a-08a6-4f1d-af80-459512bb79db · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models? , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval RULER: What's the Real Context Size of Your Long-Context Language Models? , url =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.054769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:72647b9093ced3a754c26f3853cc72b2aa6983abcbde3ebad85841c2c477924b

Observation f6c9a0c0-0791-4577-801e-6aa7211577db · outbound

This paper cites On the generalized distance in statistics , volume =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval On the generalized distance in statistics , volume =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.058693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:cb7b5091d8c9660f2a8a4ca03c4566b582b3bc94e8a7f90e0432c9cfb893717c

Observation d8b22e7f-c204-4990-966b-43cc39219d44 · outbound

This paper cites Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs , volume =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs , volume =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.062929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:a859a4ff9da5d0ed20fe118868cbfdf4862cf2ab771c743e4c38c9ce02110fed

Observation e442ee8b-7adc-43dd-b6d9-5e75a661b96a · outbound

This paper cites Lempitsky , bibsource =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Lempitsky , bibsource =

Reference 40

Resolution
verified exact
doi, observed 2026-05-18T08:12:01.882999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:56aa5e21d3b789834e73b5f01288a56878551ce59e93058442fd0420a2538a5c

Observation 4542df4b-a164-4c93-bb25-1d81b096be6b · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.067251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:66ff76059b74dafa60ae203520bd249ef37de504da3337bd5aff410f9be64fc8

Observation 8f98e092-96f5-438c-acad-8377aa688dba · outbound

This paper cites Video Google: A text retrieval approach to object matching in videos , year =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Video Google: A text retrieval approach to object matching in videos , year =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.072624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:e5f7b956196d9e1327f649bac5791d708b6074c9ca0937d5f20fb0ba69620d50

Observation 2509cfc0-0aeb-419b-aeef-a78995630b96 · outbound

This paper cites Big Bird: Transformers for Longer Sequences , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Big Bird: Transformers for Longer Sequences , url =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.077857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:aea7d6ece2ec98d59fdb907d71430c2503bafa58c2fd9e198382a0c8a3cf54b7

Observation 745738cb-b938-4ec7-af03-a2080abd3512 · outbound

This paper cites IceFormer: Accelerated Inference with Long-Sequence Transformers on.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval IceFormer: Accelerated Inference with Long-Sequence Transformers on

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.082232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:60d83bd8856750311b528b80464d54b9cd02518977063bcdee73c5ab6d7226f6

Observation a98d3c41-c6fc-4b55-b114-5464cee32fc1 · outbound

This paper cites Generating long sequences with sparse transformers , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Generating long sequences with sparse transformers , url =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.085989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:7d13b002669b43f1726e7fbcace45a9fb48730f46988da494b7093aa16028461

Observation 0ba553ff-7eda-4e11-88f4-d91db2f74560 · outbound

This paper cites Keyformer: Kv cache reduction through key tokens selection for efficient generative inference , volume =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Keyformer: Kv cache reduction through key tokens selection for efficient generative inference , volume =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.089866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:888fbc9c6095ec7a7acca5d8ed9035a5d898c8e0e8697cc84dd25a3fca3217cf

Observation fa01a263-ffba-4f9f-8243-f901132f938a · outbound

This paper cites Unlimiformer: Long-range transformers with unlimited length input , volume =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unlimiformer: Long-range transformers with unlimited length input , volume =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.094013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:993e7ed23ed646cc9189678fe5b47d93a6cecc56ac7850cdfd7e629b815b3446

Observation 5cb10304-9f9b-475e-90b5-8ff054b7368a · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.097620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:d347255f00f8b9dac8a31edc7f1103171aa2f3f08503be4121a7d0475398df62

Observation 53eeda90-c923-431e-855b-5c321770818f · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval SparQ Attention: Bandwidth-Efficient LLM Inference , url =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.101599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:9eae8f28c3bef7e6a5e84bec1466548d849dfa516ec7d64cd947206d258b9873

Observation 5f9c2d87-0097-4d12-8e4b-5a18769ed53b · outbound

This paper cites Efficient and Economic Large Language Model Inference with Attention Offloading , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient and Economic Large Language Model Inference with Attention Offloading , url =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.105381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:e2c83cea1589fa8512755308f5d2364e5867484e11b14b8c51186cb4ff0ade71

Observation f8139999-860e-43fa-a8b6-40f31b3d778b · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time , volume =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time , volume =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.109400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:f935ff69d2ee3eb5596bb56daca7f8cc6f91ef662e30a94a842e6d2322f8e899

Observation 65b1e0f9-32a7-4501-9ed6-ada258b8ce49 · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.113611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:59cb05da1b609fa61852496a453257363f81c7a55ff502106e96de6d1763e3fa

Observation 96a4dd63-b7d4-401b-9466-4e707bd9b154 · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.117249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:ae8d55c69e835558849e180f835d15960ad545c244f5de0c9df23ad790d6ad0c

Observation cef76df7-22ed-4f78-8e8f-0895d280d610 · outbound

This paper cites Loki: Low-Rank Keys for Efficient Sparse Attention , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Loki: Low-Rank Keys for Efficient Sparse Attention , url =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.121470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:47e01e02ca1e1ec6f23abb14d99eeae6542a29b7908fce066b0511bb3b08c06e

Observation ff46e5bd-3022-455a-abb6-3b5e9875e97d · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention , url =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.125463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:8e0ea437d26a325ded8e59f73873dbd2977fcd01adcd3ecfebfebde5e9af65cd

Observation a2eed4ae-aa10-49a8-8db0-b7c86e4baa44 · outbound

This paper cites Mooncake: Kimi's KVCache-centric Architecture for LLM Serving , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Mooncake: Kimi's KVCache-centric Architecture for LLM Serving , url =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.129485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:bbf159fce91db6f418da0b7d4700950b8bd698f09fea2ec6606db9d72f53e105

Observation 06e5669a-1cc8-4592-bce8-a156da499688 · outbound

This paper cites an unresolved cited work.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-18T08:12:02.132961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:18ffd125ba0387b9c87a8ade96f160323a13cac193e3a0a04bf033eff5188aec

Observation 3061713c-b5fe-4490-842c-3b0d7b0bb4ba · outbound

This paper cites Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models , url =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.136812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:3ed204a89313464cf3608902c783b47858eacc96791729e73e26b7eb0237f57a

Observation 74eb02e0-e6c8-4780-be1b-fa0dccefba6c · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints , year =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints , year =

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.140417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:7ff08b3de965a0891b8c95923b8d010c3ce265c7b06126bbcd51bbfc6c42e384

Observation 3fd397b5-19b7-4cc1-944b-d153fd9e773b · outbound

This paper cites Gonzalez and Hao Zhang and Ion Stoica , booktitle =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gonzalez and Hao Zhang and Ion Stoica , booktitle =

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.144420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:007df6dde69c9be9df24ac24451d3ad87e70db43312c5c027a03f64bcbfab31d

Observation 6b5812cf-0fcc-4015-8a02-991f2e04b931 · outbound

This paper cites How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:02.011756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:773ecdd0b4310c9914eb6eb3d5cf8dfaa3229613cbc8bae777335e3fceb4bca6

Observation b8722174-dc9e-49be-8b6f-32734276e539 · outbound

This paper cites Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval , url =

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.147790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:864b68b01b100b8ed9fad929794debfef653c84fb7713b132a27e61754e6ce31

Observation faafad50-c5be-410f-bc14-5a3bd69c8c8e · outbound

This paper cites PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest , url =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.151330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:28e21623c1bd9ff6ada935f75bd681622f5f103ca948ad269d310e4437d36393

Observation 21bb8f70-62f2-4ce7-a1d2-0004b35dd5e3 · outbound

This paper cites Non-metric Similarity Graphs for Maximum Inner Product Search , url =.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Non-metric Similarity Graphs for Maximum Inner Product Search , url =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.155121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:fdb243ecdf343b3ca105642783f019b48cbbdf30a10533b1fef74554af289c51

Observation df2a457f-61c3-4c4e-b1da-17a74a43d217 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T08:12:01.991620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:c53639883799a3cc8beeb7ca7c2ca04b2580ff8539a3b4945778bb46ea7e1f9e

Observation 42e456ea-3863-4379-8c0d-3eb35e249275 · outbound

This paper cites Yi-6b-200k.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Yi-6b-200k

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.158801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:500ad2b1ae90b525c6a5dd761fd09e7d270e65fccec79f907177d61757fdadf5

Observation cace07c2-64ad-4600-b846-d56752dbe1ec · outbound

This paper cites Yi-9b-200k.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Yi-9b-200k

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.162548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:7dae2971e17c6654dc2996128c6441e920b831fb7276893ad1eb0a114a3fa4f8

Observation fce6e1f4-0909-44ef-b5de-d21713e24e2a · outbound

This paper cites ETC : Encoding long and structured inputs in transformers.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval ETC : Encoding long and structured inputs in transformers

Reference 70

Resolution
verified exact
doi, observed 2026-05-18T08:12:01.868323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:e8992d1f441661c237664436c922f8318ef57734e46b68643faa8a6c1133c00f

Observation c001afe8-7762-4f13-88e9-592d3789f9d8 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.166476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:c336569314a04cc92e9b17417bbae5a6e302cb798f33272186bebcd63a1b9978

Observation 4fb1939a-6070-4b4f-b8ac-a5d3627aacba · outbound

This paper cites Longformer: The Long-Document Transformer.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Longformer: The Long-Document Transformer

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:12:02.017700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:c49a10417b50e430d453d530014680d4b5ad197ab81458ae47818b327bb5ac02

Observation e70b4920-e52a-4c13-b629-1af851b36e94 · outbound

This paper cites Unlimiformer: Long-range transformers with unlimited length input.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unlimiformer: Long-range transformers with unlimited length input

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.171029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:dce10b38779a33ce5a41779513f3fcd45a4eea8f925ad36ccefc22a958d1b5a0

Observation 666adf94-9fe6-462a-817f-4cb202cab9d2 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:12:01.920496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:2c373c28462af8923c8d178bfb56c19933478762af47dd2bba42a3a587537455

Observation e74fcdfc-ccbf-486d-a42a-fc3c4817577a · outbound

This paper cites Sean Wang.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Sean Wang

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:01.860771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:8fa131e9742b85f39cce9ab9b33010a435b1610d387c7f061956da1a26600f45

Observation d42c4ab1-d012-438b-b604-2a493f39504a · outbound

This paper cites Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:01.961981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:b60c1e2232f070202f0ed1c8876e8ec4b4b4d1e646f0f3137f936ad558380705

Observation 53a10c02-1a5f-4b5a-bc19-15d08887af1b · outbound

This paper cites Magicpig: sparse inference engine for llm.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Magicpig: sparse inference engine for llm

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.174975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:c70172afea89551c8b9ef7a85b729490fa6ae929202e8c298937b01c255a4769

Observation 64ba7080-2156-4d36-ab98-a3ae8415651a · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Generating Long Sequences with Sparse Transformers

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:12:01.985977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:fde7e8a4f764d140f7c9e58bdfce0aa0ed3c0e2df08a43b51bccc7cc4bb35e19

Observation 737ecc75-17e6-40af-8023-27c62719721f · outbound

This paper cites A weighted nearest neighbor algorithm for learning with symbolic features.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval A weighted nearest neighbor algorithm for learning with symbolic features

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.178971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:c34a7bcb73d889f94ffde1d4504566ed1284c7a0f5f6ad0e58a43ce002613412

Observation 419ac5d1-8298-41ff-85a9-a111fcc09f28 · outbound

This paper cites URL https://doi.org/10.1145/ 2959100.2959190.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval URL https://doi.org/10.1145/ 2959100.2959190

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:01.877928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:c5a0b0fd0e2b63cc9d627cd93a24efcf9a0e8a354f8673dd890e37d90357e124

Observation a9b44847-da35-4ee5-a6bb-3af273dba721 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.183183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:9f661f288062edf1b5e70d7500c2e1da605b3e2a3d03b20c815d5e82e0dec9a3

Observation 9a160637-9dac-40f9-b228-861aed039381 · outbound

This paper cites Attention is naturally sparse with gaussian distributed input.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Attention is naturally sparse with gaussian distributed input

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.187389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:44fbcbf676c5cbdb6c46ea106b4b160be9071c5197aada9e6efd6ed4edd46335

Observation 48ff9088-b825-41c4-9571-c07356e69a1b · outbound

This paper cites The faiss library.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval The faiss library

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.191335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:acaaa9c6f93057494e49d7406b5c3ff1eda785387db9e471daa5a6c5f8233fb6

Observation 296f2418-3c36-41d6-93fc-44a00447feab · outbound

This paper cites Model tells you what to discard: Adaptive KV cache compression for LLM s.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Model tells you what to discard: Adaptive KV cache compression for LLM s

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.195099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:91cabcbb5c0d4454c3758f0da34388ca7cc6fe7b8fc820d5c141a57b80ee9909

Observation 17bf77dc-5406-4196-afcc-c97f412ec7a8 · outbound

This paper cites Context caching overview.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Context caching overview

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.199201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:dd04dbad31107539c8580fbb75b069c6f71792103cabde71793d7be87e59a79b

Observation cd52da36-d1e3-439e-b933-908c14493742 · outbound

This paper cites Llama-3-8b-instruct-262k.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Llama-3-8b-instruct-262k

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.202739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:8e03ac0207a066621a2929f91eb54f5c08e0b53e612e140024e84421b0646d5a

Observation 02f35180-7252-4d75-9912-41a821d4638b · outbound

This paper cites Needle in a haystack - pressure testing llms.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Needle in a haystack - pressure testing llms

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.206494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:cb7d092f8eeb1114809e9f1265c128ddab68c7083098d8df2cf514875aa23d37

Observation b97d136b-87e2-4af4-b22b-08576488198a · outbound

This paper cites LM -infinite: Zero-shot extreme length generalization for large language models.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval LM -infinite: Zero-shot extreme length generalization for large language models

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.210121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:ef9535baf4d218ed5de3cdc90a902b5e4adddcda1da9ed15a55b355ccb67dc01

Observation 668d82fd-a21c-4c5d-a307-c768301643ee · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:12:01.927648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:9bd1a6c5c8f84b43c4b2228690ea723e0a6ab05f389619158f318af159822ff7

Observation ded72ace-d1a3-4e3e-bf3b-f6b0156cd3d8 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:12:01.934048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:cadb319354ee3208e8bab35f30077f250c1ae05290581cb7007ececaa5ab9467

Observation c074ab33-5a65-427e-b9fc-8b5844b8a4da · outbound

This paper cites OOD-DiskANN: Efficient and Scalable Graph ANNS for Out-of-Distribution Queries.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval OOD-DiskANN: Efficient and Scalable Graph ANNS for Out-of-Distribution Queries

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:01.941972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:7a0d7869338231c0f1bafaccc313c5199dfad664d14bba6a9da939b047e35dd4

Observation 8e637f98-6e48-42a6-8a2c-3a55a8834366 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:01.955613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:60f0a3ab24bb21f8b8920a3aa58eedbf5241424fceb3c926d881d308b5a71d1c

Observation 1a24bd02-ad16-4266-b011-9c8c2f9f01c6 · outbound

This paper cites Reformer: The efficient transformer.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Reformer: The efficient transformer

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.213887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:e76a29abcfcd884eb9aa731233d077c091811c840056459e97fbad6a1c0a06a0

Observation 7b9b8996-e134-4302-be8b-32698285da64 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gonzalez, Hao Zhang, and Ion Stoica

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.217934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:1f89153a223f73b5889f663aca054ad580ca6a95d47dac2227d99c43bed93830

Observation a974e6a3-8eab-4934-a3af-bf4b671125a5 · outbound

This paper cites InfiniGen : Efficient generative inference of large language models with dynamic KV cache management.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval InfiniGen : Efficient generative inference of large language models with dynamic KV cache management

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.221681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:b26f9e4b32bd0c8a584605ab3211d82505e1787a7493ac2f0c5aba6b39838455

Observation fb2b0e7e-821b-4fc9-8a70-7f0911754172 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval SnapKV: LLM Knows What You are Looking for Before Generation

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-18T08:12:01.980850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:25bfb93fd6607d6ac5b766b62966274dbf9c317ab2d45f2e59756eb1e062c30e

Observation c38cf58e-3af7-4ed0-9327-6a5c5f0463d4 · outbound

This paper cites Ringattention with blockwise transformers for near-infinite context.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Ringattention with blockwise transformers for near-infinite context

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.225348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:3f1b357c8c6557bc258742a55d4219f6b0164f35806041dfc29ae6ff0d6f254b

Observation f7e558ee-4788-417e-81fc-d064ec1f40cd · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.229807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:544311505c3a79acb3df310380c8dfa7505e54b434affe25f6491a142ba84ac0

Observation 1e8f0203-f860-4066-bf27-fec2c4dc8e00 · outbound

This paper cites On the generalized distance in statistics.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval On the generalized distance in statistics

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.233775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:6ebf2beebc6d518303bf1c520eaf77eb4f107aa5c77835c64593bde543b224fe

Observation 23807758-2930-4f62-a3a4-1f5d639109af · outbound

This paper cites Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.237499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:844064368d07e44009d681306773ddc70d8a8a6449c1bbbc7a707ce0c3df3a94

Observation 96cc7b0e-03d9-4440-b84d-56afb6ce0902 · outbound

This paper cites Iceformer: Accelerated inference with long-sequence transformers on CPU s.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Iceformer: Accelerated inference with long-sequence transformers on CPU s

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.241129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:9ead61022ce61edda6276b61b2494517cdfbfcab4d68c3154cef382fa1d917d4

Observation 6294af34-9e02-4a4e-bd99-1c238b2cab26 · outbound

This paper cites Non-metric similarity graphs for maximum inner product search.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Non-metric similarity graphs for maximum inner product search

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T08:12:02.244905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:358acb7bcfc6143c8b23c54ff98739c6e280f4c6f39a815da22e2b567e856dd2

Observation 8064e1cc-636f-403f-a912-a0e42e407449 · outbound

This paper cites Ghosh, N.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Ghosh, N

Reference 103

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:01.892660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:c896a493abb471d5aed9a22cac89f4aa20a62d0026277bd8d8cd599840d51987

Pith citing papers

Observation 0e5c3793-50d4-462e-903e-5c8340166dfe · inbound

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference cites this paper.

Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T11:16:31.904921Z digest=sha256:64f62ee2b975e6be0abbff645872a3a889a780a9226d4269655b6f1e995cc96a

Observation 3e10dd34-8e53-4d49-97a5-93813fe8839a · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:ce7d9c2802c771e287f43348c4c9dce5342546007e00664431d96cc4c54aae74

Observation 37d8729d-fb77-4706-967c-f7347e2aa8e0 · inbound

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference cites this paper.

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T15:59:04.724780Z digest=sha256:13f3b19cd065822a5bb538688a1dcb419289c4715268b8c8ea7eb54bf22b0f3a

Observation 93e26f7a-290d-4965-92bd-5f69f6d8c844 · inbound

Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models cites this paper.

Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T14:12:23.982580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T14:07:42.408261Z digest=sha256:f103bb20779a854a9607b8f023224b0129926a8071166edff04af61fda60b9b5

Observation 577dc9ca-44fd-4f3e-ba4d-8fc50e49423a · inbound

MemOS: A Memory OS for AI System cites this paper.

MemOS: A Memory OS for AI System RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T08:20:22.658329Z digest=sha256:e92a680997d94be6a2ff7d9a371d12614f651f37972058bbdb430ba23ffd387c

Observation 0f3ad922-f04f-4d20-8792-028b5c857272 · inbound

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference cites this paper.

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:02.786626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:02.786626Z digest=sha256:4dd8c2b0783c5c7930c583a3cf30868b303c4d3ad831164e1e7207fc4d96d8ad

Observation d9d74b36-10be-4808-90a4-be9ddf1aa955 · inbound

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning cites this paper.

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T07:25:56.953876Z digest=sha256:def4c897908ab93b41d9007b485521e8a34b2a697a9d158bffa88099cab2ac2d

Observation 8a3df2a8-96b4-4863-b7cf-3b9461b02289 · inbound

OrchANN: Hierarchical Orchestration for Skewed Out-of-Core Vector Search cites this paper.

OrchANN: Hierarchical Orchestration for Skewed Out-of-Core Vector Search RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T13:51:32.105366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:51:32.105366Z digest=sha256:c95fed82a65a8070963e57f1ccb4a67ea2385d9f9ebbcbc0f0a849bb5faf947f

Observation fb654b63-897c-4477-b8b9-d47133cae295 · inbound

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs cites this paper.

ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.127301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.127301Z digest=sha256:faf3fb08927a451aecffa4d1c322fea7175029ad11d2c51e9145b438084c8150

Observation ce26225d-95d8-40f2-983f-db61e57bddef · inbound

Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction cites this paper.

Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 7

Resolution
malformed identifier
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:11:51.802541Z digest=sha256:6a585b0e47f05d3247bf776fb0906344ec3918a5965cb8626b19a772705424aa

Observation 615c7603-3175-419f-b001-790380685650 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:49:13.353700Z digest=sha256:931a11e94f27e8ce321bb557fe1cc160a7498ec1978095707b6cc72e7d8c43bd

Observation 04a85071-d8c1-455e-bb55-4a0c42a01ec2 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:48:21.341462Z digest=sha256:5458b1ccf795f7dfbf287a9ced9d3e634f0ced62146e8b8d8c7ba24ef44ed608

Observation 88ec4cfa-dc85-417e-8ec1-ab606d2aa155 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T07:35:30.155592Z digest=sha256:35fd0b8234d2ab203cb5b9bbdb630c980ff01854d5a55c3ddae2791cbaaff871

Observation db1b1c4b-5f40-4c74-adc8-2abde36d6272 · inbound

KV Cache Offloading for Context-Intensive Tasks cites this paper.

KV Cache Offloading for Context-Intensive Tasks RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:42:39.684879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T16:41:30.021302Z digest=sha256:4ad52a752aeabfc336d7fa1009b195363b3e799a653b16e7d649b19627ca2b0c

Observation 3376f755-de3f-43f0-968d-d4a73ec6ca4e · inbound

CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference cites this paper.

CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:52:25.148104Z digest=sha256:b3a301abed2416d4f317dbc9a6931089c63ef927e82f06edc90f63e119a6ec17

Observation 9730729a-f235-4635-b290-f8efbcd46da0 · inbound

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing cites this paper.

SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:32:02.222528Z digest=sha256:40a5da342d10041a9ecec1e461d5d51854a3be1ddcae2fdc1bcee406530d6b22

Observation 22d5ceab-e1c3-4032-a1f2-3eea0fa57b3d · inbound

Graph-Guided Adaptive Channel Elimination for KV Cache Compression cites this paper.

Graph-Guided Adaptive Channel Elimination for KV Cache Compression RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:06:08.786535Z digest=sha256:d753246c95799434fde0b56e7514ccacd163781f22908046f68ad818cd76cc49

Observation d641adb8-941b-42eb-b79b-f44fa5960f3d · inbound

AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization cites this paper.

AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:09:11.684432Z digest=sha256:6054ae5bce88c85e27dec7f9961144b1f675759da3c1f979ba2256a74256bcb6

Observation 22025de9-d6e1-4a7e-bbc1-35ea733ab356 · inbound

DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference cites this paper.

DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:49:01.267556Z digest=sha256:123fadefa63f680a792b1437a016385da0d1c12f0f819d7f2217e2810ebf4191

Observation 47b902b8-c3e2-475c-85fd-c076c35d01a7 · inbound

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving cites this paper.

Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T13:15:21.201950Z digest=sha256:da458e3ea24225576e542e6756ae8645e0a0e3e30aa84b7ee9df1cb3920ede15

Observation fb27999c-de98-4a89-9908-8c25500d99a3 · inbound

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache cites this paper.

Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:15:35.871863Z digest=sha256:cc08f70ba7a4b5b775ba0053054b89e4f983cfdcb558702513ade2556049d71a

Observation 85741ffd-ef11-49db-9dd8-571b56e02b91 · inbound

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction cites this paper.

Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T05:02:25.513351Z digest=sha256:83d6950970e6526c71c4865578f01de27cf3e865075f6f64fbaaeed1b7e125a8

Observation 838dcba2-0b94-4b78-930c-748859302263 · inbound

ScaleGANN: Accelerate Large-Scale ANN Indexing by Cost-effective Cloud GPUs cites this paper.

ScaleGANN: Accelerate Large-Scale ANN Indexing by Cost-effective Cloud GPUs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:53:42.872514Z digest=sha256:bcf742453d88cc566292c03782d23d75efd08b9597a7e8d91670add7b6421606

Observation 80d8bbe1-747c-453d-8ee9-b1940ddeceba · inbound

AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference cites this paper.

AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.395521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T04:42:26.433689Z digest=sha256:710b114c1fc3cd2b9a63f379150833fe44a29586783da31220ac17188fb735ab

Observation 085003e4-696a-4ef7-87b3-c39a1326c600 · inbound

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference cites this paper.

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:18:13.967617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:13:46.098095Z digest=sha256:3d2a377a4159d38ca92f4c9f00acc081032e2b9d17b3eba54313ab3a727e5f3a

Observation 73ce100b-1fae-45ab-ab40-02bf7e9a5f2a · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 150

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:57:41.184943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:976771cb2f1f619747e3e30d449858a5faa98410fbb351e2c240419e3d26cc16

Observation f3400f28-bfc0-4fdb-a459-eddd582044c8 · inbound

RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference cites this paper.

RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T06:35:29.439088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T06:34:15.790154Z digest=sha256:7bbb5bc1dc4f4326f7297b846d420f11fcc6fbb4134e93057cadb80607074c21

Observation 792b288a-5629-4875-a7a3-0eadd081c27c · inbound

DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch cites this paper.

DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T15:06:54.254928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:06:54.254928Z digest=sha256:9f7dda542fdfc203a6050d0861d4d3713bac34e29e1823a762de4efd7f7b8fea

Observation 5b8ca242-c574-4053-a4fd-bab8a61c873c · inbound

ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression cites this paper.

ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:01:06.893261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:01:06.893261Z digest=sha256:70bdf868c5125ecf18565bff280a4b96f60ebdbd7167af8ce849bd80f1162591

Observation ace33b7f-be03-468b-b58c-17cc2a456267 · inbound

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs cites this paper.

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:55.807123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:43:55.807123Z digest=sha256:6538fec07220a61bd58893e0e3720906940fefb30b9bcdaa2dd29d06a2350b3f