Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T08:12:01.798459Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 30 inbound Pith citation observations for arXiv:2409.10516.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T08:12:01.798459Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:02.786626Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 118 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 73377b4f-3a13-438c-a642-8cbd085b289c · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Scaling Learning Algorithms Towards
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a089186-eab7-4b7b-92f1-5729d0385891 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval and Osindero, Simon and Teh, Yee Whye , journal =
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5a5c367-551b-4164-bb4a-d682582842c0 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval 2016 , publisher=
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4770016a-b290-44c3-ac4f-3bf5ed777ecc · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval IEEE Transactions on Computers , volume=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1eb601b0-52b1-4549-be1d-bd86d0ea179c · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval IEEE Transactions on Big Data , volume=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb833c48-76a8-4886-b567-fd6b592e4807 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8b466e0c-38b5-4a88-ae52-218beb04d046 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval OOD-DiskANN: Efficient and Scalable Graph
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c4bf14a-0a2f-4ae7-81df-42c4df7d4c1c · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval CoRR , volume =
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation feae6a00-4d95-46ab-b168-e16390f6b7c5 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gomez and Lukasz Kaiser and Illia Polosukhin , bibsource =
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a3d8e2c6-0682-49a6-ad39-3f9f9f04647a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval A weighted nearest neighbor algorithm for learning with symbolic features , volume =
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 68bef60b-9c91-4749-b434-cda636492075 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval seed , volume=
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eaabb11d-b223-4045-a11b-876d5acb14a1 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Query by image and video content: The QBIC system , volume =
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 52c5614c-469d-4713-9da8-b377c1b7e5c2 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval PQCache: Product Quantization-based KVCache for Long Context LLM Inference , url =
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 681e026c-d1d3-4d37-8881-d80bf084cafb · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98d952c5-9c26-47bb-b2e8-1ae74c794f5f · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 480d7388-ff99-4d8f-b3f9-5249d13b4eb5 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc65432e-90ea-4a2b-a055-9665cb4d7c1b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval The Twelfth International Conference on Learning Representations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc2b2041-fb93-4cef-bb14-1326f0f33abd · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c948570-1561-4632-8e8f-813da7fdd902 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Sean , doi =
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8136a523-356d-4264-8f3b-2d40edc397da · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Reformer: The Efficient Transformer , url =
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 532345c6-e323-439d-a52c-6b6404c76e8d · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 05401a3e-f480-4eeb-8c23-f3bc50dc22a1 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f4d7360-0853-4764-bd73-9e654bfeb271 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval FlexGen: high-throughput generative inference of large language models with a single GPU , year =
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f5e2b5c-0adf-4a7a-bbc6-8c324166bae3 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Model Tells You What to Discard: Adaptive
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5cc8c76-d3a5-444a-a66a-316f6e25d7cf · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval RingAttention with Blockwise Transformers for Near-Infinite Context , url =
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10213dac-214f-4e5d-a057-f7a61a86d039 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval and Ermon, Stefano and Rudra, Atri and R
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b945c80e-9091-46e9-abf7-51f6f0641b8d · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eff34419-a984-4160-ba45-8e5537f4592b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Splitwise: Efficient generative llm inference using phase splitting , year =
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46767789-bc80-4e7b-a090-066f3c90e449 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient Streaming Language Models with Attention Sinks , year =
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3d72fdbf-9b5a-45ee-9715-17f6b33d85c2 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0f3b649-679c-47af-89a5-3bc5c3d58726 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Snapkv: Llm knows what you are looking for before generation , url =
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4e29f8d-0352-48bb-aae7-583e9f7f56d2 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MagicPiG: sParse Inference enGine for LLM , year =
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2541b9f8-7978-4b7d-88f8-70bd248259d1 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Longformer: The long-document transformer , url =
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3d4ef521-4f59-4f76-8295-7fede33ec93c · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval InfLLM: Unveiling the Intrinsic Capacity of LLMs for Understanding Extremely Long Sequences with Training-Free Memory , url =
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 34979964-5b4d-44eb-91cc-e8261a15978a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval The Faiss library
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af1fe93a-08a6-4f1d-af80-459512bb79db · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval RULER: What's the Real Context Size of Your Long-Context Language Models? , url =
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6c9a0c0-0791-4577-801e-6aa7211577db · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval On the generalized distance in statistics , volume =
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8b22e7f-c204-4990-966b-43cc39219d44 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs , volume =
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e442ee8b-7adc-43dd-b6d9-5e75a661b96a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Lempitsky , bibsource =
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4542df4b-a164-4c93-bb25-1d81b096be6b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8f98e092-96f5-438c-acad-8377aa688dba · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Video Google: A text retrieval approach to object matching in videos , year =
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2509cfc0-0aeb-419b-aeef-a78995630b96 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Big Bird: Transformers for Longer Sequences , url =
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 745738cb-b938-4ec7-af03-a2080abd3512 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval IceFormer: Accelerated Inference with Long-Sequence Transformers on
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a98d3c41-c6fc-4b55-b114-5464cee32fc1 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Generating long sequences with sparse transformers , url =
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ba553ff-7eda-4e11-88f4-d91db2f74560 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Keyformer: Kv cache reduction through key tokens selection for efficient generative inference , volume =
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa01a263-ffba-4f9f-8243-f901132f938a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unlimiformer: Long-range transformers with unlimited length input , volume =
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cb10304-9f9b-475e-90b5-8ff054b7368a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53eeda90-c923-431e-855b-5c321770818f · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval SparQ Attention: Bandwidth-Efficient LLM Inference , url =
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f9c2d87-0097-4d12-8e4b-5a18769ed53b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient and Economic Large Language Model Inference with Attention Offloading , url =
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8139999-860e-43fa-a8b6-40f31b3d778b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time , volume =
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65b1e0f9-32a7-4501-9ed6-ada258b8ce49 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96a4dd63-b7d4-401b-9466-4e707bd9b154 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cef76df7-22ed-4f78-8e8f-0895d280d610 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Loki: Low-Rank Keys for Efficient Sparse Attention , url =
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff46e5bd-3022-455a-abb6-3b5e9875e97d · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention , url =
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a2eed4ae-aa10-49a8-8db0-b7c86e4baa44 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Mooncake: Kimi's KVCache-centric Architecture for LLM Serving , url =
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 06e5669a-1cc8-4592-bce8-a156da499688 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3061713c-b5fe-4490-842c-3b0d7b0bb4ba · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models , url =
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74eb02e0-e6c8-4780-be1b-fa0dccefba6c · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints , year =
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3fd397b5-19b7-4cc1-944b-d153fd9e773b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gonzalez and Hao Zhang and Ion Stoica , booktitle =
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6b5812cf-0fcc-4015-8a02-991f2e04b931 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8722174-dc9e-49be-8b6f-32734276e539 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval , url =
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation faafad50-c5be-410f-bc14-5a3bd69c8c8e · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval PinnerSage: Multi-Modal User Embedding Framework for Recommendations at Pinterest , url =
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21bb8f70-62f2-4ce7-a1d2-0004b35dd5e3 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Non-metric Similarity Graphs for Maximum Inner Product Search , url =
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df2a457f-61c3-4c4e-b1da-17a74a43d217 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42e456ea-3863-4379-8c0d-3eb35e249275 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Yi-6b-200k
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cace07c2-64ad-4600-b846-d56752dbe1ec · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Yi-9b-200k
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fce6e1f4-0909-44ef-b5de-d21713e24e2a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval ETC : Encoding long and structured inputs in transformers
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c001afe8-7762-4f13-88e9-592d3789f9d8 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4fb1939a-6070-4b4f-b8ac-a5d3627aacba · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Longformer: The Long-Document Transformer
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e70b4920-e52a-4c13-b629-1af851b36e94 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Unlimiformer: Long-range transformers with unlimited length input
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 666adf94-9fe6-462a-817f-4cb202cab9d2 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e74fcdfc-ccbf-486d-a42a-fc3c4817577a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Sean Wang
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d42c4ab1-d012-438b-b604-2a493f39504a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53a10c02-1a5f-4b5a-bc19-15d08887af1b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Magicpig: sparse inference engine for llm
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 64ba7080-2156-4d36-ab98-a3ae8415651a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Generating Long Sequences with Sparse Transformers
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 737ecc75-17e6-40af-8023-27c62719721f · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval A weighted nearest neighbor algorithm for learning with symbolic features
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 419ac5d1-8298-41ff-85a9-a111fcc09f28 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval URL https://doi.org/10.1145/ 2959100.2959190
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a9b44847-da35-4ee5-a6bb-3af273dba721 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a160637-9dac-40f9-b228-861aed039381 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Attention is naturally sparse with gaussian distributed input
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48ff9088-b825-41c4-9571-c07356e69a1b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval The faiss library
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 296f2418-3c36-41d6-93fc-44a00447feab · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Model tells you what to discard: Adaptive KV cache compression for LLM s
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 17bf77dc-5406-4196-afcc-c97f412ec7a8 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Context caching overview
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd52da36-d1e3-439e-b933-908c14493742 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Llama-3-8b-instruct-262k
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 02f35180-7252-4d75-9912-41a821d4638b · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Needle in a haystack - pressure testing llms
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b97d136b-87e2-4af4-b22b-08576488198a · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval LM -infinite: Zero-shot extreme length generalization for large language models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 668d82fd-a21c-4c5d-a307-c768301643ee · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ded72ace-d1a3-4e3e-bf3b-f6b0156cd3d8 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c074ab33-5a65-427e-b9fc-8b5844b8a4da · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval OOD-DiskANN: Efficient and Scalable Graph ANNS for Out-of-Distribution Queries
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e637f98-6e48-42a6-8a2c-3a55a8834366 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1a24bd02-ad16-4266-b011-9c8c2f9f01c6 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Reformer: The efficient transformer
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7b9b8996-e134-4302-be8b-32698285da64 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Gonzalez, Hao Zhang, and Ion Stoica
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a974e6a3-8eab-4934-a3af-bf4b671125a5 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval InfiniGen : Efficient generative inference of large language models with dynamic KV cache management
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb2b0e7e-821b-4fc9-8a70-7f0911754172 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval SnapKV: LLM Knows What You are Looking for Before Generation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c38cf58e-3af7-4ed0-9327-6a5c5f0463d4 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Ringattention with blockwise transformers for near-infinite context
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7e558ee-4788-417e-81fc-d064ec1f40cd · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e8f0203-f860-4066-bf27-fec2c4dc8e00 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval On the generalized distance in statistics
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23807758-2930-4f62-a3a4-1f5d639109af · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96cc7b0e-03d9-4440-b84d-56afb6ce0902 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Iceformer: Accelerated inference with long-sequence transformers on CPU s
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6294af34-9e02-4a4e-bd99-1c238b2cab26 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Non-metric similarity graphs for maximum inner product search
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8064e1cc-636f-403f-a912-a0e42e407449 · outbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Ghosh, N
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e5c3793-50d4-462e-903e-5c8340166dfe · inbound
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e10dd34-8e53-4d49-97a5-93813fe8839a · inbound
MoBA: Mixture of Block Attention for Long-Context LLMs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 37d8729d-fb77-4706-967c-f7347e2aa8e0 · inbound
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93e26f7a-290d-4965-92bd-5f69f6d8c844 · inbound
Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 577dc9ca-44fd-4f3e-ba4d-8fc50e49423a · inbound
MemOS: A Memory OS for AI System RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f3ad922-f04f-4d20-8792-028b5c857272 · inbound
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d74b36-10be-4808-90a4-be9ddf1aa955 · inbound
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a3df2a8-96b4-4863-b7cf-3b9461b02289 · inbound
OrchANN: Hierarchical Orchestration for Skewed Out-of-Core Vector Search RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb654b63-897c-4477-b8b9-d47133cae295 · inbound
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce26225d-95d8-40f2-983f-db61e57bddef · inbound
Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 615c7603-3175-419f-b001-790380685650 · inbound
KV Cache Offloading for Context-Intensive Tasks RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04a85071-d8c1-455e-bb55-4a0c42a01ec2 · inbound
KV Cache Offloading for Context-Intensive Tasks RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 88ec4cfa-dc85-417e-8ec1-ab606d2aa155 · inbound
KV Cache Offloading for Context-Intensive Tasks RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db1b1c4b-5f40-4c74-adc8-2abde36d6272 · inbound
KV Cache Offloading for Context-Intensive Tasks RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3376f755-de3f-43f0-968d-d4a73ec6ca4e · inbound
CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9730729a-f235-4635-b290-f8efbcd46da0 · inbound
SAGE: Selective Attention-Guided Extraction for Token-Efficient Document Indexing RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22d5ceab-e1c3-4032-a1f2-3eea0fa57b3d · inbound
Graph-Guided Adaptive Channel Elimination for KV Cache Compression RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d641adb8-941b-42eb-b79b-f44fa5960f3d · inbound
AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 22025de9-d6e1-4a7e-bbc1-35ea733ab356 · inbound
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 47b902b8-c3e2-475c-85fd-c076c35d01a7 · inbound
Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb27999c-de98-4a89-9908-8c25500d99a3 · inbound
Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85741ffd-ef11-49db-9dd8-571b56e02b91 · inbound
Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 838dcba2-0b94-4b78-930c-748859302263 · inbound
ScaleGANN: Accelerate Large-Scale ANN Indexing by Cost-effective Cloud GPUs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 80d8bbe1-747c-453d-8ee9-b1940ddeceba · inbound
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 085003e4-696a-4ef7-87b3-c39a1326c600 · inbound
KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73ce100b-1fae-45ab-ab40-02bf7e9a5f2a · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3400f28-bfc0-4fdb-a459-eddd582044c8 · inbound
RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 792b288a-5629-4875-a7a3-0eadd081c27c · inbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8ca242-c574-4053-a4fd-bab8a61c873c · inbound
ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ace33b7f-be03-468b-b58c-17cc2a456267 · inbound
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.