Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T15:06:54.344008Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2607.26475.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T15:06:54.344008Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4c452bd7-4d8b-4f78-9cd6-342c182fa6bf · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Gulavani, Alexey Tumanov, and Ramachandran Ramjee
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eefe7e0-0764-45e4-9a18-fab534c0bd48 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2564b0b3-500b-440b-889f-443761bd7ee2 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8efbff-a2fb-4b46-9da0-d096620cbe75 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09fd9b61-80b0-4d90-87a5-7115e6cec634 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Aly, Beidi Chen, and Carole-Jean Wu
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79958e84-a838-4fb0-9351-44aeacdc9073 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf9e74ee-e9f3-48cb-9228-bfd48a7098dc · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1183dc5a-ecb9-466f-a6ce-67d9c5e53850 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63dae347-e822-4f8a-80ad-d723fa99cf1c · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98423e8e-b5ee-4fd2-b5e6-4ba7a0d74e38 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9748d17e-21b8-4f59-9696-5cf628b8f66f · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f1939f-4852-44c6-ad80-d2c249913367 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97335c9f-704d-411d-95ee-3218daf482ac · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a07d110-f620-40dc-8898-f37a195a2a7f · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Gonzalez, Hao Zhang, and Ion Stoica
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b7382a-6022-4121-be07-fe8c73787743 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d228155d-2d7a-4eef-938f-e03a7647ae9e · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23479e2f-7291-4257-8865-668f9463c5ce · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e252f059-912e-4cba-8b5d-bd908f8bb083 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 792b288a-5629-4875-a7a3-0eadd081c27c · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd00e49-b66c-4432-a6a7-56c0a2b9944f · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc9f0db-8cf0-45c3-a86e-272a458abfc8 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09843e44-e451-4ba3-a1d9-21da068b0745 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch The Llama 3 Herd of Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f66a4f-aa3b-4e19-a567-13069e5ccb7d · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Qwen2.5 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcad6095-135b-4efc-87cd-ba66181bb4aa · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f20d5d-1fcf-4a96-b5e5-37bfa9edeb1a · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Accelerating Large Language Model Decoding with Speculative Sampling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d44c84-d350-481f-a96f-059836c33170 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc52335d-66ae-48ec-9852-793794367816 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Fast Transformer Decoding: One Write-Head is All You Need
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 269daa92-ecca-401d-9b8c-7f3ab60f3e19 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90decea-93a0-450f-92de-c7c09d3ec2f9 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6edad205-356a-4f9b-823e-2bc6a5ea71cc · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4181d1-05e3-476a-9459-c8495ce10012 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31308b7-64c8-4cf0-813a-f2be4ad5523c · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c852c2a-63c3-4e5f-a077-7ffccd88ea75 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17f8006-8d23-4ff8-abf5-8c441c6c2a92 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch 2026.SpeContext:Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c202f01d-ffa8-44ac-be14-31cc044adcde · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e32c5e24-579e-4c09-af90-242c64cbed12 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b37f6c-fe01-46b1-8791-52e31648eb5f · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d8e877-0643-41ce-b2dd-c6820612372d · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47f875e5-b963-4b55-b0fe-01d743783f1e · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25dcbd49-b8cc-43c3-b0d0-4ae08f66900d · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Barrett, Zhangyang Wang, and Beidi Chen
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 856fe18b-3dc1-44a3-bc39-d3e7fc27f804 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee8b9ccc-c701-4315-8646-a4c81e53dcc3 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02836f98-6a69-4f83-ab3f-17d8e6771212 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch InProceedings of the 29th ACM Symposium on Operating Systems Principles (SOSP)
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b5298d-e673-4f36-a5f4-058153f598b5 · outbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.