Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2407.00079.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:11:15.969489Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
13
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 352b9913-3e7c-4066-8ad5-660f610f4814 · inbound
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7528d19c-e4e7-4b6a-b19b-c5c1cff684c4 · inbound
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1597fe2b-4566-4862-95e8-7fe26d69ce6c · inbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b0771189-3fea-4e80-91d1-60c3bcb1ea70 · inbound
From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3a129b94-c23c-47d6-b7e9-d5193030acaa · inbound
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ebb2090-66ea-44ff-9cb0-72c5a1ba1f84 · inbound
On Evaluating Performance of LLM Inference Serving Systems Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe5dd24-aa07-44fe-8361-135161d09c58 · inbound
Past-Future Scheduler for LLM Serving under SLA Guarantees Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42bc596e-6c1b-4d5c-90ba-f02346f3abcd · inbound
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e423ebc-4e6b-4f56-876d-b671ecf41ed6 · inbound
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e7aad83-de96-4a2d-a74d-e8aea9d126dd · inbound
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92be0001-d2d4-4b27-b8bb-b84232124710 · inbound
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259df76e-5863-4d72-8fd7-f6bd7ca33955 · inbound
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 358e0bb5-b56f-4ab7-91df-b8ce2596b5a0 · inbound
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a9981e83-bb93-4c9b-b8c6-f65b2d426fc6 · inbound
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2c786606-936f-41f7-9486-acba5244a115 · inbound
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation be385161-c75d-4823-ac03-f1f4d621b35d · inbound
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1324798b-10f5-435d-887e-e830606d2d60 · inbound
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ce4ab281-2187-4221-b962-1aa18009bdbf · inbound
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 72ce4e52-8721-4738-a326-af124bf4f46b · inbound
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0c86b6ee-87f1-4c9d-9069-54b6c96603dc · inbound
PreFT: Prefill-only finetuning for efficient inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6e62323-3a28-47cb-85f4-a532990b95b5 · inbound
Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 87eee698-7063-4fc6-a4ec-335ac8a4eaf7 · inbound
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da99d0d7-d3d8-4dc5-96d0-80e2439e0af0 · inbound
Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b99726db-bcc4-451a-a747-db9dd14c3604 · inbound
Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6a7026a3-6fad-4535-9b78-648b29be30b2 · inbound
SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e304f344-56ad-4675-bbe0-9395700334cb · inbound
Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ef0ee4aa-1e77-4571-83ad-aee61d130932 · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 176
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a1bed497-9edd-4dc1-a395-b959a7092a3d · inbound
Token-Operations-Oriented Inference Optimization Techniques for Large Models Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 176
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d4f8d0-d408-420f-aa10-6104605c5c63 · inbound
Human-Less LLM Serving: Quantifying the Human Tax on Throughput Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b1e392b9-f9d0-46e6-ac5c-b3c2d2f25b59 · inbound
Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d063ebc-92c8-4e23-a621-98eefecb3671 · inbound
Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50f5c434-9115-4a2f-b9d4-92fe427608c8 · inbound
Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2898e7b-bcdb-45b4-8c5e-0560faa4ad5a · inbound
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 69825baa-86d1-4130-823e-eadfc4c44395 · inbound
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36bbb44a-af48-4c3d-9700-7d0f15b55344 · inbound
OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d8b16788-b522-4b7e-ab51-38379c76484a · inbound
CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b019a78b-e277-413a-8f33-392ccf9cb36d · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bdbfc038-faf2-431a-afd8-67b23139ec73 · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d36af72a-b138-47a3-8cc6-7be560f40b68 · inbound
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6a9d0331-1118-4278-9fc4-43b1397b7ba6 · inbound
DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2f41ac6-904f-47b6-9e19-4a5677ee4d3c · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43c055a7-a8eb-4daf-8c8d-f31ee9ecee5d · inbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f6ad3a4-c991-4d54-9e5d-9ccff413a265 · inbound
Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e5e8596-d161-4b47-ab91-067c04b95cb9 · inbound
Topology-Aware Data Movement for Disaggregated GPU Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed69102-1abf-4a82-988c-ccbeda51369b · inbound
Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.