Pith. sign in

Paper Citation Record · LEDGER

Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2407.00079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.00079 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:11:15.969489Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 352b9913-3e7c-4066-8ad5-660f610f4814 · inbound

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval cites this paper.

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:02.005662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T08:12:01.798459Z digest=sha256:97560583725e0c04ad0fa8e0c391e1336357351fca79a3cffec5f220d05e8a76

Observation 7528d19c-e4e7-4b6a-b19b-c5c1cff684c4 · inbound

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching cites this paper.

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:58:11.891568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T16:57:46.645061Z digest=sha256:0bcfaa61313a259b4222838c9b50ebde8b675ddb6dc849e4651d3c1f5df37300

Observation 1597fe2b-4566-4862-95e8-7fe26d69ce6c · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.720450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:6e61cb82bb2dce1de291edce831ad791b856ea6fb2e8f91e50525dd347a8fe1b

Observation b0771189-3fea-4e80-91d1-60c3bcb1ea70 · inbound

From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs cites this paper.

From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:05:09.843818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T11:05:09.588491Z digest=sha256:947deb17a5bdc34ca54f9893aedd20f73579751ac9d041002668558d3ddbcba6

Observation 3a129b94-c23c-47d6-b7e9-d5193030acaa · inbound

Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving cites this paper.

Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.476761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T13:14:26.628447Z digest=sha256:817ba478a69f41dce3772db05315b9e9e535dee862515549c0fa89b0bd2d13c8

Observation 9ebb2090-66ea-44ff-9cb0-72c5a1ba1f84 · inbound

On Evaluating Performance of LLM Inference Serving Systems cites this paper.

On Evaluating Performance of LLM Inference Serving Systems Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:15.969489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:11:15.969489Z digest=sha256:fb186888ba58ca8f27b825dc0d0e730b2e0810e5dc02ea73b04f037cf7c942bb

Observation 2fe5dd24-aa07-44fe-8361-135161d09c58 · inbound

Past-Future Scheduler for LLM Serving under SLA Guarantees cites this paper.

Past-Future Scheduler for LLM Serving under SLA Guarantees Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:20.294435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:20.294435Z digest=sha256:bcb38ca8b5cbb599becf40eeca54b21e4af2ef2fb20255eaa77202cca11221d1

Observation 42bc596e-6c1b-4d5c-90ba-f02346f3abcd · inbound

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs cites this paper.

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:33.971861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:06:33.971861Z digest=sha256:ba36b6caac85c920251a1e13692be6b17a21ec4aec77827736c0d190ca53f79e

Observation 7e423ebc-4e6b-4f56-876d-b671ecf41ed6 · inbound

Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving cites this paper.

Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:11:43.902548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T15:08:15.328986Z digest=sha256:970c86f6c65c58353d43c5adcfad642b527fab06e7d6b5297c87ff3feab7c803

Observation 0e7aad83-de96-4a2d-a74d-e8aea9d126dd · inbound

HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling cites this paper.

HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:41:51.752300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T21:39:02.560962Z digest=sha256:63e688733e94fcf603da5c5f4ed375ea223e5d6c285229a40dcf35daebf77c91

Observation 92be0001-d2d4-4b27-b8bb-b84232124710 · inbound

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference cites this paper.

Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:38.950601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:38.950601Z digest=sha256:126fa0b69550b4b0879aa0f9d955e976b204895ae335ad26dba7d8d830623d44

Observation 259df76e-5863-4d72-8fd7-f6bd7ca33955 · inbound

TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications cites this paper.

TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:20:38.741005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T21:17:58.275784Z digest=sha256:4d25b734e4af59034a1477e12a86ce1b8e0ab12f64c71fca3a8b8ba4ad328d71

Observation 358e0bb5-b56f-4ab7-91df-b8ce2596b5a0 · inbound

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse cites this paper.

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:05:39.054797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T02:04:46.559881Z digest=sha256:5684ea67b2f1563e0b98fe8a1ef58f92849e791606e0538e4208e876b99729c8

Observation a9981e83-bb93-4c9b-b8c6-f65b2d426fc6 · inbound

The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project cites this paper.

The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:45:12.282525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:40:27.945478Z digest=sha256:77d90ba0e5702db5ff778905215d9ec9b8f6d87148745a8bfaa134a036253786

Observation 2c786606-936f-41f7-9486-acba5244a115 · inbound

JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency cites this paper.

JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:28:09.832457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T19:26:38.134505Z digest=sha256:ea678fb53a2cf9160b6cfb8aed180a5810642f4ee9be5943ca77569d568da43b

Observation be385161-c75d-4823-ac03-f1f4d621b35d · inbound

MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs cites this paper.

MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:47:12.427852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T07:45:17.043107Z digest=sha256:6ab5f0167c765118b7ca922c771178986a70ea680db09f98cd2a45b00dcaa6b6

Observation 1324798b-10f5-435d-887e-e830606d2d60 · inbound

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding cites this paper.

Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:18.872688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:56:39.124969Z digest=sha256:39a32cfab87ca7d3dce8cd114646069dc27fddd5a0df3864110db781e8774731

Observation ce4ab281-2187-4221-b962-1aa18009bdbf · inbound

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference cites this paper.

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:14:10.662004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:01:02.726796Z digest=sha256:82feda7e12fb0a6bc8d9f43b871d73631081e1b65aa36e7fe9e110fee3f05132

Observation 72ce4e52-8721-4738-a326-af124bf4f46b · inbound

Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving cites this paper.

Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:10.413581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T16:19:33.613685Z digest=sha256:8efeec6fbdde3103368563ac6c5ed5159037087492bdc9bfe1b73cdc180a3954

Observation 0c86b6ee-87f1-4c9d-9069-54b6c96603dc · inbound

PreFT: Prefill-only finetuning for efficient inference cites this paper.

PreFT: Prefill-only finetuning for efficient inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.631370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T02:10:14.721584Z digest=sha256:5a62b2bc6b12d3e8abcef003aa13c5cf3da5eed73b79a902fcdc9e83f63b2d20

Observation f6e62323-3a28-47cb-85f4-a532990b95b5 · inbound

Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture cites this paper.

Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:09.325179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T22:12:13.114405Z digest=sha256:548e6398724adcdd20a7340f242fafbce490767933ae0084df93bcfe1ef68fe6

Observation 87eee698-7063-4fc6-a4ec-335ac8a4eaf7 · inbound

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI cites this paper.

Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:13.602230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:30:56.324289Z digest=sha256:347a1ec6b839ccf6e4bda1a33be77299b13d6bf76a308c34b4cd4aa77cb4f13b

Observation da99d0d7-d3d8-4dc5-96d0-80e2439e0af0 · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.446547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:15a1c8940f03081a94439bc374e975b8086cc95467a35e47cff762476a155893

Observation b99726db-bcc4-451a-a747-db9dd14c3604 · inbound

Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching cites this paper.

Semantic Cache Distillation: Efficient State Transfer via Reuse and Selective Patching Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.081015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T22:43:25.637631Z digest=sha256:e03673619a491fd08efce5d38052b684ea80a638daa48f7662152f4b5e3d5b31

Observation 6a7026a3-6fad-4535-9b78-648b29be30b2 · inbound

SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving cites this paper.

SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:27:26.298208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:50:50.066713Z digest=sha256:7b2e5bbf3b50ad904131acbf7eeb02d715f9f675195a8755b4ff5474e329ee91

Observation e304f344-56ad-4675-bbe0-9395700334cb · inbound

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation cites this paper.

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:11.785527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T08:45:42.781160Z digest=sha256:2108194981d61f7141547db9e78653e5d301f3599777a9284feda16c977edd9c

Observation ef0ee4aa-1e77-4571-83ad-aee61d130932 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.750859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:7dcb2ce3a5c7d242a7d9214ca4451d88fd00e45c5fca372089ab4ad3a40f382e

Observation a1bed497-9edd-4dc1-a395-b959a7092a3d · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:19.644055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:19.644055Z digest=sha256:f2815b2db851aa7c2b1381020e03055c2940f6d29b628bb93dd457f3fa8d75b0

Observation 96d4f8d0-d408-420f-aa10-6104605c5c63 · inbound

Human-Less LLM Serving: Quantifying the Human Tax on Throughput cites this paper.

Human-Less LLM Serving: Quantifying the Human Tax on Throughput Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:11.798434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T00:44:55.517266Z digest=sha256:1b401987d37c88b6cfe70199f1a9c8f0270fd61f73f29e545b24d3c741050d0e

Observation b1e392b9-f9d0-46e6-ac5c-b3c2d2f25b59 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:45.338293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:585773a2c4af894e6feca916ee50c741e201dc08f06a5f8fd14a6203543518f5

Observation 1d063ebc-92c8-4e23-a621-98eefecb3671 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.140206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:b8970ac5e809166fa385dda38169ac01518ea92195af539f5e1802741ccca7b6

Observation 50f5c434-9115-4a2f-b9d4-92fe427608c8 · inbound

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference cites this paper.

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:35:43.612928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T04:13:53.722933Z digest=sha256:62d4df40302194e9cc5340e2364624db80958ecc2b62aaf69870731dbc948224

Observation f2898e7b-bcdb-45b4-8c5e-0560faa4ad5a · inbound

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving cites this paper.

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:56:43.946753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T06:40:53.505889Z digest=sha256:7fb7d22fc71e6332d58f92b2ec60e32da1448cea44afa0f1ce7f6f6862834fcd

Observation 69825baa-86d1-4130-823e-eadfc4c44395 · inbound

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving cites this paper.

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:58:50.500809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T18:58:44.830766Z digest=sha256:b8322c556611e561dea3cdff3532e1530e64f60d34e6b2ac50fc097a5483000e

Observation 36bbb44a-af48-4c3d-9700-7d0f15b55344 · inbound

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters cites this paper.

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:37:42.198474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T06:30:30.308713Z digest=sha256:e7a3cd33699e035f1d2952bef26ce49b472a7f1d9f39609f5b25f2441c0152ab

Observation d8b16788-b522-4b7e-ab51-38379c76484a · inbound

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving cites this paper.

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T21:08:24.159706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:08:24.159706Z digest=sha256:d6393a120415a591d81463e3f65e8467367bd356621d5a790891cef97469a85f

Observation b019a78b-e277-413a-8f33-392ccf9cb36d · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.129136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:5dcf05241e169669fe9d18b2a41e8debae178a14de9834724de3e962b3962c58

Observation bdbfc038-faf2-431a-afd8-67b23139ec73 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.505587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:138664bd844a44a0aec938753385181c8c8592b7ea46f513e0580ecde1cd0f86

Observation d36af72a-b138-47a3-8cc6-7be560f40b68 · inbound

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference cites this paper.

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 103

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:17:45.637015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-11T00:12:13.830917Z digest=sha256:e2a0258b448170e90b7e96fd86dc8e1196b518289ebfab292e85add606a47def

Observation 6a9d0331-1118-4278-9fc4-43b1397b7ba6 · inbound

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression cites this paper.

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 117

Resolution
verified exact
local_arxiv, observed 2026-07-08T03:14:31.493985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-08T03:07:23.648382Z digest=sha256:eb8cdf935b465d641c08d3169e7d1bbc995283ed1b8932eb6466895ee1b970d0

Observation b2f41ac6-904f-47b6-9e19-4a5677ee4d3c · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.325524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:b1135f371f7fc02fd0d93bf0e7bdc79f8dce19aa820fc53bd9807e0e7ffb179c

Observation 43c055a7-a8eb-4daf-8c8d-f31ee9ecee5d · inbound

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization cites this paper.

Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:56:40.960364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T00:55:52.215700Z digest=sha256:1f7427c4b4ee64f085b1d6e124b647c6d450a919fcbfdb4815218229c595d1b8

Observation 2f6ad3a4-c991-4d54-9e5d-9ccff413a265 · inbound

Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices cites this paper.

Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T07:32:31.866525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:32:31.866525Z digest=sha256:214a19cc8cbb77a8550f891ab0ed13924b8e0f087968b57a4800965b172420eb

Observation 1e5e8596-d161-4b47-ab91-067c04b95cb9 · inbound

Topology-Aware Data Movement for Disaggregated GPU Inference cites this paper.

Topology-Aware Data Movement for Disaggregated GPU Inference Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:34.428229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:34.428229Z digest=sha256:b3d3b1e7d9e9cd5c986e7872fec0c1cf72bd8792384485be022d2276c4a70552

Observation 8ed69102-1abf-4a82-988c-ccbeda51369b · inbound

Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework cites this paper.

Spatial Prefix Caching for Wireless Edge LLM Inference: A Stochastic-Geometry and Queueing Framework Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:33:26.100297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:33:26.100297Z digest=sha256:d45d8a0b4f6cd2cc4a8bafabec44a483973cf251d5ee2a660c4c187e3db60d70