Pith. sign in

Paper Citation Record · LEDGER

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

As of 17 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 16 inbound Pith citation observations for arXiv:2507.07400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07400 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:45:42.411640Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:58:39.204380Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ddc67721-7ff7-45d2-a005-486c7bc7f317 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows React: Synergizing reasoning and acting in language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.222956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.222956Z digest=sha256:3d7ba8a8c3acbd68f6964f5f21dfd81b0aec8b8f85b168f831e54af2b89eadb3

Observation a60f3c9d-5f10-40f5-95ac-be25f82fab4d · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.227937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.227937Z digest=sha256:5bb86f71396e99a87d210c1bd9c2e2eb2a8ad855b192dd0a5253920e5a62faca

Observation 55921259-8778-4990-a660-09036a7a2743 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.233314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.233314Z digest=sha256:60256dcf34f5aeab826cb05a2d3e103c3e95f310d851725fee76a07cbad61b27

Observation 7914f024-1158-4497-b3e8-1fb12c7a18a9 · outbound

This paper cites Camel: Communicative agents for" mind" exploration of large language model society.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Camel: Communicative agents for" mind" exploration of large language model society

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.238391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.238391Z digest=sha256:2e88820eb9c6d87f3c8fbb6819a58163bd74baa1155dc51443d94f4d27d38d4c

Observation d96f7d36-be33-47db-9b86-ca3418007a30 · outbound

This paper cites PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:45:42.772387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.243742Z digest=sha256:db3323d366608be1e77ed63c0c37580e9f6de344f06f73c91bf0c6801bea5a7e

Observation 90d6b685-6b40-4beb-9549-a1b1c9b8e75f · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.248910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.248910Z digest=sha256:ced93144c8e71d9faa2cb2cc1261fc5e87df538bed3369b6274380fc2a2efbbb

Observation f7eb983f-f5fe-46e0-b19b-ab1d3674e4a6 · outbound

This paper cites Gptswarm: Language agents as optimizable graphs.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Gptswarm: Language agents as optimizable graphs

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:43.022975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.255938Z digest=sha256:58eb4501e69e09719c5939680f07f0a98cfd8dbfa7ee994f7a02a5bcf0ac422d

Observation bc9def31-20cf-4e04-a077-65e5c6cd55e1 · outbound

This paper cites AFlow: Automating Agentic Workflow Generation.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows AFlow: Automating Agentic Workflow Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.261460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.261460Z digest=sha256:efb3ae94bc850db84c8a8f40a31921e7cce84980fb1071f2985dac14f081ab22

Observation 7c2ce70a-94c3-4c42-affe-fe97b883e641 · outbound

This paper cites Very Large-Scale Multi-Agent Simulation in AgentScope.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Very Large-Scale Multi-Agent Simulation in AgentScope

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.266656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.266656Z digest=sha256:adb9e5b6a022b42fadd9b2dc87e3cd86f7353604d231cb3455aadfbcdfff37d9

Observation ce19083b-5ba9-4f51-b18a-a88d0bb19648 · outbound

This paper cites Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.272015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.272015Z digest=sha256:8b58b34d19ec5412343294f09bee064643fb4b527c385689f26b39e3cf547297

Observation 96a5f89f-c5d8-4ad5-abee-3a93b3096963 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Efficient memory management for large language model serving with pagedattention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.276704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.276704Z digest=sha256:d94dbd7636c8edf725e45b97911415f5e83a1c14d9fabdd363387d4dbe195e9c

Observation bda1cd92-185a-472c-9c29-957d9edb573a · outbound

This paper cites Gonzalez, Clark W.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Gonzalez, Clark W

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.996931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.281347Z digest=sha256:c2e6af9713617b5f07a5b6562c117431422ded461e4b64806e0f9f7baf57c682

Observation c7c4aa97-e930-41cd-bf27-9bdc2e4bce7c · outbound

This paper cites TensorRT-LLM.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows TensorRT-LLM

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.982123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.286228Z digest=sha256:a7aa6ea69cd1f1baaf19beaf33e8a53f9bf9f1903df6ae49e37f1f0fd5f0929a

Observation 687a0462-43bb-40c4-8c89-d89cf734eba8 · outbound

This paper cites Automatic Prefix Caching.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Automatic Prefix Caching

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.966950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.290936Z digest=sha256:db581aded2ada67edbad9cb2c8b7a275dbbd329e604a0ba08a81cebcc0c263bd

Observation 74408f15-ecf0-4635-b034-f7e1da15f41e · outbound

This paper cites RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.296114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.296114Z digest=sha256:e3795af1e51df0665d45ea52e581d65503241d1be8da7e328ac13643e5b55b50

Observation fb353792-c84a-45fa-98b5-ff09cd5c5a15 · outbound

This paper cites {Cost-Efficient} large language model serving for multi- turn conversations with {CachedAttention}.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows {Cost-Efficient} large language model serving for multi- turn conversations with {CachedAttention}

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.951768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.302064Z digest=sha256:585747e3713ab060615387a73f034112d734e6928e92770ace65015ef32ca515

Observation 6426a51c-213d-4d06-b5d7-2c56a4c8a539 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Generative agents: Interactive simulacra of human behavior

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.306995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.306995Z digest=sha256:3e03a0ef0a93ace29259763fabad7e49261f5622b9a9f7f751c9cd0280fd3adf

Observation 96815315-538e-4beb-b77a-fa09f0deda9a · outbound

This paper cites A survey on large language model based autonomous agents.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows A survey on large language model based autonomous agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.311794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.311794Z digest=sha256:b7509c172fea87eda38bae40c2fa3c8a643e39e3919ce4f71261d3afd6aee914

Observation 8cf33364-7348-4f83-9dcf-2c133442380a · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.316209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.316209Z digest=sha256:fc4d145d88eed18837b6cf6055ddf714ea5cacb6e5d189c5ef95483219bc16b2

Observation 6ec7bd6f-7b16-49e6-be9b-08466e82c288 · outbound

This paper cites Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.320701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.320701Z digest=sha256:a385fa3591d811b5b1d55c9a8724ae95fb3f6572ff9a713318df0bb13fd3d1bf

Observation 007a7681-db51-4dd4-9127-da403a12d622 · outbound

This paper cites Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.325909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.325909Z digest=sha256:b80a22618915770a45e742541c1ac9d4396411efa9150b267b68d10d77f4120f

Observation efa75a61-7f40-4e16-8f3e-6bc7ec6816df · outbound

This paper cites MAGE: A Multi-Agent Engine for Automated RTL Code Generation.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows MAGE: A Multi-Agent Engine for Automated RTL Code Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.330477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.330477Z digest=sha256:0efbb5249a300ac2b851cf8801bd9643d2e5b84503b4e095f60ebc12fe1cbc36

Observation 0af19f0d-9703-4ccb-8a29-fe0a9137a7e4 · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows ChatDev: Communicative Agents for Software Development

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.334764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.334764Z digest=sha256:a48c7bc412f7d50699e649c9f097d47fe60152ff1fcfbcbed5c85220b00bc031

Observation 071ae83c-0559-41d4-9882-9ab8bc5e9250 · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Improv- ing factuality and reasoning in language models through multiagent debate

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.339249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.339249Z digest=sha256:251fcb09d078ffe538302f9cd1f63033581c5d826b05b243dc7675915795683f

Observation c39ff46b-ee40-459a-8bda-3d5cb2e53ba3 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Accelerating Large Language Model Decoding with Speculative Sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.344014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.344014Z digest=sha256:654c8c15b9bb0754ab92dbb2c0b1020723f20bd7eec261c7a96d63856e935e11

Observation 5a0484f4-020a-4727-8d2e-9342846f4c11 · outbound

This paper cites Fast inference from transformers via speculative decoding.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Fast inference from transformers via speculative decoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.348481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.348481Z digest=sha256:6c3577b94c75eb626e268cdb52caec5f2dd1508b4504b64844b44656342144f4

Observation 3ba568b7-6fb8-4c60-afe1-3312768270a2 · outbound

This paper cites Prompt lookup decoding, November 2023.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Prompt lookup decoding, November 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.894668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.352902Z digest=sha256:850ef33726b9f815a7ce39046cc65378b3bb5684ed921fcc339caff5ea3fbd3f

Observation ef4c5fae-3d8f-48b2-a03c-032fc212b8eb · outbound

This paper cites Efficient streaming language models with attention sinks.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Efficient streaming language models with attention sinks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.879053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.357195Z digest=sha256:467785463af3aa43563624deed6ad72207f0449971c0b5fcbae83a98fd878844

Observation cf829135-f430-4acc-8754-43b503bed880 · outbound

This paper cites Efficiently Scaling LLM Reasoning with Certaindex.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Efficiently Scaling LLM Reasoning with Certaindex

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.361717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.361717Z digest=sha256:b565584155593050011c8db32d35791f470c3ebd9c25e58fdfd930f9b0f35791

Observation 218cb1c7-c682-42ad-a1f2-aa2bae8309cc · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Orca: A distributed serving system for {Transformer-Based} generative models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.366575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.366575Z digest=sha256:d0e94b56ea965ff8b387f6b784af163be47540638769e444a7f85b688138e374

Observation 8fc98626-c13a-466f-aff9-12bf892ad6e8 · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Fast Distributed Inference Serving for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.370910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.370910Z digest=sha256:422d939114d4b0ea3cc602ab005ac19dcd29dc5f04b28aa40fbe34937202f0b1

Observation 66d38722-36b3-4010-bb3f-fde70f928c6d · outbound

This paper cites Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.375175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.375175Z digest=sha256:42d947dce76f2e951e54cef3dac9d2e4c3d0b6d2ee934fcb3e35cfbaa8a443e3

Observation 39572001-5ffb-47af-b551-e1bd8474c506 · outbound

This paper cites Stateful large language model serving with pensieve.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Stateful large language model serving with pensieve

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.853980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.379599Z digest=sha256:ae8512ebdec170866e27ec4e45460dfa363bdc9f6b166d443acbc29cc85b45df

Observation 519ae05f-2cce-46bd-92ab-866c62479560 · outbound

This paper cites InferCept: Efficient Intercept Support for Augmented Large Language Model Inference.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows InferCept: Efficient Intercept Support for Augmented Large Language Model Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.384225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.384225Z digest=sha256:d9719ba92b8a374ad26f5f8f6c9f0aa7d7acfdd88d320271d30a3f759bc3b6c0

Observation f300aba2-02cf-48da-a93b-cbe612a86f26 · outbound

This paper cites Autellix: An Efficient Serving Engine for LLM Agents as General Programs.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Autellix: An Efficient Serving Engine for LLM Agents as General Programs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.388683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.388683Z digest=sha256:3def124b3570705cedd8c44b5536b4603e38b7c91ac8fa40b3ee62054f958de2

Observation 99c684cf-6b44-4846-b990-b15ee6860f83 · outbound

This paper cites Parrot: Efficient serving of {LLM-based} applications with semantic variable.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Parrot: Efficient serving of {LLM-based} applications with semantic variable

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.837952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.393131Z digest=sha256:31551a757fa70106bfaae0e3422ed718efe6b9aca182d932e5f0314ffb43441a

Observation 72716aa1-a196-4ab0-9289-d329a85fc8d1 · outbound

This paper cites LangGraph.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows LangGraph

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.821121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.397557Z digest=sha256:33797d9f76416a1ef029a1ee0ae996fe2deeab52afbad4fe85b9c112517a1c9f

Observation 9e15dc4e-6d66-4013-92a9-70b26b4662fe · outbound

This paper cites Building effective agents.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Building effective agents

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:45:42.805802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T18:45:42.402250Z digest=sha256:ec117aa28f16e84f213597c2b5b5b50bd02993e94c884254dfd44d1831a3261a

Observation 13f354ff-166e-4bb5-ace0-ee2b53f11b46 · outbound

This paper cites AgentScope: A Flexible yet Robust Multi-Agent Platform.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows AgentScope: A Flexible yet Robust Multi-Agent Platform

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.406791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.406791Z digest=sha256:a4dba950fe5c4f7cb5e996250b1d55b5277dadbd5f4ec46388e71d1c6639a3ed

Observation 887322ec-44c8-4927-a54c-ed5fc257900b · outbound

This paper cites Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems.

KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:45:42.411640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:45:42.411640Z digest=sha256:cdec793603fe6598512123e189152a763636b5f954309f44282c1066a4d5c4eb

Pith citing papers

Observation 26d3d1de-fff5-4a5c-9e1a-38e9331db713 · inbound

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving cites this paper.

Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T22:20:51.838309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:20:51.838309Z digest=sha256:326713541e4720b5e922f44ede2073721aac8fdf863046179f3df329dc721243

Observation 8504fea4-ac1a-4c6f-ad49-df4739e91feb · inbound

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache cites this paper.

ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:46.907699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:16:49.292491Z digest=sha256:dd34b169638224c8b8616472e309370c79023e14a71351085d0f1a3a5f1ec08d

Observation 9bd72f02-e53d-4bdd-a705-6e85b90aa4a2 · inbound

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling cites this paper.

Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:36:36.686395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T06:31:36.776819Z digest=sha256:81ab41dd09bb7b6efc25c418e55c14fd5f7f81ff2d46c1ccb6a7499397d636e8

Observation ed36350b-6157-449b-bb3f-bb64896b7591 · inbound

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference cites this paper.

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:34.256580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T03:55:32.585495Z digest=sha256:fd2f92746fddf9aaa1478d598b83489faecfb598ddc799760ad3e38b2534b529

Observation d380eb5e-e830-48b5-bdc7-32b31d024ce3 · inbound

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving cites this paper.

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.127977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T14:16:46.241126Z digest=sha256:cd8d171d9792198afddc2a631007c23af71960ad26bc1668f9ff261fb6d719e0

Observation 903503ab-0c69-45a8-bbe7-430b6ddc9c2a · inbound

PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design cites this paper.

PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.474892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T00:53:50.578344Z digest=sha256:c942c5a8c2e685c78fbeddeb294f9408b0e878644414dc64128a800862314889

Observation 3d13b212-72be-4504-8766-0d483000950c · inbound

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference cites this paper.

VeriCache: Turning Lossy KV Cache into Lossless LLM Inference KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:12:51.215622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T22:09:15.782104Z digest=sha256:5b4fd820552bcf2ed2c1855ad6e91388a9707b5c383276ee97ecb1aa0aadda8c

Observation d7ec5bef-2697-4fe9-be9a-fa4b5f8e50a3 · inbound

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches cites this paper.

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.016867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T22:08:45.016196Z digest=sha256:542527474ab3c16ef573341890e2e454fe0d533316025b90875041192a11c1b6

Observation 061ca800-78f4-4b34-b6dd-832b565c2b15 · inbound

VineLM: Trie-Based Fine-Grained Control for Agentic Workflows cites this paper.

VineLM: Trie-Based Fine-Grained Control for Agentic Workflows KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T23:56:22.405561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:56:22.405561Z digest=sha256:e90635335c5166663b2a0b3b3be9a8be03d2dda727f3f7bbc3605579440330e5

Observation 22dab550-d3c8-4bb7-9865-251ed0bc1eaf · inbound

Resident KV Claims: A Conformance Contract for Future Reuse under Active KV Pressure cites this paper.

Resident KV Claims: A Conformance Contract for Future Reuse under Active KV Pressure KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:24:45.202521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T14:18:35.209571Z digest=sha256:5abed3b79c4f54441c61ffc684d024b44ab65b08796aff8a44e4b24354b9684b

Observation 57a31b54-9b25-48a8-92a6-c6e5d950bfc7 · inbound

VikingMem: A Memory Base Management System for Stateful LLM-based Applications cites this paper.

VikingMem: A Memory Base Management System for Stateful LLM-based Applications KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.635254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T07:30:42.104455Z digest=sha256:aa95873b64882d0fdf20bf6270a17f991ec4d16ddd89f72db06574ffcdca97b7

Observation aec54df1-77a5-4946-b6ae-20d28c31d373 · inbound

Streaming Communication in Multi-Agent Reasoning cites this paper.

Streaming Communication in Multi-Agent Reasoning KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:06:47.918375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T06:24:07.052856Z digest=sha256:db83a8ae159fb4bf1139b4309cf233b24ce238755a69ccb02e51a75272fe31bb

Observation 4ae630f1-8f27-4dba-8663-26525168fd5d · inbound

Streaming Communication in Multi-Agent Reasoning cites this paper.

Streaming Communication in Multi-Agent Reasoning KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:58:39.204380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:58:39.204380Z digest=sha256:2cab66190789359d8d02a9f0073eea166bba6db5302cefd598bb37a3582ae1ee

Observation 4e8cb190-cd0a-4d88-8fe9-f9124eebd6a0 · inbound

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering cites this paper.

SmoothAgent: Efficient Long-Horizon LLM-Based Agent Serving with Lookahead Context Engineering KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.694422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T17:19:53.560039Z digest=sha256:733748b3bbeabd4bd0564f5669bf6b561fac4e8ecebbb16ca0a1faca794dc5eb

Observation bd742d16-48fb-4143-9404-57f04d3397a2 · inbound

Workload-Aware Caching for Multi-Agent Systems cites this paper.

Workload-Aware Caching for Multi-Agent Systems KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T11:22:57.020958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:22:57.020958Z digest=sha256:7f31837a5b4303333da1b2eb48a43e940eff24c4c73d4dce9d3365c4d77645e2

Observation 0a65ee2b-195e-4223-ba1d-95c47b4382bf · inbound

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework cites this paper.

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:14:12.063684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:14:12.063684Z digest=sha256:83782479de11bdda1504809376915c3781789475ea24410ec6aaa42a94f884ff