Pith. sign in

Paper Citation Record · LEDGER

Efficient Remote KV Cache Reuse with GPU-native Video Codec

As of 6 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2602.09725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.09725 v3

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T05:21:04.555356Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact28
  • verified fuzzy52
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 16b0dc06-25ab-4d1a-af35-6aa1eb66a78d · outbound

This paper cites https://github.com/vllm-project/aibrix.

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://github.com/vllm-project/aibrix

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.173618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:c0cbd1cb07fae05e6c627d65e916a334adb09c9f330e98784d695a84bfdb9eaa

Observation dea069a6-609e-418c-b942-1c74fd2ee5b4 · outbound

This paper cites https://aws.amazon.com/ec2/ instance-types/?nc1=h_ls.

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://aws.amazon.com/ec2/ instance-types/?nc1=h_ls

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.124589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:172974c696a5a7f6e14f515c951088fd21b9a3b5be06fe095122ebfc0c33d176

Observation e3fd3335-166e-499c-a5c8-a064a61f4a71 · outbound

This paper cites https://gpuopen.com/ advanced-media-framework/.

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://gpuopen.com/ advanced-media-framework/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.206723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:6dcc4b3091f8ee6a1b269ad64c18b7cce4b7a798a1b4dc91cb67400ebe77f4b9

Observation 886ed6d6-4035-4621-9a94-d7ccb1d15925 · outbound

This paper cites https://www.intel.com/content/ www/us/en/architecture-and-technology/quick-sync-video/ quick-sync-video-installation.html?wapkw=quick%20sync%20video.

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://www.intel.com/content/ www/us/en/architecture-and-technology/quick-sync-video/ quick-sync-video-installation.html?wapkw=quick%20sync%20video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.167702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:f97051edc35de96ed089205210bc064d2bed640c8fb643be07ef048590c2c1fa

Observation daa9268c-7b3b-4f93-a83a-e6ae688ae5e3 · outbound

This paper cites https://developer.nvidia.com/video-codec-sdk.

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://developer.nvidia.com/video-codec-sdk

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.149402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:20be4fd7b0278961394c1143d68d4d9efcf88c81c59cddc5a18036236b73e303

Observation 145b81ba-1ffa-4656-a596-1f94928ea03e · outbound

This paper cites https://huggingface.co/01-ai/Yi-34B, (Accessed on 02/04/2026).

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://huggingface.co/01-ai/Yi-34B, (Accessed on 02/04/2026)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.164373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:3a58e3d11d961018ff29800e7e1c77eb0656078557b954ce3a10c988e1e74047

Observation bfe42054-dd66-4c44-9eda-48d30fefb8db · outbound

This paper cites https://cursor.com, (Accessed on 02/04/2026).

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://cursor.com, (Accessed on 02/04/2026)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.139553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:88557b8a912e0536abf97e6fda5e2d1b13c5c62b8b5ad032fcac3c6c7d9d9ddf

Observation 88544977-f3ca-4af4-9bf7-8cc0fd9c43b4 · outbound

This paper cites https://www.ffmpeg.org/, (Accessed on 02/04/2026).

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://www.ffmpeg.org/, (Accessed on 02/04/2026)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.152280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:610a293bb80b7eed05d9ce51d3993214ef55143b01b1b1ed4b89b3f6b8c3d964

Observation b9f18ed9-c9fb-4629-a897-9235ef199702 · outbound

This paper cites https://gstreamer.

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://gstreamer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.121443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:f469a745390ebb7f352431a38f1af1c9e907a7f807b55d39d1fabd42627ce1f3

Observation 5a267a88-acd9-4a4e-b9ad-51c31eb0d472 · outbound

This paper cites https://huggingface.co/ LargeWorldModel/LWM-Text-Chat-1M, (Accessed on 02/04/2026).

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://huggingface.co/ LargeWorldModel/LWM-Text-Chat-1M, (Accessed on 02/04/2026)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.213402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:7e42264138987ecb5f0b668296d8f7a806c768ec026f0a1145d8c02d3493618e

Observation 50341256-09f7-457f-b0e1-513aeb9ed437 · outbound

This paper cites https://huggingface.co/ meta-llama/Llama-3.3-70B-Instruct, (Accessed on 02/04/2026).

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://huggingface.co/ meta-llama/Llama-3.3-70B-Instruct, (Accessed on 02/04/2026)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.133594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:bc2429ec0e4ba968127a53b62e1734f479a1d709dc1e0bf465e2f87ed866791c

Observation 70d88b88-8b0e-400d-8b5c-543093e823c3 · outbound

This paper cites https://www.anthropic.com/news/claude-4, (Ac- cessed on 07/14/2025).

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://www.anthropic.com/news/claude-4, (Ac- cessed on 07/14/2025)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.216595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:7c37df44d2940232b3281ff0ef1cd77c304040f492c4ff21b418270f3753ca7a

Observation 7463500e-f9f6-4f72-a982-5daa2a272365 · outbound

This paper cites https://github.com/LMCache/LMCache, (Accessed on 07/14/2025).

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://github.com/LMCache/LMCache, (Accessed on 07/14/2025)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.081024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:0aad11aac1176bb2ba0d64cdb58b1ce9868e807684a168328268f44385f514fd

Observation f28bebbf-25ab-450e-a47c-bbbc2120d215 · outbound

This paper cites https://docs.anthropic.com/ en/release-notes/system-prompts#august-5-2025, (Accessed on 08/05/2025).

Efficient Remote KV Cache Reuse with GPU-native Video Codec https://docs.anthropic.com/ en/release-notes/system-prompts#august-5-2025, (Accessed on 08/05/2025)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.142812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:421043364b600c5740c66155aa882687b63756e20bcefb1d44f1bb27aac5fdc2

Observation a332ce39-8548-49f3-ac42-e366a32fd4fe · outbound

This paper cites Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve}.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve}

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.136650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:a02982e07bfae637312c29de0b207764bd7dec14b368071bf50233628dedf14b

Observation 16bb9141-1fe6-4ac9-86c1-ea184d8d94b6 · outbound

This paper cites L-eval: Instituting standard- ized evaluation for long context language models.

Efficient Remote KV Cache Reuse with GPU-native Video Codec L-eval: Instituting standard- ized evaluation for long context language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.088146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:28ac3cf95a62591081b4169f15da4cedf7e352a40dc5c26c342fffcdd400f90f

Observation 1a04a44e-f524-4d10-bf81-3e756d3149e6 · outbound

This paper cites Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Longbench v2: Towards deeper understanding and reasoning on realistic long-context multitasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.241311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:12591cae58321a0b24838b88e00d10376fb43cd36cf9e00599a545c7ae7ccd86

Observation 2d24e5df-0909-4e2e-8829-4b1b957dfb21 · outbound

This paper cites Longformer: The Long-Document Transformer.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Longformer: The Long-Document Transformer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.793916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:4c51c9f9a5c92a466f8e71e317150fa56cd9dc7ee5229e908934ab948131003f

Observation 30c4df66-ecdd-4a14-a62f-51cd8acff193 · outbound

This paper cites Improving lan- guage models by retrieving from trillions of tokens.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Improving lan- guage models by retrieving from trillions of tokens

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.180036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:6a0d33678e605eb742c6833cea70ac9480e6102a8fd520204f98c7915e40d2fd

Observation 95241c4d-399d-412f-baec-58c58ebd7c46 · outbound

This paper cites Moe- lightning: High-throughput moe inference on memory-constrained gpus.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Moe- lightning: High-throughput moe inference on memory-constrained gpus

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.091507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:bad2d7350e7873537b9a9f335bea916e3648c9dab87c34daa20c17e2a5d42663

Observation ec7be858-6faa-44cd-9403-22f9a60797da · outbound

This paper cites Locality-aware Fair Scheduling in LLM Serving.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Locality-aware Fair Scheduling in LLM Serving

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.829821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:48f7210970be4968fce52419047d90ab8c47b1fceb7d83ef69fc3dbfafcb53b5

Observation 567c1bd2-f5fa-4467-8056-805f3bbbc3e1 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Evaluating Large Language Models Trained on Code

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.887472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:615e014b6457764c1921aa7fdba2887190e25c35f87691305f84eac2a6e84602

Observation 614ea6f7-c76e-4745-a97e-a29f13cde320 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.892859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:98f699d04fd44dc7eda8bf0b144041f6df48f67a3cdcff977fd1aa3fb31d97f6

Observation a648d72c-3d0a-4e5b-b6d0-a0ffe05a45bb · outbound

This paper cites Dynamo distributed kv cache manager (nvidia dynamo sdk v0.2.0).

Efficient Remote KV Cache Reuse with GPU-native Video Codec Dynamo distributed kv cache manager (nvidia dynamo sdk v0.2.0)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.225697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:e75307d88a28fb4a631acdf8c719db8b61eb568768a68586079f5a1b1b737110

Observation 2fa5b435-5951-4660-9e16-1018a5ce1714 · outbound

This paper cites Llm.int8(): 8-bit matrix multiplication for transformers at scale.Ad- vances in neural information processing systems (NeurIPS).

Efficient Remote KV Cache Reuse with GPU-native Video Codec Llm.int8(): 8-bit matrix multiplication for transformers at scale.Ad- vances in neural information processing systems (NeurIPS)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.210391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:ddc90e0dc5f71a74a14e63fc14b7f1e899f06afc211d6effbbaed51b2a3fe30f

Observation e3f58898-ff02-4e6d-93a3-2df863fba275 · outbound

This paper cites arXiv preprint arXiv:2510.03215 (2025).

Efficient Remote KV Cache Reuse with GPU-native Video Codec arXiv preprint arXiv:2510.03215 (2025)

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.787925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:9161425bea1913bc70ba8eccbfe3d5906f84b6c695233743064ec80bf00d542c

Observation eaa49f0c-d138-4300-83cf-0a5869018229 · outbound

This paper cites In18th USENIX Sym- posium on Operating Systems Design and Implementation (OSDI 24), pages 135–153.

Efficient Remote KV Cache Reuse with GPU-native Video Codec In18th USENIX Sym- posium on Operating Systems Design and Implementation (OSDI 24), pages 135–153

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.228885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:f059c5e412efe7efc6b68b7b9f4c5baa73f76379473fec96f2a0f4107bbe06a6

Observation 7e51f887-4df5-4016-950a-705822ccc5d7 · outbound

This paper cites {Cost- Efficient} large language model serving for multi-turn conversations with {CachedAttention}.

Efficient Remote KV Cache Reuse with GPU-native Video Codec {Cost- Efficient} large language model serving for multi-turn conversations with {CachedAttention}

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.183327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:3053d3ef06355af3fc2561bcd0c8cea6a4fedae9f9f828b70bb0a834c5e5e4ab

Observation a53159e2-8608-4db6-bf4a-446ddfc79123 · outbound

This paper cites Fast state restoration in llm serving with hcache.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Fast state restoration in llm serving with hcache

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.170680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:83378d2332c0d8134403f5ffc80f3358c77b704369082f45a83febd814049962

Observation 0963909e-d08c-4d8d-a699-ec5caa12d701 · outbound

This paper cites Prompt Cache: Modular Attention Reuse for Low-Latency Inference.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Prompt Cache: Modular Attention Reuse for Low-Latency Inference

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.751793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:c99011bd5ff6f20c766bf6a110fcbb332675543c77e158b21c43d7cda6e7dab7

Observation c0a2eecc-af91-47a5-88c0-6e4ffbf90283 · outbound

This paper cites M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity.

Efficient Remote KV Cache Reuse with GPU-native Video Codec M\'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.819213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:39ea2adf4e02f4b9156bf429949737a51976672fe165222d5d39acd4f505b545

Observation eeda7cff-0e9c-496d-8e8c-669ae375bfd2 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Efficient Remote KV Cache Reuse with GPU-native Video Codec DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.813297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:d71eb3e64b1669c1dee76ebd4d00d59b0487a02a52c93bbfdde5c3d0dbf82870

Observation 45f92272-9686-4023-b9f4-62d136b0e0bf · outbound

This paper cites Metagpt: Meta programming for a multi-agent collab- orative framework.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Metagpt: Meta programming for a multi-agent collab- orative framework

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.102436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:d833fe9ca89a6c6312dcb288c42d8d3f40b996d9ea588896420d64bc16d643e9

Observation 8974ef50-d1c1-4ec2-8af2-5059d1492dfc · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quanti- zation.Advances in Neural Information Processing Systems (NeurIPS), 37:1270–1303.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Kvquant: Towards 10 million context length llm inference with kv cache quanti- zation.Advances in Neural Information Processing Systems (NeurIPS), 37:1270–1303

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.193754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:b90f30ad68f9d46b246ae2315723992c50a0eb7f2b05879685c2a6af107987c1

Observation 47bd091b-30b8-416a-a99f-bbff14be2f8f · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

Efficient Remote KV Cache Reuse with GPU-native Video Codec MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.835574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:59c424704a87d718aca05dde36456196a6db2f7b48bb681e6d734946beeb77c9

Observation 589b64f2-1f67-4ac8-a289-76ca53d075eb · outbound

This paper cites Accelerating llm serving for multi- turn dialogues with efficient resource management.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Accelerating llm serving for multi- turn dialogues with efficient resource management

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.130640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:8ee448dff3a19d1e91dfe1003d70d90cbe1565372b1ba5bcc46bf17154ab4fc1

Observation d444ba08-02db-4664-8dfd-df17a1634937 · outbound

This paper cites Efficient kv cache spillover management on memory-constrained gpu for llm inference.IEEE Transactions on Parallel and Distributed Systems, 37(1):90–105.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Efficient kv cache spillover management on memory-constrained gpu for llm inference.IEEE Transactions on Parallel and Distributed Systems, 37(1):90–105

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.105825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:7ec4eef2940f6b3371b766db4ec0686249c80ae35cd8bb0abc8f883e1eaaa234

Observation a5d1a4ba-4053-429c-9b86-4183c040b589 · outbound

This paper cites De- mystifying cost-efficiency in llm serving over heterogeneous gpus.

Efficient Remote KV Cache Reuse with GPU-native Video Codec De- mystifying cost-efficiency in llm serving over heterogeneous gpus

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.127724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:ac415e7800e335526387d203197ae0639379da7438fa885144c8fa225e60434b

Observation aaef1185-2944-46aa-a466-43a06277731e · outbound

This paper cites RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation.

Efficient Remote KV Cache Reuse with GPU-native Video Codec RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.800333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:b1b48cdafee0474e021a0bab15ed097ae0bb488c13c818c4940469d27cf63759

Observation 444be752-dc4d-438e-80ba-f0b259a576f9 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Gonzalez, Hao Zhang, and Ion Stoica

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.155312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:4287997f85375af86e14bfb8acedfe0587d4fe0d8bfe420f3c5f4c3c73c0402f

Observation f0efd002-92e0-498a-89f4-640e8b124f82 · outbound

This paper cites Robomemory: A brain-inspired multi-memory agentic framework for lifelong learning in physical embodied systems.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Robomemory: A brain-inspired multi-memory agentic framework for lifelong learning in physical embodied systems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.095917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:912c85309adcb42d946d64ab394c470f3d202b9466d1e801ab60fff2a683f5d7

Observation 2ba173be-c2d4-4f9f-9a69-969fbc338cac · outbound

This paper cites fabric-lib: RDMA Point-to-Point Communication for LLM Systems.

Efficient Remote KV Cache Reuse with GPU-native Video Codec fabric-lib: RDMA Point-to-Point Communication for LLM Systems

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.734761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:22ca01899e59f5e6e780f6c04f686a6d0a741e67d87285caacd89e39d35f6bc4

Observation da0b93bd-fe80-43ac-b4df-b32a3ef2febc · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.746402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:67e5878e51148ef252dd29c24d8a4ef70bedfea839d76dd5efe5f73f7e1a4db6

Observation 9931bd86-5c74-4e87-a399-d56e8b1d0925 · outbound

This paper cites Parrot: Efficient serving of {LLM- based} applications with semantic variable.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Parrot: Efficient serving of {LLM- based} applications with semantic variable

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.238170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:ad01a4aba24f51596f7507f07fd0f6a7773697b1502956790b18c09437d2edb2

Observation 3d4b8d82-e24c-43cf-b841-26eb029c6102 · outbound

This paper cites Onetwovla: A unified vision-language-action model with adaptive reasoning.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Onetwovla: A unified vision-language-action model with adaptive reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.870988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:40667580faeab6535e584718ddefcbf04d6f04fc50a0b12e891ddc26d6de269f

Observation 3b1c03a8-84d9-4c24-b10a-2dad4910ebb4 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.203197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:a41dbf337840435e095f0a1dd83d09a41a9ead9fe6caf660daa8fb7999d0ddb7

Observation 40db504b-8c33-481f-954a-cbad20b45977 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

Efficient Remote KV Cache Reuse with GPU-native Video Codec KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.806239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:a0f612d59ac914245e2c2ef4376f67050a42783943349be5832fb88da921b8df

Observation de4ab9a6-4c5d-4870-967f-76bf61ebfb02 · outbound

This paper cites Sky- serve: Serving ai models across regions and clouds with spot instances.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Sky- serve: Serving ai models across regions and clouds with spot instances

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.112260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:df8b8eb8c1d674a24145c83ab633264a9660735f8cff0ec09000e7b6c5a1aab2

Observation 88b1fc74-c64e-4e3d-b7ed-56c0522653bf · outbound

This paper cites Helix: Serving large language models over heterogeneous gpus and network via max-flow.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Helix: Serving large language models over heterogeneous gpus and network via max-flow

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.186493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:6114f7a831321b2b4a3f34b2a34b2f58684993c7ef3d58339651017a2da92797

Observation 0b396630-4092-4376-8719-aa3068bd1fbc · outbound

This paper cites Spotserve: Serving generative large language models on preemptible instances.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Spotserve: Serving generative large language models on preemptible instances

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.244109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:233909e7c7eb5825bfdd6f01391046a45467a9b8c21486e699724ba66b140da8

Observation da8dd679-23b6-47f7-b5a0-7e1daeef5325 · outbound

This paper cites InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference.

Efficient Remote KV Cache Reuse with GPU-native Video Codec InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.775393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:0b6a25b4a167cfb249eb9e1040c29e09b4227bf9a794372081d80d65fa25a822

Observation b1be52ca-d1ad-4165-aa3d-c9a04788717f · outbound

This paper cites Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Llmlingua-2: Data distillation for efficient and faithful task-agnostic prompt compression

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.196968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:7e00f2dbaa1ff05f53b0e9c3da269a5cc52e03a0ef8dc8aa18be5457c5cd57f8

Observation 04a6d11d-bb54-4a54-b13d-b5accdc00b3a · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

Efficient Remote KV Cache Reuse with GPU-native Video Codec ChatDev: Communicative Agents for Software Development

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.876576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:d457c0ac9008ac3be43defa6c7840e1cec88a875cb6e372129277620ae378c40

Observation d196fdc6-70d5-4ee0-b1cb-d0396b260e33 · outbound

This paper cites Mooncake: Trading more storage for less computation — a KVCache-centric ar- chitecture for serving LLM chatbot.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Mooncake: Trading more storage for less computation — a KVCache-centric ar- chitecture for serving LLM chatbot

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.222898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:563b78b6a95d97a7e6d1f1f189f0daf1c9dc61beda4d7abea945bb8c84ec9246

Observation 0776aa80-46d5-430c-9654-fecfcdbe81d7 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Code Llama: Open Foundation Models for Code

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.853630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:27359b2d347e25e78a2281a59d9d23ebba14239b8b5a831a3218b03d7e943cd7

Observation edaeb9ba-f87c-40b8-843d-63fab71702a7 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

Efficient Remote KV Cache Reuse with GPU-native Video Codec MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.847893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:de768cd6989728a0c95e51b3ddd864f6520be8cfb8025b891096c9f9d32b4919

Observation 80d92b4e-2a3b-48e4-896d-9da1d4791dab · outbound

This paper cites Preble: Efficient Distributed Prompt Scheduling for LLM Serving.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Preble: Efficient Distributed Prompt Scheduling for LLM Serving

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.865185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:9e3167e63ef9d025cb9612e25580629a104a2150333f4e5e8f1f4f67e8311d86

Observation 73f9dbaf-db55-4583-a641-b1285cbd8d41 · outbound

This paper cites Déjàvu: Kv-cache streaming for fast, fault-tolerant generative llm serving.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Déjàvu: Kv-cache streaming for fast, fault-tolerant generative llm serving

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.109078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:fbf9fe055afc3e40b2e371dee4f6613b1b8eaa612340bb5adbcf9ad4625bef1b

Observation c467e934-24af-45f1-bda0-9945043202ea · outbound

This paper cites Overview of the high efficiency video coding (hevc) standard.IEEE Transactions on circuits and systems for video technology, 22(12):1649– 1668.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Overview of the high efficiency video coding (hevc) standard.IEEE Transactions on circuits and systems for video technology, 22(12):1649– 1668

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.115279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:ecf951d3195b34d094c8a73524bd7a51ed9e966e89030854d93ab5511ebff93f

Observation e09c6b35-6379-405f-bd33-d46117d8d38b · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Llumnix: Dynamic scheduling for large language model serving

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.235126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:0e7682a0db940f6af2f42973eb0aff41558f900f2251befff22d568c309961db

Observation c50e5ae4-3a50-491d-bba8-d258b56d7e75 · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context llm inference.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Quest: Query-aware sparsity for efficient long-context llm inference

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.099299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:8db54728d7a2d268738f56f4cb16f7adea3aa6f9e65b4950cedd349113600c53

Observation f2c86f34-1ba7-4c0f-81ca-23777a0c8d1c · outbound

This paper cites Tetrisched: global reschedul- ing with adaptive plan-ahead in dynamic heterogeneous clusters.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Tetrisched: global reschedul- ing with adaptive plan-ahead in dynamic heterogeneous clusters

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.176878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:f9e7699dbd350b6be33bd92ef0a49c14d29e346eca4b0797b5ab617bc7682d7b

Observation 0f27d7c7-cf12-4e89-9ce8-86d4dd62538e · outbound

This paper cites AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents.

Efficient Remote KV Cache Reuse with GPU-native Video Codec AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.780975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:c85e52703340dde6c6b03bf0956ed73f4b457d84c8ffebca4ae9e2ada95c6681

Observation b26a37f7-def2-4cd3-a671-072a01779de7 · outbound

This paper cites Kvcache cache in the wild: Characterizing and optimizing kvcache cache at a large cloud provider.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Kvcache cache in the wild: Characterizing and optimizing kvcache cache at a large cloud provider

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.145957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:46e08d75e1c4b74015fd09695a6fdc6ed404924c66e8b156b4db2aceeb21f39d

Observation a4e3188a-afa1-4c01-bbf3-d72c1fdd7935 · outbound

This paper cites Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.158763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:efc694695e3b0d04f4955445878232ce2b22f44f5c507a39a01b6c4bd774f083

Observation 6c1e2dd8-680a-4307-8f16-834f8f853e2d · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Fast Distributed Inference Serving for Large Language Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:23:46.844514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:454820c59b38d7867968ad140186cd876bb7ad58d55fbd78381e2af0dcf937b3

Observation 6e298152-0a0d-4908-bb18-ff496d1783b6 · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.762421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:272b5720303016a0dd5de2768c3e3c36878b8e3c5b7500be5ea7e30ff3d41a99

Observation 5b77d895-1f23-4596-a1d4-94bcc2ecce8a · outbound

This paper cites Shadowserve: Interference-free kv cache fetching for distributed prefix caching.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Shadowserve: Interference-free kv cache fetching for distributed prefix caching

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.859844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:4ddbd1bf44533dc757fe84a5b54335f5917f4805e7c068bea89a75686c29c612

Observation c64de8e7-7e33-486b-9c10-1ef52947c9d3 · outbound

This paper cites Efficient streaming language models with attention sinks.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Efficient streaming language models with attention sinks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.231938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:b9e195fe7b47d6ac7515d57097efc552b5c075ed1ae9f67d5bafe6ba1890279d

Observation bca2aa03-58fd-45a2-adcc-31e71317e029 · outbound

This paper cites LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management.

Efficient Remote KV Cache Reuse with GPU-native Video Codec LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.824323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:74efd642da874358211477c39b4ed52f46ba3d3b7f3b76963dbc2137aca2b5c0

Observation e45a3712-8758-4215-9544-bccd879a65a2 · outbound

This paper cites an unresolved cited work.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-16T05:22:23.161625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:5df702288b6b852542e5c34cd4d51c9c5e3aa67fa2e7fadac2bb363338a390de

Observation 242b5473-934b-455b-ada8-44d2fc4aa2d4 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.190414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:4f9c98049ac73fb4e559f54fd64e788f669aa93cc4098712d1e2dbebeb6655dd

Observation bc12ec78-6281-4597-b4d6-568c6019802a · outbound

This paper cites ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition.

Efficient Remote KV Cache Reuse with GPU-native Video Codec ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.841918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:bbb19dc5a291995c26a64212baa506f43ebbfade28d276dc76c8fb2f1704ceb8

Observation 39fd06e3-2d01-4f35-bfcf-2a712a50c161 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Yi: Open Foundation Models by 01.AI

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:22:22.740147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:cf63250702cbf5673a3fd7a6914565decfe5f0c1ddcca097960411423f20c995

Observation 1b010bf1-3a32-4ea7-aec2-f73cacb8ecd6 · outbound

This paper cites Orca: A distributed serving system for transformer- based generative models.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Orca: A distributed serving system for transformer- based generative models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.118468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:5fde10889a8c2056ac9e3c979842cf16d952d73ff9ce2a9922f484493fbe4770

Observation eb0bfdc0-1eb3-4a0e-bd04-26d3a55565fb · outbound

This paper cites Stateful large language model serving with pensieve.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Stateful large language model serving with pensieve

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.219727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:f1be795160dbf25cc4f4aa14fdc1fd5e95c6eb2911bcd8f349c72edf666d9968

Observation 72cd7fc8-7cd6-4738-900c-4226315e166c · outbound

This paper cites LV-Eval: A balanced long-context benchmark with 5 length levels up to 256K.

Efficient Remote KV Cache Reuse with GPU-native Video Codec LV-Eval: A balanced long-context benchmark with 5 length levels up to 256K

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.757266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:cf17cae82ffae4e557ed68540e6d01312941c3b9a2bd3d0ed259d366830e6e14

Observation d0fdf0c6-27a8-4d76-89dc-bbace123bdfc · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative infer- ence of large language models.

Efficient Remote KV Cache Reuse with GPU-native Video Codec H2o: Heavy-hitter oracle for efficient generative infer- ence of large language models

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.200124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:aae84cb991a2748bedc3f17ce3ba77d0f69da841e5b233532a386315e31f4a7a

Observation 5d00a158-4ba2-461e-90b7-2779c067af74 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Gonzalez, Clark Barrett, and Ying Sheng

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.249632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:c77b232ec3e54cc039339c88a577758fbdcdc2bd3d68602610ed6742bdbf80a7

Observation 7c227fd4-2eef-4924-987d-240212455ccd · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:22:23.246944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:bd745a1a94c296ac268232f270719ea0fa995c9ca9146fed515cedf015130e2e

Observation 38174ce9-9863-439c-8ea3-8e0e9220c8d3 · outbound

This paper cites Self-Taught Agentic Long Context Understanding.

Efficient Remote KV Cache Reuse with GPU-native Video Codec Self-Taught Agentic Long Context Understanding

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:22:22.882351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:21:04.555356Z digest=sha256:6e7ee3f01d82d4aee85b4e4ab6371e50b4d15dd6febdda44bbef35e8a5192a47

Pith citing papers

No inbound Pith citation observations are available.