Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2607.08057.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9b862f76-056f-4a5d-acfc-141e73d7b3bf · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Reza Yazdani Aminabadi, Samyam Rajbhandari, Am- mar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, and Yuxiong He
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a4e0bd6-55d9-4de2-b720-7ea7ca35cb23 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization QAQ: Quality Adaptive Quantization for LLM KV Cache
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 037ea3a9-2cf1-49e0-a4c6-45d98885b7a1 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f879b246-8df6-4573-88c7-a7473d278f62 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Deltakv: Residual-based kv cache compression via long-range similarity
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5c2f566-6573-44f2-83fe-7b5390964c35 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69d14e1c-8af3-4b6c-a72b-c7d151241648 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization A Survey on Large Language Model Acceleration based on KV Cache Management
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b1535c11-1425-4769-8eca-0e811aa67576 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 43c055a7-a8eb-4daf-8c8d-f31ee9ecee5d · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c82e1a2-e934-47ec-9a94-dfe6b9bd0471 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b20c001a-d8b3-4706-a16c-cf3d439c60e5 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Idris O Sunmola, Zhenjun Zhao, Samuel Schmidgall, Yumeng Wang, Paul Maria Scheikl, and Axel Krieger
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed5e878e-f0e7-4d67-bf39-9f4e66f85dd2 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Surgical Gaussian Surfels: Highly Accurate Real-time Surgical Scene Rendering using Gaussian Surfels
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 972cf1f8-48ed-4582-8b16-a949690228db · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization A Survey of Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d8bcde0-6f1a-4a16-8ae6-3343673db9b7 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization A Survey on Efficient Inference for Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 409dd5e7-b0a0-475c-8bd1-abdab80afb19 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization retain” vs. “evict
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10db3a9b-75f7-4d17-b12d-d05574827fab · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization BALI (Jurkschat et al., 2025) measures LLM inference across six frameworks or acceler- ation approaches
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 758c8532-d83a-4b3d-9265-a37e5dcbd14e · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Workloads.A future sKis benchmark should cover the following three workload types to stress temporal, spatial, and structural KV behaviors:
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9588fd59-2e21-448c-9c23-766403368ad9 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9112735-8597-47ad-b223-5ed99926fb14 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 794abd52-105c-4163-aa50-f59968e0376f · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Reporting standards.In addition to the basic in- formation like model, hardware, and configuration, we recommend the following reporting standards for sKis benchmarks:
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 34921c23-a415-4ec5-bdfc-e223dcde3917 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4db01e9f-5aa5-4761-95aa-ca750a5371d5 · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization memory curves for structural methods to reveal the trade-offs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1abb0e95-aa48-47ef-973d-575eacdfdf2d · outbound
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.