Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:08.498820Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 4 inbound Pith citation observations for arXiv:2501.06709.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:58:08.498820Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:45:32.847427Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
44 of 44 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 70838d38-4447-49df-a558-80af7a61cc6b · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Brown, B
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e9b3d3ab-052a-46ce-b479-6114c41f6f03 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2642f22a-d560-412d-a3b7-041a69b91997 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LLaMA: Open and Efficient Foundation Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b27bae-ca69-4c32-aab8-045993c3a9e1 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Characterization of large language model development in the datacenter,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a9b7d5f8-4f80-4113-9703-d45e9911c44a · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Orca: A distributed serving system for Transformer-Based generative models,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5a53d823-1a65-4e4a-859f-72126ad12211 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Splitwise: Efficient generative llm inference using phase splitting,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d3401967-7f6b-4079-82e4-878eb18686a4 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Optimus: Warming serverless ml inference via inter-function model transformation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 995dcd7f-b0f1-4a81-8ef3-f9f7b29ef31b · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Otas: An elastic transformer serving system via token adaptation,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c51ead8a-cd79-482f-9ad7-581802a112da · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 49e0f30b-0397-499f-9da2-81df2420cc4a · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficiently scaling transformer inference,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b98744f7-ea30-4b0f-8c29-375784e0e37e · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LongloRA: Efficient fine-tuning of long-context large language models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c4a6f511-e017-42b4-aa5f-014cac17659c · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient memory management for large lan- guage model serving with pagedattention,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 36fe6834-c623-4a4a-b6ee-baf3aa39282c · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d00bb195-ef32-4aaa-85a2-30ee35428d73 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management H2o: Heavy-hitter oracle for efficient generative inference of large language models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f871c23b-05d6-4ad8-9056-f3344e591c07 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient streaming language models with attention sinks,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c370e43a-eb65-47c6-908d-b817162cd06d · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f6791e28-2a64-42b7-901d-bd756b1bba73 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d8c44a-c3e6-421b-9e35-83e93800a937 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Model tells you what to discard: Adaptive KV cache compression for LLMs,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 92d5fa76-fd6b-45b5-8513-c5cddeeb8fac · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Flexgen: high-throughput generative inference of large language models with a single gpu,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 21a3dae1-09ea-4da9-9f75-686ecf9d61e7 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Infinigen: Efficient generative in- ference of large language models with dynamic kv cache management,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1668b1ff-fff7-482d-9975-177f6099fe7d · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Cost-Efficient large language model serving for multi-turn conversations with CachedAttention,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 17a3a078-f1d3-4e2a-a5b9-b39cccd494d5 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Deepspeed- inference: enabling efficient inference of transformer models at unprece- dented scale,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e7f660de-3832-4453-8725-a33dbb518132 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Llumnix: Dynamic scheduling for large language model serving,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9057a506-236a-438c-90e5-4d64452995a6 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Serverlessllm: Low-latency serverless inference for large language models,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19df26f4-727a-4b8a-b54f-a3a7326ff70c · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Turbotransformers: an efficient gpu serving system for transformer models,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7e83071a-3a68-4526-94d9-b8d3d42306be · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Taming throughput-latency tradeoff in llm inference with sarathi-serve,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3cc9f462-3e53-4d19-876a-a09fa998e019 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a9042608-ac4d-4cde-97b3-35f7cd5fc350 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3405f68f-6281-44eb-ba11-2d40878f4747 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management dLoRA: Dynam- ically orchestrating requests and adapters for LoRA LLM serving,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7546fca3-7b42-40bd-9489-c5b6ea9e92ce · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management LMSYS- chat-1m: A large-scale real-world LLM conversation dataset,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2f10efc9-75f1-4584-8455-1b3a8a0032fb · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Wildchat: 1m chatGPT interaction logs in the wild,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 207f5195-686a-466c-a0d1-ae5395559b42 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Judging LLM-as-a-judge with MT-bench and chatbot arena,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7b768cf5-06c4-4aab-8dc5-4266b15924aa · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Koala: A dialogue model for academic research,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37831d52-d151-4bfd-b0d0-10ecd0a38cdc · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Response length perception and sequence scheduling: An LLM-empowered LLM inference pipeline,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 53e40a07-cfdc-4642-8716-f2bc105a4cce · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f145516f-1315-48cf-b63d-90fa4948dcbb · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Adaptive resource provi- sioning for the cloud using online bin packing,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4abe209a-994f-41df-9d29-5bf5a9aac970 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Efficient online strategies for renting servers in the cloud,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 71d93163-586b-4e8f-8e20-fef44de7b11d · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Powernap: eliminating server idle power,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c54bf0-4ef5-498b-9932-3f39f792bc10 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Easy, fast, and cheap llm serving for everyone,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e58c2e49-e9eb-416a-a8f6-336b99cdc3ec · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Ray: a unified framework for scaling ai and python applications,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8ecb158a-0e6c-4f84-a2c1-6053038575de · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Gloo: Collective communications library with various primitives for multi-machine training,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a36461d3-72c8-4209-b669-8f95285db950 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Openai platform document,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8f967afb-874f-4d4b-957e-8dee2255b2e4 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Anthropic platform document,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4a2d838e-b267-4d3b-8d9e-a7e1500860a8 · outbound
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management Available: https://github.com/ray-project/ray
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 01ad7691-2c93-438b-9311-2acf41d527b1 · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
Reference 260
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80459404-6387-4874-b2a6-aef638b43d33 · inbound
Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b31f61f-9a83-4899-bc23-8de4f04438e7 · inbound
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 87c72b18-1423-4970-92e9-e9b572cdfcb1 · inbound
LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.