Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:57:01.881688Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 6 inbound Pith citation observations for arXiv:2505.03756.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:57:01.881688Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T21:47:00.295144Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T22:05:06.251185Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d77a65af-c5f5-4158-8fbc-7d35818c28b1 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f34dadd4-cb6a-456b-a374-2f65acd35538 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70718582-7ee3-4eda-a965-d2cc5ed84ee7 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Introducing apple’s on-device and server foun- dation models, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f4768bd2-f376-46a5-a399-07b5d37ae985 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management The Costly Dilemma: Generalization, Evaluation and Cost-Optimal Deployment of Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a72d59f2-eeb1-45fc-b5cb-bfa4892ebc69 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Taskmaster-1:toward a realistic and diverse dialog dataset
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1410a725-0eb3-4e76-8d83-afebd3df86f7 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Punica: Multi-tenant lora serving
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4c54425e-6112-41a9-9e66-d24d841ff385 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f314f6-b18b-4715-9eae-df39be8b22a8 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Palm: Scaling language modeling with pathways
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c3a67f-f606-4125-99ec-2ff937571ba5 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Introduction to tpus, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b3a5b57-94c8-458f-8990-c692ddf520e8 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management sglang: A fast serving framework for large language models and vision language models., 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2fe70934-7afc-4584-a05b-49a4a7ce9db3 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management High bandwidth memory, 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation edc0a231-ccfd-4d35-84b1-a2034f778590 · outbound
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8e80c258-b45c-41fb-89d2-9f22bce18848 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Nvidia a100 tensor core gpu, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4167d71a-fbd4-4219-9538-4efcdb01e611 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Qlora: Efficient finetuning of quan- tized llms
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39b5a8a-f1b8-4a92-a694-5d52b6a1857c · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Gpt-3: Its nature, scope, limits, and consequences
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eb411b4-7279-4a1f-9c5d-93c1778190cd · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9486e9e-d77f-4486-85da-e34bb823c605 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Prompt cache: Modular attention reuse for low-latency inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ae4ccf8e-dd1a-4b43-a95e-5831a7b1955a · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LoRA: Low-Rank Adaptation of Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2c5c4d-d506-407e-adc0-556294f77415 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e365f6ed-7408-49b3-a184-184c30990a42 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Chameleon: Adaptive caching and scheduling for many- adapter llm inference environments
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3c5753f6-b4ac-4728-bc25-464a14aa2bc9 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Efficient memory man- agement for large language model serving with page- dattention
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc91c4d-7915-4744-8885-301a44c8a8af · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28475c05-f47c-42e4-969c-d3453d406675 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fff902a-b521-4faa-a448-88f410897a0b · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Prefix-Tuning: Optimizing Continuous Prompts for Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b5eff5-021e-4f5a-9f8e-6772002df314 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb81c90b-306d-477c-9f2e-4d716beb127a · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management DoRA: Weight-Decomposed Low-Rank Adaptation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0398a13-59c0-4c8b-be41-69a2b40332ce · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Instruct-tune llama on consumer hardware using alpaca-lora, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eef52f8a-75f9-4197-9171-bbebf9c66ade · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26d9b191-8666-40f4-9cc5-16cb73f88a6b · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Language Models are Few-Shot Learners
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f6700a8-b82d-4592-a1b5-7e9fd585d7b7 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1a54de16-150e-42eb-a09b-f09f61482d29 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Chatgpt, 2020
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e426f054-1e11-4d0a-a0cb-c532dd4c1ebf · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management torch.stream — pytorch 2.0.1 documentation, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bd27a102-1e9d-4a79-bff1-db2cc4a90ac6 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Moon- cake: Kimi’s kvcache-centric architecture for llm serv- ing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e806169b-dbc6-4605-81f9-914cbc8931d9 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 53bc2ed7-fd1f-43e7-ac63-079b8163d743 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Fast Transformer Decoding: One Write-Head is All You Need
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2197089-74ee-4b9d-8be4-289cfa67e394 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Slora: Scalable serving of thousands of lora adapters
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 29180fcf-06f0-49f7-b4fb-cbbf98c2855d · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bfd6808-f4ef-406d-9234-86793841d94b · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de06853e-ebe8-446f-b414-d009d559bd45 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Discovering finance keywords via continuous-space lan- guage models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 71c9a847-92c9-4836-9e38-d7dcd8350fd1 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management vllm: A high-throughput and memory-efficient inference and serving engine for llms
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b0b5f971-9565-49b9-a289-803c287750b0 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LoRA-Pro: Are Low-Rank Adapters Properly Optimized?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4704b4f-dfea-4f00-bff9-862ee3e83cee · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management {dLoRA}: Dynamically or- chestrating requests and adapters for{LoRA}{LLM} serving
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bd84af8d-a272-4a4e-8be5-ebd876a11f20 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f863be8-a115-493b-86e9-1b1388c262ce · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Orca: A distributed serving system for transformer-based generative mod- els
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c247a40-3f34-474e-9abf-c3d79632cc3c · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Stateful Large Language Model Serving with Pensieve
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2844c1-3634-45de-b2a6-b6c669d62409 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 704ea79e-3b71-4257-89f0-78163f6d5e8b · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111a0041-f61d-4710-8eb1-2b8f4258e3c5 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea5a426-71ac-40ec-8787-d347aea43be1 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Faster and cheaper serverless computing on harvested resources
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af570e82-87f7-43d2-a34f-3514cbb703b2 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be2da8b7-7aae-430a-80f0-7c12da66791f · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Judging llm-as- a-judge with mt-bench and chatbot arena
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4000ee5f-5a5a-4768-85d1-4e767c690e33 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Efficiently programming large language models using sglang
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab5219c-9d8c-4cfc-97d3-ab0f472ae1f9 · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7075f73d-61e6-4f70-85b1-d89fbb0c83ae · outbound
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7de1d06-576c-40c4-9f24-4a0533830e9c · inbound
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 73fcc33d-1d8e-4967-b5f1-28bf0cada3ec · inbound
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 17a2c2dc-19e2-408e-b1c8-aedd740e52ba · inbound
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4c19f3af-b35e-44ce-9b08-b052bb18a265 · inbound
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9980b588-d301-4673-af44-6c390837a1b1 · inbound
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d4ac3560-e4c6-4f14-a0bf-fe5cb1507393 · inbound
PreFT: Prefill-only finetuning for efficient inference Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.