Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T20:10:09.020308Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2602.24044.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T20:10:09.020308Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a233e350-3aef-4861-bfb5-7c5561a4e573 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Vidur: A large-scale simulation framework for llm inference
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc99142-7a9c-4d2d-8acc-efddddf537ab · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Taming throughput-latency tradeoff in llm in- ference with sarathi-serve, in: 18th USENIX Sympo- sium on Operating Systems Design and Implementa- tion (OSDI 24), pp
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c588857d-a1c0-4dbc-bee2-680e866c66f0 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Clean sharegpt dataset
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f21e5e5f-995e-4756-b5e2-0308f70745f3 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Language models are few-shot learners
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67101bfd-3a65-42ee-a873-c8bef006c2de · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2e137e-b616-469a-bce4-2465627afe24 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Punica: Multi-tenant lora serving
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9976b25a-4d08-4f3f-8498-aff7bd17953b · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Llmservingsim: A hw/sw co-simulation infrastructure for llm inference serving at scale, in: 2024 IEEE Inter- national Symposium on Workload Characterization (IISWC), IEEE
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3564582c-e288-4c54-abbf-e373ff3ab183 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Computers and Intractability: A Guide to the Theory of NP- Completeness
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1993d999-a665-403d-9414-96f8f4695b9a · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4615fc-356e-4dd0-aa26-c0a0f3b6827b · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9075a214-5de2-4296-858d-e129efa4f2f7 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Parameter-efficient transfer learning fornlp, in: Internationalconferenceonmachinelearn- ing, PMLR
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd75c6e9-a505-4c97-a0f4-ad899115142f · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Lora: Low-rank adaptation of large language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08683555-6cef-4227-b645-c18032096df5 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Chameleon: Adap- tive Caching and Scheduling for Many-Adapter LLM Inference Environments
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c88794-6735-49cb-836d-c4514b9275ed · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Fast algorithms for bin packing
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5de80dd-d827-4444-b286-57fe4a32d0fd · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Ef- ficient memory management for large language model serving with pagedattention, in: Proceedings of the 29th Symposium on Operating Systems Principles, pp
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12d10cbb-89b8-458b-8b9a-8802fbb884c2 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Numba: A llvm-based python jit compiler, in: Proceedings of the Second Workshop on the LLVM Compiler Infrastruc- ture in HPC, pp
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a2a520-729a-4c87-a011-486123706c54 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37946ea0-ddb1-44b4-a58d-8d8a77ebfa97 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Prefix-tuning: Optimizing continuouspromptsforgeneration
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c53d03-6a39-407a-9f01-f00723626e73 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Few-shot parameter- efficient fine-tuning is better and cheaper than in- context learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8aa8b5f-7ea3-414a-8524-ce4954c1acf8 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec86f556-fabc-4bfc-888f-a8f2483cce29 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving DeepSpeed-MII
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0938c5d-4b1e-49fd-add6-667da2e48b4a · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving TensorRT-LLM
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70a7674a-31bc-4549-94cf-3e7a42334af4 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Scikit-learn: Machine learning in Python
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf9b79c7-9b4b-4f42-a965-abcf0052d886 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a8fd1af-78de-40ca-bcd3-5c1640ec3b85 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Mind the memory gap: Unveiling gpu bottlenecks in large- batch llm inference, in: 2025 IEEE 18th International Conference on Cloud Computing (CLOUD), IEEE
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd1c87a-613d-49c5-98f9-2dd29576e5da · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Towards pareto optimal throughput in small language model serving, in: Proceedings of the 4th Workshop on Machine Learning and Systems, pp
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6265eb50-4119-41ef-b976-97a1edf34905 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d040ed-fc9e-4e9c-8d29-4b62ded7c17b · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving EdgeLoRA: An Efficient Multi- Tenant LLM Serving System on Edge Devices
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56fc4404-708a-42d0-96d8-7f36aae052dd · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c28a3b1-8b30-40c5-893a-e75f15e5a616 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Lst: Ladder side-tuning for parameter and memory efficient trans- fer learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc657ccd-90b0-4be8-815d-7c2b060a9bb7 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 891765e7-c9dd-440d-8513-9a076fdbe722 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Parameter-efficient fine-tuning in large language models: a survey of methodologies
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8f7eb87-51b8-4e3e-ad3e-5bc027b201ca · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Finance lora adapter for llama-3.1-8b instruct
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e5b63a-373b-4efc-ad73-542a674052bd · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving dlora: Dynamically orchestrating requests and adapters for lora llm serving, in: 18th USENIX Symposium on Operating Systems Design and Imple- mentation (OSDI 24), pp
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d48b7ccf-1f00-41ef-b96d-406399b665e8 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Qwen2.5 Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11e7cb8-9054-44ef-8339-f5fcbc1de873 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Sql lora for llama-2-7b
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfefe2b4-30b1-4f6a-bbf2-5108a1e93d95 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Orca: A distributed serving system for transformer-based generative models, in: 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22), pp
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8dc3b4a-6280-417c-9c12-19a398a439de · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Shepherd: Serving dnns in the wild, in: 20th USENIX Symposium on Networked Systems Design and Imple- mentation (NSDI 23), pp
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6436fb1a-26a4-47bd-8d82-aff6012b05d2 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Medical lora for qwen2.5-7b- instruc
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6931a87-de69-4e14-86cd-972d2714ead2 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Unresolved cited work
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c5a457-55d4-4e16-b8be-eed0b6a16d19 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Proceedings of Machine Learning and Sys- tems 6, 296–311
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312734b0-5dc8-41d7-8b9c-7ce034606633 · outbound
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving Poster session, San Diego, CA
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.