Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T06:27:23.580445Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 151 outbound references and 2 inbound Pith citation observations for arXiv:2604.17227.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T06:27:23.580445Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T21:14:49.980623Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 151 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c48d021b-c541-4a70-97a1-acbdd17d45ab · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb4b2a59-1731-492c-8188-b74ba53c7337 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Cloud container technologies: a state-of-the-art review.IEEE Transactions on Cloud Computing.2017;7(3):677–692
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87d57ad5-9bb0-4aa3-b1f4-e7e99f4f2c17 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Attention is all you need
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfbd3ce2-6896-4075-b0a1-30ca5b82bf50 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Scaling Laws for Neural Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 921ccf4f-c350-484b-8d0b-9fc327f4461e · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Deep learning.Nature.2015;521(7553):436–444
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42cc0610-3505-46a8-94f3-434d6cea221d · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4c780b8-a6aa-4ea1-9cde-96d2911ed3d5 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.; 2020
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4414f9a-de00-4c4a-8ab8-d5105e610d3c · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Curran Associates, Inc
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation edd5c27a-92dd-4792-8e0c-2ec5752cc93f · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d1ee9dd-94c6-4734-a7a8-0f3a543aff82 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation abbd93f5-bdde-4ef1-86ff-452886a8afc1 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Towards end-to-end optimization of llm-based applications with ayo
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 191292db-c5f5-46bb-a001-ca180fef44a1 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In-datacenter performance analysis of a tensor processing unit
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 71e6585a-b4c3-484f-8bc6-464319c0caf5 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76f87055-c3fd-41ea-8f8c-b79c8a5a175f · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Efficient memory management for large language model serving with pagedattention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00b1b090-037b-4b7b-a68e-63993c042d0b · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Open issues in scheduling microservices in the cloud.IEEE Cloud Computing.2016;3(5):81–88
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6832a14-856f-4a83-8538-6dd4db7e69ee · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04b02cac-0fc8-495f-88f3-a0b7207abf99 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Flexgen: High-throughput generative inference of large language models with a single gpu
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31c66a56-e34f-4b51-9dca-7a324c2a4559 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 16th USENIX symposium on operating systems design and implementation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7592b4d9-ee42-4e58-99f8-b94283ad7729 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Evaluation and benchmarking of llm agents: A survey
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f62d3047-60a2-4612-a197-c53594eb37af · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0dcb709-9ef9-4b4f-8ceb-a3b11b8cbe72 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7de80109-f300-4869-8291-f4851a0591e6 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Process modeling in web applications.ACM Transactions on Software Engineering and Methodology.2006;15(4):360–409
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 674efc27-9095-43ba-8783-20e44dd606c0 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Challenges in deployment and configuration management in cyber physical system
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e8760e8-7ba6-4e9a-989d-1cdc00472b7d · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Parallel processing systems for big data: a survey
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87b97a7f-290d-432b-9472-6b79a89b36a8 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3266407b-5c1d-4cd9-a523-4b9831d03494 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Deep learning workload scheduling in gpu datacenters: A survey.ACM Computing Surveys
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5b06068-db71-42aa-a74e-63855aa317e8 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications.ACM Trans
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aee1276b-26da-4cb1-9be2-40515646f7fe · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda LLM Inference Scheduling: A Survey of Techniques, Frameworks, and Trade-offs.Authorea Preprints.2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ed6541f-855a-4c8a-ad05-0de53dbfd62c · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Cloud Native System for LLM Inference Serving
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6219f06-6e3a-45c4-bb18-ba43623f8455 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda {NanoFlow}:Towardsoptimallargelanguagemodelservingthroughput.In:Proceedingsof the 19th USENIX Symposium on Operating Systems Design and Implementation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a1a0b92-22c1-45d6-9c25-9a4134562f43 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Flashinfer: Efficient and customizable attention engine for llm inference serving.Proceedings of Machine Learning and Systems.2025;7
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b70ee6af-dddb-4229-9208-ca4bc3086d6c · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2025:446–461
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a32948a6-e47b-4542-8725-c4ae40f46be8 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda throttll’em: Predictive gpu throttling for energy effi- cient llm inference serving
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 136e5946-975c-43d6-948e-22a5ef67db2d · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f61e7be-90b1-4fc5-a2e0-3b76f985d0dc · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Extracting training data from large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 52b95004-ad1a-4f20-b241-b1c70af874c5 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Quantifying memorization across neural language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d4861b5-39cc-417a-9dce-1fd76660fbb7 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Deep learning with differential privacy
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79988a33-79ae-4a19-802b-1b733f770298 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Communication-efficient learning of deep networks from decentralized data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9fb2a544-f8fc-4640-8e94-99a86ad8a14a · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Oblivious {Multi-Party} machine learning on trusted processors
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84cbbcb9-0f0e-4a32-858f-d4609f09995c · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7f52107-b027-4e1e-90c6-7726ee0fbdc8 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Stealing machine learning models via prediction APIs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3897ff2d-81d3-4c5b-a502-d1e187ecc349 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 867c3056-3558-43b0-9978-e4746d394dbf · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda ACMComputingSurveys
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 483a2e16-488a-47b8-a178-8cebbb345a8d · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda IEEEInternet of Things Journal.2022;9(11):8364–8386
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d42d3bd2-9270-46c3-acf1-06fc45c39d89 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Spatial big data architecture: from data warehouses and data lakes to the Lakehouse.Journal of Parallel and Distributed Computing.2023;176:70–79
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation feaab007-1c53-4b10-aeed-098ea0d3f8c6 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda IEEETransactionsonKnowledgeand Data Engineering.2023;35(12):12571–12590
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 90ed30bf-8d6e-483d-9703-f031ef266de1 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 56a3e7f3-5af7-45b2-890e-e365d97dc068 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Computational cxl-memory solution for accelerating memory-intensive applications.IEEE Computer Architecture Letters.2022;22(1):5–8
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39e63fcf-045e-4efd-83d7-d7d1c1292230 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a28dbc49-7171-4db8-b8d2-9f37d0ef3ead · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 139a995e-2eeb-4bd5-859c-619552c03f31 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda New Trends in High-D Vector Similarity Search: AI-driven, Progressive, and Distributed
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f894ca1-b279-46c6-a7cd-50a60b4ee277 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unleash llms potential for sequential recommendation by coordinating dual dynamic index mechanism
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5a4194cc-fc9f-4b09-be6f-094e7a7c59d4 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1ddf85f7-28e1-48f3-baad-46b027d7885f · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda A Generative Caching System for Large Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e54daba5-3418-4a40-b476-013c1f153ae3 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Tracing the Data Trail: A Survey of Data Provenance, Transparency and Traceability in LLMs
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95e1ea6d-628d-4f72-81cb-e5e5314f365b · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 41st IEEE/ACM International Conference on Computer-Aided Design
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5a6e6b58-1395-41da-809e-8aa2ee6b3b5a · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2023:189–201
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cda81a75-03f0-4941-9fff-fc5a3bcba3a8 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Packt Publishing Ltd
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5acdf7b7-5833-4142-93cd-19bb6cea4ad3 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Llm-pilot: Characterize and optimize performance of your llm inference services
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c10931d-66ce-424e-9566-8d6835130bad · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda How hungry is ai? benchmarking energy, water, and carbon footprint of llm inference
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a560c185-448e-4cc3-bbb9-8a339d009fe9 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda A survey on federated fine-tuning of large language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9bde166-b12b-4be7-b4bb-c6e04bb594d3 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda IEEE Transactions on Computers.2026;75(4):1636-1649
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 470a91df-ce65-4816-aec1-95220eeb9386 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2026 , issn =
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4cdca6a3-738d-4eb3-b33a-e108ece95454 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda ShuffleInfer: Disaggregate LLM inference for mixed downstream workloads.ACM Transactions on Architecture and Code Optimization.2025;22(2):1–24
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b8d8b10-2c56-495d-b1c0-603b5941ac09 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Splitwise: Efficient generative LLM inference using phase splitting
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c02a5c08-c33e-4da5-b4cf-530662a421fc · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Fast Distributed Inference Serving for Large Language Models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf870d12-d482-4314-8454-4d57614fcc5a · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e2d305a-eb41-4beb-b51b-bc43a6fd9001 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity.arXiv preprint arXiv:2512.03416.2025
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ebbb320-0c7b-4aa0-9864-c55af95745a8 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c720b1a-abd1-4dfe-bb2a-3d0af8021e15 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07cb87e2-5ea5-4db6-9d49-76b258da44c6 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda MegaScale-Infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6f1961f-65db-4611-baea-e25520c1d722 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Punica: Multi-tenant LoRA serving.Proceedings of Machine Learning and Systems.2024;6:1–13
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 252ce8d0-2450-4d20-ab5a-a1d0bc76e187 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda AlpaServe: Statistical multiplexing with model parallelism for deep learning serving
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89e93491-6fd0-4de9-9676-507bdd68670d · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Niyama : Breaking the Silos of LLM Inference Serving
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b768dfc0-cc88-4817-ac1e-a89250b66ff3 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Dis- aggregated LLM Serving in AI Infrastructure
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d9e69d42-00ab-4fce-92f1-1ebd1daf33cd · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda DynamoLLM: Designing LLM Inference Clusters for Perfor- manceandEnergyEfficiency.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b769b3bd-5ab1-4d73-b565-4ea4b0e547f8 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2024:207–222
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 34368300-537d-4d58-9c1d-09b685e1df89 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e021acf-ec94-43f0-ad98-bddc56040aaa · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48d79405-9708-478b-8961-ebd6a4bdf4a2 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Fairness in serving large language models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4db4c170-dc7f-4925-9d9c-198787d691ab · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters.IEEE Transactions on Services Computing.2024;17(6):3473-3484
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f59ea1a-7706-4412-8db4-9cf016b23707 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the 2023 USENIX Annual Technical Conference
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0855d5e9-181c-4cf3-b976-b5aa139b186d · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2024:929–945
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6b3a9a2-364a-4b87-a085-8f8509fef4a1 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 80a2d610-850d-461a-b19e-f84086bd8d0a · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa266403-5627-4fff-a767-dca513225439 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda USENIX Association 2024; Santa Clara, CA:135–153
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1edc71a-d655-4c6e-a2bd-beff4a77fe74 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda TAPAS: Thermal-and power-aware scheduling for LLM inference in cloud platforms
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26b0ea8f-3cbd-4e8c-a4c6-80e7b21ac568 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda throttLL’eM: Predictive GPU Throttling for Energy Effi- cientLLMInferenceServing.In:Proceedingsofthe2025IEEEInternationalSymposiumonHighPerformanceComputer Architecture
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b59494f5-6e33-4c66-a087-a3e4edb67882 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Medusa: Accelerating Serverless LLM Inference with Materialization
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54a4aea7-0576-4baa-b0db-10a051ecd40c · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04ae0b47-ae62-40d0-b1e8-1a3e30cfa5e5 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda 2025:415–430
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72bfd0a9-e6db-497c-a0ff-9edcd31fa148 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c6c0afc-6793-4665-89fe-b29eadc0bca2 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda FPGA-Based Sparse Matrix Multiplication Accelerators: From State-of-the-Art to Future Opportunities
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2a152618-5713-4a2f-8ee9-7e371ce6e873 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda A survey on hardware accelerators for large language models.Applied Sciences.2025;15(2):586
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f49a0d9-8768-41c1-a0af-0ce3f0c5d290 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unleashing the Potential of LLMs for Quantum Computing: A Study in Quantum Architecture Design
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b92a5b2-d53e-4cbb-90aa-981378b81a51 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Efficient Large Language Models: A Survey
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1402783-282e-4a23-b33e-67340aeed24a · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5152f1ce-90b0-4831-8771-e6593bd6f7a3 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e806a45e-cdc7-457f-a409-f22f2c480eb0 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda LLM-based cost-aware task scheduling for cloud computing systems.Journal of Cloud Computing
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 32bdcac2-74ec-4c92-80f4-a718d4355b33 · outbound
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda Green scheduling for LLM workloads with model and data reuse across geo-distributed data centers.Digital Communications and Networks.2025
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4804a744-832c-4fc8-acef-932896006489 · inbound
BrownoutMoE: Structure-Aware Expert Grouping for Efficient and Accurate LLM Web-based Services Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab4d15bd-1b8b-4b4a-8b7c-eedd876d3522 · inbound
CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.