Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T21:57:12.142867Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2512.16056.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T21:57:12.142867Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T05:30:33.693227Z
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a1512afd-1f2e-4827-a3ab-bb35a7d8634a · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48d2025b-c6e0-404a-ae87-50c2b4b82577 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Advanced Micro Devices
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d7481eb-1776-4adb-bc86-40358fbfdbea · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Advanced Micro Devices
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e5f7dd5e-13f7-4efa-9215-0114396729dc · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Deepspeed-inference: enabling efficient in- ference of transformer models at unprecedented scale
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2cf82dc-bc87-4c6e-a579-2554c3aee017 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Qwen Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T16:38:14.101106+00:00.
Observation 7f3bb36f-358e-44ec-94d2-3d7d9f0e9c42 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 310ba4a2-ed39-4f40-b360-86edd3eb2c6e · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services {PipeSwitch}: Fast pipelined context switching for deep learning applications
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eb15c589-e7a2-4955-9b6d-3439f5f42411 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Overlapping data transfers with computation on gpu with tiles
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ddb17191-178f-42d2-b8a3-9868ae585bef · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services NVIDIA H20 GPU Specifications
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e07bb5a2-e984-4ace-ac87-c84fa74f5b1e · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services MP-RDMA: enabling RDMA with multi- path transport in datacenters.IEEE/ACM Trans
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 441cb060-0925-4e7f-87fa-a17e4508c910 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Lmcache: An efficient kv cache layer for enterprise-scale llm inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 371f2c2c-782a-4ef7-9724-bc68ec9dde9b · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Liminal: Exploring the frontiers of llm decode performance
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cceb7542-ddcf-4471-bc72-697c3d9c0b79 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d1506e8-3b68-4bd1-98c5-17b93ee2a7a5 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services {Cost-Efficient} large lan- guage model serving for multi-turn conversations with {CachedAttention}
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e056e66c-bdf6-4623-9a8e-ecae4a64f8b1 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Prompt cache: Modular attention reuse for low-latency inference.Pro- ceedings of Machine Learning and Systems, 6:325–338
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c6646b8f-7bd8-46ac-80fc-b3f47c3709aa · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Accelerate: Train- ing and inference at scale made simple, efficient and adaptable
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 026b2508-a03a-44f9-92d9-3f63239a9459 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Elsevier
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a2db5fe7-fac7-4534-a938-4d2fd0429f50 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 87–101
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 32fd17b9-fcf1-4339-ae4d-b7db99498a3a · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Ragcache: Efficient knowledge caching for retrieval-augmented generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 70f74dba-9826-46ee-aa2c-81c5d008f5c6 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services In- put/output memory management unit with protection mode for preventing memory access by i/o devices, Jan- uary 14 2014
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dafb295d-15b6-44eb-a207-961b9c88ad66 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Efficient memory manage- ment for large language model serving with pagedatten- tion
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9e23028c-9366-4374-80a2-e0e627c92fd4 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Tuccl: Tailored and unified configuration optimizations for high-performance collective communication library
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a2b88dfe-c9fe-4015-839c-682ca6669fa6 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Reducing gpu offload latency via fine-grained cpu-gpu synchroniza- tion
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 79106038-1976-4d4b-8dc4-aeeff2c198aa · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Mlp-offload: Multi- level, multi-path offloading for llm pre-training to break the gpu memory wall
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c99002f3-e5c4-4b3a-8ed7-5176e83740b5 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services AzurePublicDataset: Azure LLM In- ference Dataset 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cce9ab10-0984-43d3-b0a8-5f1031e7d7b7 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for?Queue, 6(2):40–53
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 41890653-6936-48b1-9be6-f801bb00176f · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services NVLink and NVSwitch
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9a3dfa21-96ed-430f-9a82-8c784f89f533 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services NVIDIA NVLink 4.0 Technology
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 162c1464-a930-465d-b04d-edc251fde8d1 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Chatbot Arena Conversations Dataset
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 843df7a1-6ebe-4d5e-8af9-c9770fa13f3c · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Marconi: Prefix Caching for the Era of Hybrid LLMs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6b2507e1-9e05-41dd-854a-cd7eaa912f0a · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Pci express® base specification revision 5.0 version 1.0
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4bc0f6bd-9623-4d58-a1c9-3122aabcdcf4 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8266e2cf-d829-4b34-a41e-2a95666e08ed · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services An i/o characterizing study of offloading llm models and kv caches to nvme ssd
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2a81e73d-0f78-4d56-99f9-bbb1da58781d · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Enabling efficient GPU communication over mul- tiple NICs with fuselink
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5967116-cafb-43bb-bcb4-87eacc2ed466 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services SparQ Attention: Bandwidth-Efficient LLM Inference
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b6bc2ed-f353-49a8-8f2a-f66c4094a8f1 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Flexgen: high-throughput generative inference of large language models with a single gpu
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a7960dbc-5723-4890-9be3-85f620d00910 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Efficient intra-node hierarchical par- allelisms and dynamic load balancing strategies on het- erogeneous systems
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b2c301ad-db6a-43f1-a76b-d1f37faf58f9 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Accelerating intra-node gpu communication: A perfor- mance model for multi-path transfers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 496af748-083d-4023-9184-ad22689cafbc · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Collabora- tive bandwidth-efficient intra-node allreduce
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40fb733f-c921-4685-8dcc-06c691006112 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Enhancing intra-node GPU-to-GPU perfor- mance in MPI+UCX through multi-path communication
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4615ee21-cc2f-4e5f-8846-f911100d65f1 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Sur- vey of intra-node gpu interconnection in scale-up net- work: Challenges, status, insights, and future directions
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7d2bb17b-3976-4b20-bba6-05076a2c7257 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Engine-agnostic model hot-swapping for cost-effective llm inference
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation da24bbc2-1653-4d9b-a3bf-4962e6087438 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Scalable and efficient intra-and inter-node in- terconnection networks for post-exascale supercomput- ers and data centers
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 623f91de-72ee-4c44-98d6-e34d6b100ede · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Under- standing intra-node communication in hpc systems and datacenters
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1f8a2a01-f2bf-4fbc-bed3-14fa25c98d45 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Efficient multi-path NVLink/PCIe-aware UCX-based collective communi- cation for deep learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8102c47c-503d-49b4-945d-10e1e3a2f206 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Performance models for cpu-gpu data transfers
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14b70501-5ab6-42fa-86ee-ac22a3efbf95 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Sleep Mode
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cdaa05ba-84ab-49e9-be66-8a0e38219325 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Design, implementation and evalua- tion of congestion control for multipath {TCP}
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7e193dbb-1e27-4f50-bbaf-379040edb276 · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Qwen3 Technical Report
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation abca3b2d-459c-4ccb-8696-d3852c16c6df · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Learned prefix caching for efficient llm inference
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 272a3b40-91b7-49b6-b0ad-88847301d1ef · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Sglang: Efficient execution of structured language model programs.Advances in neural information pro- cessing systems, 37:62557–62583
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 77adca2e-93ff-407f-97cf-6c52547b147b · outbound
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 718b15f1-63ac-4ccd-af80-703bb99b4d5b · inbound
HybridQC: Hardware-Grounded Simulation of Tightly Integrated Hybrid Quantum-Classical Systems MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.