Pith. sign in

Paper Citation Record · LEDGER

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2512.16056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.16056 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T21:57:12.142867Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:30:33.693227Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact9
  • verified fuzzy41
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1512afd-1f2e-4827-a3ab-bb35a7d8634a · outbound

This paper cites an unresolved cited work.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-16T21:58:36.486195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:f9631e24e61ccf4ece9f502e96127245f158c4e439c6c0a14217c819c0f63b31

Observation 48d2025b-c6e0-404a-ae87-50c2b4b82577 · outbound

This paper cites Advanced Micro Devices.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Advanced Micro Devices

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.488386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:931a22ccf94cb10127a5b4f35767a471fae0f255da502fbea69acf66030ea90c

Observation 1d7481eb-1776-4adb-bc86-40358fbfdbea · outbound

This paper cites Advanced Micro Devices.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Advanced Micro Devices

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.493061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:8a50181755f7529ce028800835c625d1fa8b244171b4cdf41ed397e47b225f7e

Observation e5f7dd5e-13f7-4efa-9215-0114396729dc · outbound

This paper cites Deepspeed-inference: enabling efficient in- ference of transformer models at unprecedented scale.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Deepspeed-inference: enabling efficient in- ference of transformer models at unprecedented scale

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.490728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:1bdce3ddb5d75bf8ce5cf6c224717c087443265c84f2ba3d8b06e0da38c7cc72

Observation c2cf82dc-bc87-4c6e-a579-2554c3aee017 · outbound

This paper cites Qwen Technical Report.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Qwen Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:58:35.765384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T16:38:14.101106+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:ddc9dea0f198b3fbbe92d87d7ff5e78fe5fbea08890d9be0e0f31aa108aecb37

Observation 7f3bb36f-358e-44ec-94d2-3d7d9f0e9c42 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:58:35.743753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:179b96d2e52c4ac6f62540b4723f1c43ce0fa5341439c3b0bed5c5f5bdd625fc

Observation 310ba4a2-ed39-4f40-b360-86edd3eb2c6e · outbound

This paper cites {PipeSwitch}: Fast pipelined context switching for deep learning applications.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services {PipeSwitch}: Fast pipelined context switching for deep learning applications

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.529197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:be9810614066a22d5bbb68470594a7e8941c7ec0f51182a77080dde1b791c39b

Observation eb15c589-e7a2-4955-9b6d-3439f5f42411 · outbound

This paper cites Overlapping data transfers with computation on gpu with tiles.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Overlapping data transfers with computation on gpu with tiles

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.531268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:9f996354cfb705796ccab9b71428d626e79f4dfd039ea1439f19a354c6d32efe

Observation ddb17191-178f-42d2-b8a3-9868ae585bef · outbound

This paper cites NVIDIA H20 GPU Specifications.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services NVIDIA H20 GPU Specifications

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.535412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:fdbf788dc1c31a0b8ba281d0c64769e87f5cfe80282d82d5a880791200515acf

Observation e07bb5a2-e984-4ace-ac87-c84fa74f5b1e · outbound

This paper cites MP-RDMA: enabling RDMA with multi- path transport in datacenters.IEEE/ACM Trans.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services MP-RDMA: enabling RDMA with multi- path transport in datacenters.IEEE/ACM Trans

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.539868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:e93cdb233e0154884dae86a9cdc538a2d7125dbc70bb78d23ab8b1c6da057d7f

Observation 441cb060-0925-4e7f-87fa-a17e4508c910 · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Lmcache: An efficient kv cache layer for enterprise-scale llm inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:58:35.747666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:eb0554f299c1deb6f906366f5f5c5eaac2eeae9a6cee62d4aa30912cc72e62c8

Observation 371f2c2c-782a-4ef7-9724-bc68ec9dde9b · outbound

This paper cites Liminal: Exploring the frontiers of llm decode performance.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Liminal: Exploring the frontiers of llm decode performance

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:58:35.736890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:93dc83e6ec111038c62a82cc33ac02cac2f73ada973ce2aae4d1ecf3d93f9ce0

Observation cceb7542-ddcf-4471-bc72-697c3d9c0b79 · outbound

This paper cites Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:58:35.762214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:6f6a026b43328a8f459c917907c61725f81e34158c8ef0a268f38404e5930c94

Observation 8d1506e8-3b68-4bd1-98c5-17b93ee2a7a5 · outbound

This paper cites {Cost-Efficient} large lan- guage model serving for multi-turn conversations with {CachedAttention}.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services {Cost-Efficient} large lan- guage model serving for multi-turn conversations with {CachedAttention}

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.495681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:b0bcdbe57dfe3394ec7228e3614072b109b115d7e957083b9bea9606ed8a5d27

Observation e056e66c-bdf6-4623-9a8e-ecae4a64f8b1 · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference.Pro- ceedings of Machine Learning and Systems, 6:325–338.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Prompt cache: Modular attention reuse for low-latency inference.Pro- ceedings of Machine Learning and Systems, 6:325–338

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.584662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:e0083dc43cc418d498229abd16454c2634194370389f18d70c8d5d33dd25f10c

Observation c6646b8f-7bd8-46ac-80fc-b3f47c3709aa · outbound

This paper cites Accelerate: Train- ing and inference at scale made simple, efficient and adaptable.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Accelerate: Train- ing and inference at scale made simple, efficient and adaptable

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.586818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:879edb9b67587414fc6adaa5ae1f1718be96ac3e6b8b6f4859e8a02cf2d633b3

Observation 026b2508-a03a-44f9-92d9-3f63239a9459 · outbound

This paper cites Elsevier.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Elsevier

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.578185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:f0458a3357ce49fd6f7c7ac24bc8dd7c62bbc92a1c320eac01c0d75e4b01273b

Observation a2db5fe7-fac7-4534-a938-4d2fd0429f50 · outbound

This paper cites In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 87–101.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 87–101

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.582402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:34d2b0ed8251596abd4257e1f41eb5593e2c214661b9dee8e74307fd1e1adeee

Observation 32fd17b9-fcf1-4339-ae4d-b7db99498a3a · outbound

This paper cites Ragcache: Efficient knowledge caching for retrieval-augmented generation.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Ragcache: Efficient knowledge caching for retrieval-augmented generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.569355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:debf367ef4836cab8b0f14ea6b7d0744473aaef99186be6f9de4beed1b674b8c

Observation 70f74dba-9826-46ee-aa2c-81c5d008f5c6 · outbound

This paper cites In- put/output memory management unit with protection mode for preventing memory access by i/o devices, Jan- uary 14 2014.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services In- put/output memory management unit with protection mode for preventing memory access by i/o devices, Jan- uary 14 2014

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.564948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:4db0e8c7a4b04c01aad89f925f53f96cf1e701b2f34ed1564b3480f72484ff77

Observation dafb295d-15b6-44eb-a207-961b9c88ad66 · outbound

This paper cites Efficient memory manage- ment for large language model serving with pagedatten- tion.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Efficient memory manage- ment for large language model serving with pagedatten- tion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.571780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:c78bbaebf5727f1d9bbc83d7ec121e0f96ea32dbcc387abf89c499a811a732c7

Observation 9e23028c-9366-4374-80a2-e0e627c92fd4 · outbound

This paper cites Tuccl: Tailored and unified configuration optimizations for high-performance collective communication library.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Tuccl: Tailored and unified configuration optimizations for high-performance collective communication library

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.562583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:70d27daeb2095dc40bced5e1ef2d464fba50959a106c5f1b2141424bdf900791

Observation a2b88dfe-c9fe-4015-839c-682ca6669fa6 · outbound

This paper cites Reducing gpu offload latency via fine-grained cpu-gpu synchroniza- tion.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Reducing gpu offload latency via fine-grained cpu-gpu synchroniza- tion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.555180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:613e287904f9505650c43f7456b5c191bdb1fa8d2ca7d81e2cad3a8bb28084ef

Observation 79106038-1976-4d4b-8dc4-aeeff2c198aa · outbound

This paper cites Mlp-offload: Multi- level, multi-path offloading for llm pre-training to break the gpu memory wall.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Mlp-offload: Multi- level, multi-path offloading for llm pre-training to break the gpu memory wall

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.557553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:54b4d9b1abd988e548dda5e7f6cfc8b7fbf37b57f0f1bdd439807e3548bd65b0

Observation c99002f3-e5c4-4b3a-8ed7-5176e83740b5 · outbound

This paper cites AzurePublicDataset: Azure LLM In- ference Dataset 2023.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services AzurePublicDataset: Azure LLM In- ference Dataset 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.559940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:da590facdeaefbba75503b43793de1134ad08678212ab7f1833ec77b375c2da9

Observation cce9ab10-0984-43d3-b0a8-5f1031e7d7b7 · outbound

This paper cites Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for?Queue, 6(2):40–53.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Scalable parallel programming with cuda: Is cuda the parallel programming model that application developers have been waiting for?Queue, 6(2):40–53

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.567284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:c48fb4f9c2738153b835af4f862ae8de1f4e5626419ad662e6f4bf9652c6a01b

Observation 41890653-6936-48b1-9be6-f801bb00176f · outbound

This paper cites NVLink and NVSwitch.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services NVLink and NVSwitch

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.580068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:2bba3078475942858fd80877b55d7b98af27ec6d2bc7d2423f8ac4885f99c945

Observation 9a3dfa21-96ed-430f-9a82-8c784f89f533 · outbound

This paper cites NVIDIA NVLink 4.0 Technology.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services NVIDIA NVLink 4.0 Technology

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.550031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:730573b1c4294d87f3df368d5e29c14ad5c209774edc4f4ee8183e47c271dd64

Observation 162c1464-a930-465d-b04d-edc251fde8d1 · outbound

This paper cites Chatbot Arena Conversations Dataset.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Chatbot Arena Conversations Dataset

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.547646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:44976d014cab0c506c636599e9b4395d41867da6d2955377df036924122504bb

Observation 843df7a1-6ebe-4d5e-8af9-c9770fa13f3c · outbound

This paper cites Marconi: Prefix Caching for the Era of Hybrid LLMs.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Marconi: Prefix Caching for the Era of Hybrid LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:58:35.758757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:671d4464c93454d37b6a12b83dd579649e64acc89da26eb64852e5b25d0b2576

Observation 6b2507e1-9e05-41dd-854a-cd7eaa912f0a · outbound

This paper cites Pci express® base specification revision 5.0 version 1.0.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Pci express® base specification revision 5.0 version 1.0

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.533391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:e83e37a12c95e3f09fd44324c85845e8200e2b05d718bd049bd5a3499e24b65c

Observation 4bc0f6bd-9623-4d58-a1c9-3122aabcdcf4 · outbound

This paper cites Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.537501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:5b18efc73acabeb6f3aefc9a5347ed2d73181993d4f4fea6044850da56d86ac0

Observation 8266e2cf-d829-4b34-a41e-2a95666e08ed · outbound

This paper cites An i/o characterizing study of offloading llm models and kv caches to nvme ssd.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services An i/o characterizing study of offloading llm models and kv caches to nvme ssd

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.552527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:51b5521d603f7d55ae68e3e3554b9f27911e4988d369d236b22a726250667d66

Observation 2a81e73d-0f78-4d56-99f9-bbb1da58781d · outbound

This paper cites Enabling efficient GPU communication over mul- tiple NICs with fuselink.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Enabling efficient GPU communication over mul- tiple NICs with fuselink

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.574062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:93b26423a9b9884b3bfeef662875e1d2654578026faf9ffab73930549fe60e13

Observation a5967116-cafb-43bb-bcb4-87eacc2ed466 · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:58:35.755161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:855b4bdf828a73b06defc6d6adc9cc726cd6d208b0e2a833c9bdcff401db903e

Observation 9b6bc2ed-f353-49a8-8f2a-f66c4094a8f1 · outbound

This paper cites Flexgen: high-throughput generative inference of large language models with a single gpu.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Flexgen: high-throughput generative inference of large language models with a single gpu

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.521390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:c85b8f9466435198669a95f560bf0f3e906e11ea74bae42f00bc5a85cfca1c8d

Observation a7960dbc-5723-4890-9be3-85f620d00910 · outbound

This paper cites Efficient intra-node hierarchical par- allelisms and dynamic load balancing strategies on het- erogeneous systems.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Efficient intra-node hierarchical par- allelisms and dynamic load balancing strategies on het- erogeneous systems

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.515852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:af7b8c2770fe9b8235d498b199f9c35de81551cc54bea6a7dfd57ea0103d9633

Observation b2c301ad-db6a-43f1-a76b-d1f37faf58f9 · outbound

This paper cites Accelerating intra-node gpu communication: A perfor- mance model for multi-path transfers.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Accelerating intra-node gpu communication: A perfor- mance model for multi-path transfers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.524390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:448acbe573ed278d54bfaab2d2c14d6b2f2d9648b21fc0de2aac60f0544b914d

Observation 496af748-083d-4023-9184-ad22689cafbc · outbound

This paper cites Collabora- tive bandwidth-efficient intra-node allreduce.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Collabora- tive bandwidth-efficient intra-node allreduce

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.510549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:1cc06f33747ef0d0d4f4bca11b3c6abc2afd5bdeda18acef6f7c7d8988c87e68

Observation 40fb733f-c921-4685-8dcc-06c691006112 · outbound

This paper cites Enhancing intra-node GPU-to-GPU perfor- mance in MPI+UCX through multi-path communication.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Enhancing intra-node GPU-to-GPU perfor- mance in MPI+UCX through multi-path communication

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.507739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:20176ceec64c4952964027eb3f9f9575a80c7a7f5e528b42a1ffa4874442bf01

Observation 4615ee21-cc2f-4e5f-8846-f911100d65f1 · outbound

This paper cites Sur- vey of intra-node gpu interconnection in scale-up net- work: Challenges, status, insights, and future directions.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Sur- vey of intra-node gpu interconnection in scale-up net- work: Challenges, status, insights, and future directions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.501381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:79fee1b5737140f92ba98894cb091a3171e70fb3d5d43c07c8036a11a5611005

Observation 7d2bb17b-3976-4b20-bba6-05076a2c7257 · outbound

This paper cites Engine-agnostic model hot-swapping for cost-effective llm inference.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Engine-agnostic model hot-swapping for cost-effective llm inference

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.504565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:9c9210fb9d6935bcead13e6c77b7e52648a52776408efef8159831ce1dbd53b8

Observation da24bbc2-1653-4d9b-a3bf-4962e6087438 · outbound

This paper cites Scalable and efficient intra-and inter-node in- terconnection networks for post-exascale supercomput- ers and data centers.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Scalable and efficient intra-and inter-node in- terconnection networks for post-exascale supercomput- ers and data centers

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:58:35.740286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:7f0285eefbad37ac462c96273267538769cc42a37938f45657928038f5bd8b3a

Observation 623f91de-72ee-4c44-98d6-e34d6b100ede · outbound

This paper cites Under- standing intra-node communication in hpc systems and datacenters.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Under- standing intra-node communication in hpc systems and datacenters

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:58:35.733063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:a82d508cd1817c4d614c8a6b9b7917088e8151c8b6e6fe73a8f3485985e61012

Observation 1f8a2a01-f2bf-4fbc-bed3-14fa25c98d45 · outbound

This paper cites Efficient multi-path NVLink/PCIe-aware UCX-based collective communi- cation for deep learning.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Efficient multi-path NVLink/PCIe-aware UCX-based collective communi- cation for deep learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.513247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:d7498c0e4d656c2fb628fafe75042ada855fc021420adc3767812bd74b185b92

Observation 8102c47c-503d-49b4-945d-10e1e3a2f206 · outbound

This paper cites Performance models for cpu-gpu data transfers.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Performance models for cpu-gpu data transfers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.518719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:bc23f8def7328616cfc34e38224ca20bd8edf9bf678249f9a0870ea4b989ad38

Observation 14b70501-5ab6-42fa-86ee-ac22a3efbf95 · outbound

This paper cites Sleep Mode.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Sleep Mode

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.526863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:6421743bc2af3e1e7a19bdb094d6dadb8979ae7f8b72738b0076a65bc7dd4bdd

Observation cdaa05ba-84ab-49e9-be66-8a0e38219325 · outbound

This paper cites Design, implementation and evalua- tion of congestion control for multipath {TCP}.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Design, implementation and evalua- tion of congestion control for multipath {TCP}

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.576292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:10a06767ca253f38e5ac7360da6348dfd481e25d0db686d6296ff6ec991ec028

Observation 7e193dbb-1e27-4f50-bbaf-379040edb276 · outbound

This paper cites Qwen3 Technical Report.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Qwen3 Technical Report

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:58:35.751577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:c94198a6f77f7ec704d5d374936770f37ae730f51a918cbed20e4523dd2f30f1

Observation abca3b2d-459c-4ccb-8696-d3852c16c6df · outbound

This paper cites Learned prefix caching for efficient llm inference.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Learned prefix caching for efficient llm inference

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.542420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:e11a747f695e12a84c24f3a59e256a4ead8f2c141be8383bd11a33e3a695ad0a

Observation 272a3b40-91b7-49b6-b0ad-88847301d1ef · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information pro- cessing systems, 37:62557–62583.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services Sglang: Efficient execution of structured language model programs.Advances in neural information pro- cessing systems, 37:62557–62583

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.544891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:2245f32d2540f8b29398110dc9b461d49b72754946da9fa7c09341def9a1b46f

Observation 77adca2e-93ff-407f-97cf-6c52547b147b · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving.

MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:58:36.498710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T21:57:12.142867Z digest=sha256:dfe23ade0e266113a43a315138767e82f38fec1d514aaddb6623d8df82c9913f

Pith citing papers

Observation 718b15f1-63ac-4ccd-af80-703bb99b4d5b · inbound

HybridQC: Hardware-Grounded Simulation of Tightly Integrated Hybrid Quantum-Classical Systems cites this paper.

HybridQC: Hardware-Grounded Simulation of Tightly Integrated Hybrid Quantum-Classical Systems MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:33.693227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:30:33.693227Z digest=sha256:63e51c59648af7ed87663c56f028065097b0208131c250ca811f0911995eca72