Pith. sign in

Paper Citation Record · LEDGER

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs

As of 5 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2607.23115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23115 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:40:12.723229Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved83
  • parse uncertain2
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e81fbd40-63ff-4fc2-bd55-05114131724b · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 1

Resolution
parse uncertain
no resolver link, observed 2026-08-01T03:40:10.999156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:10.999156Z digest=sha256:52b2374aa6b3a7cb7ccea447be1609f11420002d8bd3278b02cd68b62179e4df

Observation 4da32b21-7b4a-4b7e-b44f-f20dc4ec28b5 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.005260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.005260Z digest=sha256:2caa203a0893f4304aca70df4d50cb0a14d2a9abbce2c8cd38d661738a284aca

Observation a5039eec-7bfa-4d8e-85b5-96275f47ad23 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.010282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.010282Z digest=sha256:4d6a67b87c8b46a1a26c4a8a114b5a8aae15cc6928b3ee39813d61d19b4d64da

Observation ee08d559-9968-4f12-9b7c-842efcc8f3e2 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.016717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.016717Z digest=sha256:184b0603ad61c2d241ad25e930e8ecc5052def0d54212a30d990489d912e6a80

Observation 75dcedb9-39d9-49cb-8803-3983dd7e0a82 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.021558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.021558Z digest=sha256:ac68aa0e3cdf18004a36c6907c38181323e60414738594917e450cbb38e9fc5c

Observation acf32578-7480-4dd9-8ee9-a728a93cabdc · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.026430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.026430Z digest=sha256:a80e5eec01270701a3d3ae431575b1867a14ed7f9a28844106f30cf0e3a11585

Observation 398ec4d2-4f66-46a8-8a2b-aebc9527be8d · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 7

Resolution
parse uncertain
no resolver link, observed 2026-08-01T03:40:11.031727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.031727Z digest=sha256:4aa36e96705debebc747e3c3c86eb8ffbfd69f0b9b4fae96cf3a2a5168e63938

Observation d6d493f9-cf75-428d-9f6b-79d4780f6d6c · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.035939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.035939Z digest=sha256:f53e359bacb0f87fed412ec31456532e6419bc9a3c9ae28713e81571f8dc1d17

Observation f91212a2-8ee2-4d32-b5d3-de5ef048aaf6 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.040074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.040074Z digest=sha256:5326d97afd724a0ecae0c3c60749d5abc75704e7f6e98eccc67c1de297c1969a

Observation faf139c7-f5ae-4003-a1d8-05b1a98d601e · outbound

This paper cites Scissionlite: Accelerating distributed deep learning with lightweight data compression for iiot.IEEE Transac- tions on Industrial Informatics, 20(10):11950–11960, 2024.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Scissionlite: Accelerating distributed deep learning with lightweight data compression for iiot.IEEE Transac- tions on Industrial Informatics, 20(10):11950–11960, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.044193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.044193Z digest=sha256:2b469c705cf8c8f527a44b95751123fac6a21c8961a95b1aa55ad5648d9d2f5a

Observation cfd4fa3d-553e-4145-a395-0831c5726fd3 · outbound

This paper cites Tooth: Toward optimal balance of video {QoE} and redundancy cost by {Fine- Grained}{FEC} in cloud gaming streaming.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Tooth: Toward optimal balance of video {QoE} and redundancy cost by {Fine- Grained}{FEC} in cloud gaming streaming

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.048531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.048531Z digest=sha256:d763e7833c0269a1c49877a8d57cbfddfb8edaef2231c566e1f8d5b5e865f39a

Observation a98f0dba-9b93-43cf-930d-3a588ffe6c00 · outbound

This paper cites Crux: Gpu-efficient communication scheduling for deep learning training.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Crux: Gpu-efficient communication scheduling for deep learning training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.053431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.053431Z digest=sha256:cade0face3f8a6832223df15c4599bec9dd2d858f967228cb689feac3c087571

Observation 7aba9b77-b011-46a8-995d-5a2d2eb23de6 · outbound

This paper cites Eva: Cost- efficient cloud-based cluster scheduling.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Eva: Cost- efficient cloud-based cluster scheduling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.057720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.057720Z digest=sha256:04bf82a1b7d78af01f7504495eaff43d4ca4acca601040d1b274586879ae17c9

Observation 5af2baa5-b736-4b88-9b16-8515905d1d06 · outbound

This paper cites Kernel oper- ations on the gpu, with autodiff, without memory over- flows.Journal of Machine Learning Research, 22(74):1– 6, 2021.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kernel oper- ations on the gpu, with autodiff, without memory over- flows.Journal of Machine Learning Research, 22(74):1– 6, 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.061940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.061940Z digest=sha256:95492c84ad8a63f9c151c9d15b800c6561aa034ec9e97df939954c0161be1eae

Observation 7bbd9110-e6f2-46af-85ee-526807ff3686 · outbound

This paper cites Remote procedure call as a managed system service.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Remote procedure call as a managed system service

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.066423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.066423Z digest=sha256:33fad29e6fd5359381698efab0b73195003bd5d0bb717384dfd51f849f3ee8cd

Observation 591726b1-51c1-47dc-b966-fdb425e9b7e1 · outbound

This paper cites Multiplexing dynamic deep learn- ing workloads with slo-awareness in gpu clusters.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Multiplexing dynamic deep learn- ing workloads with slo-awareness in gpu clusters

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.070517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.070517Z digest=sha256:6cca6f97fc1677bccf0c5ef537e685dfe884bf7e542ffb2ebcf37b7d79a47327

Observation 45221d94-c5ea-4cb3-a447-c93037871049 · outbound

This paper cites {GRACE}:{Loss- Resilient}{Real-Time} video through neural codecs.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {GRACE}:{Loss- Resilient}{Real-Time} video through neural codecs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.074609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.074609Z digest=sha256:950ef93b3eef46cc3b03651c1c1e597ba93eb05f1faf116b0a07ada2926f29a9

Observation 7379fdeb-82e1-4252-8bb5-d5414e62298a · outbound

This paper cites Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.078564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.078564Z digest=sha256:0818d901c281248446aae7af6f4ed1ac988b60b113dcb36de26b4745325f8963

Observation e4812f51-687c-4fc5-a585-03122aa6c47f · outbound

This paper cites Oneadapt: Fast adapta- tion for deep learning applications via backpropagation.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Oneadapt: Fast adapta- tion for deep learning applications via backpropagation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.083135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.083135Z digest=sha256:a30c66fbca81f08c6260f1249c76ab33ed00dc569c3ff19b37661d797593426c

Observation 85a1258d-f65f-47ea-86ae-05737e63e0c4 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.087432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.087432Z digest=sha256:eb20f6c02f2e7085fc43475315188344ff409369613924373ff1ceb653ecd5af

Observation 6b7fc549-6faf-42cb-82e9-bc59ca50a6b4 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.091621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.091621Z digest=sha256:cec714828d2f048d0ec93edd3d319c64400ebea965eb9372bd5a52f2a998a5b5

Observation 379c6fbb-fec8-47e1-9d20-3021728b2615 · outbound

This paper cites Dgsf: Disaggregated gpus for serverless functions.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Dgsf: Disaggregated gpus for serverless functions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.095749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.095749Z digest=sha256:50046b3a7a17c116a98086b48b5fab8489da8d8999536e67793a581142834c70

Observation 818656c9-a51e-4238-b955-4067ca201da4 · outbound

This paper cites Rdma over ethernet for distributed training at meta scale.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Rdma over ethernet for distributed training at meta scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.099856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.099856Z digest=sha256:f11eb51241e5404a53f30ecd34043e1de473684a0d58ac5d0dc4ae4931e5662a

Observation d63b93d0-c65c-4979-947e-2eec9e855e89 · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.104095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.104095Z digest=sha256:eb937c278548a61ba3cc24efb39ebcf5e9627f9809f9b9cd0c1b057fdb6dcb17

Observation ee927be3-9751-4415-8937-0d7a837b8199 · outbound

This paper cites A gpgpu transparent virtualization component for high performance computing clouds.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A gpgpu transparent virtualization component for high performance computing clouds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.108698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.108698Z digest=sha256:74e99cb2e69614135ecda796cb4cd4fb0a7a0ab086550698dcca694cfe90a527

Observation ab9ae254-bd5c-408d-9972-c3ec21c2066f · outbound

This paper cites Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.113329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.113329Z digest=sha256:35a1b29588913186aa14564b8630ab928c42c21541b00a1fb0c4da05d7f7408c

Observation 9e82ee4d-322c-478b-b98b-eddd16370dcd · outbound

This paper cites Kace: Kernel-aware colocation for effi- cient gpu spatial sharing.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kace: Kernel-aware colocation for effi- cient gpu spatial sharing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.120130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.120130Z digest=sha256:0915a40e930db7b75e510dd04362c6608d6b6e79d089cd2ef553f02e82272cdb

Observation b4ca0d1e-f0ff-4e5f-a691-c506cc9959ad · outbound

This paper cites Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.125670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.125670Z digest=sha256:dfa3f6e949d904483caab34f2936cc14548ad4e1ce477d7d838bcb1fb11a66b3

Observation 834d32cb-73f6-4e79-96b7-042367f4f229 · outbound

This paper cites Multi-agent collaborative infer- ence via dnn decoupling: Intermediate feature compres- sion and edge learning.IEEE Transactions on Mobile Computing, 22(10):6041–6055, 2023.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Multi-agent collaborative infer- ence via dnn decoupling: Intermediate feature compres- sion and edge learning.IEEE Transactions on Mobile Computing, 22(10):6041–6055, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.130167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.130167Z digest=sha256:35f1bb43a3f56007e9bce6c01b36146064aa1d440b4cd4e51012d1734bd90b7a

Observation 6a4fff8d-2f02-44bf-a2e0-5806d25bd8c6 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.134612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.134612Z digest=sha256:8ca397e94ee6dc1826a01a92ebb547db044d8c600e889cd455418a31d3e2c094

Observation 6af4123f-65db-4204-9fe1-e3ee783c1dfa · outbound

This paper cites In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 87–101, 2023.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 87–101, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.138890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.138890Z digest=sha256:e9fb09cc1ec646098fc777258b098c19b7d5d112e762f70519bf2621a998b8b5

Observation 895f8f53-7860-40f3-8e5f-32e8a2ff6f53 · outbound

This paper cites Charlie Hu, Xiaojun Lin, and Nan Deng.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Charlie Hu, Xiaojun Lin, and Nan Deng

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.143422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.143422Z digest=sha256:2d58d9a93fd9dbb51b04635d8c9014f5d605e45d3ca2f7aba9a5bdb0a6c55e4e

Observation a12eb9f5-097c-4c74-ade9-57c9a012db8c · outbound

This paper cites A house united within itself: Slo-awareness for on-premises containerized ml inference clusters via faro.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A house united within itself: Slo-awareness for on-premises containerized ml inference clusters via faro

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.147743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.147743Z digest=sha256:e43ad4c84857f29586e0ce30d5db81cc056faeb0e82fb8c00c30aacf4d0328de

Observation 817605af-c388-4aec-bca1-2b9a568f9d27 · outbound

This paper cites Mor- ley Mao.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Mor- ley Mao

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.152064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.152064Z digest=sha256:79111c24b659d85206a68327bc5f664678d71e462b6c40ac5ff9efc983ef8e3a

Observation 6d512668-de1e-402f-84e4-697d9c7ef85a · outbound

This paper cites Deepum: Tensor migration and prefetching in unified memory.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Deepum: Tensor migration and prefetching in unified memory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.156399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.156399Z digest=sha256:6a74bccea4c7b1ad4083e9ea0dcc84269997cce0dad92e03d5f9da4627feca34

Observation b6ab778c-3b3f-4b23-b486-80fd42e4dd23 · outbound

This paper cites A neural-network- based realization of in-network computation for the in- ternet of things.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A neural-network- based realization of in-network computation for the in- ternet of things

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.161088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.161088Z digest=sha256:7b71dcce5959d7e537e3577d0a731e680d910ed856c678a6b5fe8ffb52eeee7e

Observation 5ee45b07-34d1-42ca-b61d-e5cee2fea889 · outbound

This paper cites {SuperServe}:{Fine-Grained} inference serving for unpredictable workloads.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {SuperServe}:{Fine-Grained} inference serving for unpredictable workloads

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.165512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.165512Z digest=sha256:8bb0603ddce033c2b2f23d1cc74934285d5b6e63ac70f5a5cf44156bb0f6d59b

Observation 54e34d1c-f630-4ea1-b661-51f6b00f7317 · outbound

This paper cites A survey on in-network computing: Programmable data plane and technology specific applications.IEEE Communications Surveys & Tutorials, 25(1):701–761, 2023.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey on in-network computing: Programmable data plane and technology specific applications.IEEE Communications Surveys & Tutorials, 25(1):701–761, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.169742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.169742Z digest=sha256:97e8474a65df4e9a005e826fc97cfc503a6843f45579ee623c9cedd5cec3aa7c

Observation d438c829-33e9-49ed-9b5d-8fc735a80211 · outbound

This paper cites Navigator: Dynamic multi-kernel scheduling to improve gpu per- formance.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Navigator: Dynamic multi-kernel scheduling to improve gpu per- formance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.173949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.173949Z digest=sha256:f98713734f72e761505a9e31a12d2d99c25cabd9b553362bbaaa123fb4c6d931

Observation 57aef14b-de47-4fa5-a3d9-552485dd5783 · outbound

This paper cites Efficient memory manage- ment for large language model serving with pagedatten- tion.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient memory manage- ment for large language model serving with pagedatten- tion

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.178278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.178278Z digest=sha256:610b04181658ec7a9e745b71c33484aa43eba73e6f88a532af5d67c6b62de815

Observation 2784287b-4418-4329-bb95-87bea9ede7f6 · outbound

This paper cites Forecasting gpu performance for deep learning train- ing and inference.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Forecasting gpu performance for deep learning train- ing and inference

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.182444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.182444Z digest=sha256:c07b8b26a9fc4b264e2d74e63d1a420f3ca0a3e910ee32bf5ad52f0dbd53b810

Observation 93854fcd-8bc9-495f-8ba3-0e35228838d0 · outbound

This paper cites A survey on large language model acceleration based on kv cache management, 2025.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey on large language model acceleration based on kv cache management, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.186651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.186651Z digest=sha256:b305e2906ff9e5821d63bcf74f9f55924a3ce0499e9315f8fd4d84371f492d52

Observation 59f34cb6-8c73-46c1-bac8-e7ddd46290ff · outbound

This paper cites THC: Accelerating distributed deep learning using ten- sor homomorphic compression.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs THC: Accelerating distributed deep learning using ten- sor homomorphic compression

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.190669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.190669Z digest=sha256:ddab7f0330a42e4046ac5cccf2e313dbcf71d6bc92808ab45eee5adaf8ea928c

Observation 8832f8da-a903-4fa7-a36e-7e876239de19 · outbound

This paper cites Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.195265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.195265Z digest=sha256:7a9d7576d0313173b52b4bb052ab3fd2e19f72086915779da283d0b1de185dba

Observation 4f757e3e-db2e-4b27-86fc-7c7bf30891d0 · outbound

This paper cites Incbricks: To- ward in-network computation with an in-network cache.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Incbricks: To- ward in-network computation with an in-network cache

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.199805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.199805Z digest=sha256:1747c50d1d40ab488a60974da2c62dd3ee72b91bc50d7bea7e91980c762f97a4

Observation 9cc3512f-8635-4f4a-a61d-ab588b7d3118 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large lan- guage model serving.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Cachegen: Kv cache compression and streaming for fast large lan- guage model serving

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.204037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.204037Z digest=sha256:4cb15a644f0d15e1f084837064ca5bcb77a951eacf9f5d4c54396256ec5b3739

Observation 74a2a691-484f-422e-ad78-ae137fa5c5cf · outbound

This paper cites A convnet for the 2020s.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A convnet for the 2020s

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.209596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.209596Z digest=sha256:6a250a0681884fb255738cc243b6240f1f4b53bcc307736e5decf0c87c945cfc

Observation d225d4c4-314e-4207-a6cf-e764f6337ada · outbound

This paper cites A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.216782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.216782Z digest=sha256:51306c553869755b899dc83715d602b520e24a876a9442b3ef8039ebfbd489f1

Observation 5147f1f6-46fc-4e50-8ab4-02134dfaec9b · outbound

This paper cites A survey of storage systems in the rdma era.IEEE Transac- tions on Parallel and Distributed Systems, 33(12):4395– 4409, 2022.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey of storage systems in the rdma era.IEEE Transac- tions on Parallel and Distributed Systems, 33(12):4395– 4409, 2022

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.228099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.228099Z digest=sha256:72976037a7f1b4b6917b9687159c83557ab475baaa6950e6efc15a3e445f3951

Observation 803fc393-b9d1-4a27-9bda-af2fae79b699 · outbound

This paper cites Skyserve: Serving ai mod- els across regions and clouds with spot instances.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Skyserve: Serving ai mod- els across regions and clouds with spot instances

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.240104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.240104Z digest=sha256:0200fa2bcf2dce9fab29d3909272b29f2f06c02ca87e5c12dc908e1fa96e5171

Observation 3f7e2269-9f65-4a92-9f83-c07e307cbf42 · outbound

This paper cites Efficient scheduling policies for Microsecond-Scale tasks.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient scheduling policies for Microsecond-Scale tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.252541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.252541Z digest=sha256:b6ecec2de9c4db836ac1f41f76a40fce9a508c48359da2f029f3014e6fdf12fd

Observation 5bf9d1f6-3277-4920-99de-e1685840341c · outbound

This paper cites To- ward performance-portable petsc for gpu-based exascale systems.Parallel Computing, 108:102831, 2021.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs To- ward performance-portable petsc for gpu-based exascale systems.Parallel Computing, 108:102831, 2021

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.266736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.266736Z digest=sha256:90d141a589fd2fa87ba27205e095fd304d71963ef5557ac0b4f8ab6dea31622c

Observation 5220bb3a-f75b-48e1-ad6c-dcef1b447bc0 · outbound

This paper cites Porting warpx to gpu- accelerated platforms.Parallel Computing, 108:102833, 2021.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Porting warpx to gpu- accelerated platforms.Parallel Computing, 108:102833, 2021

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.285770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.285770Z digest=sha256:d34111a2d46da0d55f4e2e81cedb0136341277a8268e810b55cac26cb2e3553b

Observation ba9cc23b-1df0-4c96-a5f0-166389d8f973 · outbound

This paper cites Jellyfish: Timely inference serving for dynamic edge networks.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Jellyfish: Timely inference serving for dynamic edge networks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.314644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.314644Z digest=sha256:08f3ef10fac537e0434336f0a1b370ecb0f0f3451354b07bb6645363835b82bc

Observation 0bf2f227-ab26-4dc9-9b0a-8a3d888cce04 · outbound

This paper cites Bringing umap closer to the speed of light with gpu acceleration.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Bringing umap closer to the speed of light with gpu acceleration

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.330029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.330029Z digest=sha256:d027496db72f0912c3731280c7af8e77c3c93b88d4fdd227d308678caed1d91b

Observation e333a6ea-f011-4074-9c60-244140e14361 · outbound

This paper cites Nvidia nvswitch: The world’s highest- bandwidth on-node switch.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Nvidia nvswitch: The world’s highest- bandwidth on-node switch

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.337088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.337088Z digest=sha256:9822d412b1fb6e21e14b2b24680431485f3b148fd0d4d35c0a849725562e0128

Observation e97e016e-0e66-4ac1-a005-0ec3c96f368a · outbound

This paper cites Gemel: Model merging for memory-efficient,real-time video analytics at the edge.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Gemel: Model merging for memory-efficient,real-time video analytics at the edge

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.344509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.344509Z digest=sha256:112a8f9f1e946794d32f8a0c249d84f3a83971c366df4697a8c469558c229b00

Observation ba9ad0f5-3866-4e5e-8ef5-184567811aed · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.352874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.352874Z digest=sha256:a04eeb96ac28f972aaf88f06fb6fb854a3025fd0e1f8cd0303fe1cc6014abfaa

Observation 5e56c7da-083f-4b4d-9639-22c3682c3071 · outbound

This paper cites {CASSINI}:{Network-Aware} job scheduling in machine learning clusters.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {CASSINI}:{Network-Aware} job scheduling in machine learning clusters

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.364647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.364647Z digest=sha256:96845d6b8fc67c0ec1f4fc74876ae00b8219f3e1038523f6c13d7381761b1028

Observation b84179bc-e235-4d4c-ba79-78842ad7de83 · outbound

This paper cites Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.372960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.372960Z digest=sha256:ed809affb82a3dc21c2a3c914f7eec5d439fbb430e61b2eb0f1b151b81b40588

Observation 2900bced-7c98-4149-8e37-75d00f579e1d · outbound

This paper cites Enabling large dynamic neural net- work training with learning-based memory manage- ment.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Enabling large dynamic neural net- work training with learning-based memory manage- ment

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.379969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.379969Z digest=sha256:ae20434b6a8a0a2e3dbe8e56d4dbea5c8067ce6ae10668c5cc4be203506f39ea

Observation 748e008d-6c97-475a-b8f2-1ed5e89acc24 · outbound

This paper cites {Cloud-LoRa}: Enabling cloud radio access {LoRa} networks using reinforce- ment learning based {Bandwidth-Adaptive} compres- sion.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {Cloud-LoRa}: Enabling cloud radio access {LoRa} networks using reinforce- ment learning based {Bandwidth-Adaptive} compres- sion

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.387513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.387513Z digest=sha256:6fc0642ce70110e414f74013ede140a494f4004c300a1f300fef44c2f465eb95

Observation b275ef95-bf4f-448f-9b27-da8cd39c9bc1 · outbound

This paper cites Exploiting simultaneous communications to accelerate data parallel distributed deep learning.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Exploiting simultaneous communications to accelerate data parallel distributed deep learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.394196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.394196Z digest=sha256:2bd6a5a7820dcbe6f48104e25fa568e98bb96768ce6ebfe94dedb81e0aa26df1

Observation 693cf1f8-96ac-43b0-b349-607ef9f9dd10 · outbound

This paper cites Orion: Interference-aware, fine-grained gpu sharing for ml ap- plications.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Orion: Interference-aware, fine-grained gpu sharing for ml ap- plications

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.400877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.400877Z digest=sha256:64f9163f570f52a9b9c9f47af0cca2c6a43d5a3708e78466f68cfdeeb5cd1d7e

Observation 1e32dc8d-be47-41bc-89af-2f2e08ef8975 · outbound

This paper cites gremote: Cloud ren- dering on gpu resource pool based on api-forwarding.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs gremote: Cloud ren- dering on gpu resource pool based on api-forwarding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.416566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.416566Z digest=sha256:b79c69d89cf20d9db1b84c42b40b7189f2af54235925a680cf48a72464e26782

Observation 188d3cf4-d525-4a75-9005-8c8caff3d0ab · outbound

This paper cites Lammps-a flexible simu- lation tool for particle-based materials modeling at the atomic, meso, and continuum scales.Computer physics communications, 271:108171, 2022.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Lammps-a flexible simu- lation tool for particle-based materials modeling at the atomic, meso, and continuum scales.Computer physics communications, 271:108171, 2022

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.440322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.440322Z digest=sha256:c948160dd2984bcd8311f767ed17e90ef88ed5862befae3228885416c676da30

Observation d9075125-c546-4c9d-8edd-367ac0d72b9c · outbound

This paper cites Plssvm—parallel least squares support vector machine.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Plssvm—parallel least squares support vector machine

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.468036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.468036Z digest=sha256:7e96995945fb73e6f447fa0e7dfb1383145acb536dce56a062834703f29f7690

Observation 5771c190-475e-4cae-95be-a852315c3cea · outbound

This paper cites Aqua: Network-accelerated memory offloading for llms in scale-up gpu domains.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Aqua: Network-accelerated memory offloading for llms in scale-up gpu domains

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.522658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.522658Z digest=sha256:2aecfce12c93715c1316a296a3148acbfc1709467cda28515979decd91fc2476

Observation 9ecc8377-7758-414f-9506-64ac79e2ed9a · outbound

This paper cites Coflow scheduling for llm training.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Coflow scheduling for llm training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.576555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.576555Z digest=sha256:7e3d69386ef0d78ef740ed7050126a67c0266dc8afa4387925dff345b1ae3dd3

Observation 158cf9d6-e008-4c7c-a59e-73473353b2b2 · outbound

This paper cites Characterizing Network Requirements for GPU API Remoting in AI Applications.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Characterizing Network Requirements for GPU API Remoting in AI Applications

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.630913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.630913Z digest=sha256:8d5240f1b9f2aee73ca5e8cdf7ccbf591a94581b8c85d45d7a7d04b91e3314fe

Observation b7884337-5363-4d05-b8a1-1ae263e751ab · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.675274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.675274Z digest=sha256:3ff52d0bfdc535e1a34b6d010b37794ebdb6ba4e2e0687519a2ab8c3634889bd

Observation b25a7077-5fa2-4b4e-b832-c3e0c4d9b323 · outbound

This paper cites Beware of fragmentation: Scheduling {GPU-Sharing} workloads with fragmentation gradient descent.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Beware of fragmentation: Scheduling {GPU-Sharing} workloads with fragmentation gradient descent

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.755151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.755151Z digest=sha256:4b540e962a4867b96ee490ad61424f9d228805ce8ffebd8ed2e5b7bb961904ee

Observation ea9fd065-ddb7-48c9-a4c8-7b2ef3390bcb · outbound

This paper cites Transparent {GPU} sharing in container clouds for deep learning workloads.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Transparent {GPU} sharing in container clouds for deep learning workloads

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.798443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.798443Z digest=sha256:e9141df9b389f8191eceede0b11c2d5a81ad7331c279eda84fef263932d20eb8

Observation e46584d6-1c14-4058-bc99-72391d454c51 · outbound

This paper cites Yan, and Junchen Jiang.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Yan, and Junchen Jiang

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.872455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.872455Z digest=sha256:3a1b7b3046040491193fa5c2aa3d2ee04a759319c205c13e99fe8bcd18c66d6e

Observation 04c59993-3ce9-4598-a883-0405c0d8b2a1 · outbound

This paper cites Efficient tensor offloading for large deep-learning model training based on compute express link.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient tensor offloading for large deep-learning model training based on compute express link

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:11.914851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:11.914851Z digest=sha256:4f6e905f2e96e5dc92379cd744518350c556322b19fd8428ab01c4b9274e0e7d

Observation f0699efe-20a1-4124-a9e0-7c8f547898d5 · outbound

This paper cites Infless: a native serverless system for low-latency, high- throughput inference.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Infless: a native serverless system for low-latency, high- throughput inference

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.019245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.019245Z digest=sha256:2dcc49564ca953718cbe786843488a30ce6fc6dba8749f2a15abe08dd596f63a

Observation e87e1bfe-f4ef-4c12-85e7-b6ea4ff2bef8 · outbound

This paper cites Deep compressive offloading: Speeding up neural net- work inference by trading edge computation for network latency.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Deep compressive offloading: Speeding up neural net- work inference by trading edge computation for network latency

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.078779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.078779Z digest=sha256:38a551ebf7f74db6197203883c3932bb33c4c0980293bfd25245d2c492f98e1b

Observation d9dce3a7-97b3-4d3a-8f78-165427374f2d · outbound

This paper cites Horus: granular in-network task sched- uler for cloud datacenters.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Horus: granular in-network task sched- uler for cloud datacenters

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.186797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.186797Z digest=sha256:bfd83a962d78d3f16a9f9552cbc3f7be6ded24f3b408e09aff9bc4b0099a8b52

Observation 0525d198-f573-46b9-bfc1-ec895eb25d05 · outbound

This paper cites Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.308277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.308277Z digest=sha256:8ed1d01beae1fbd2b676821bce32d509e64dd1a269da01abc4faf1fdd76ecaef

Observation 32543cfe-c61d-4c10-80df-6a54ead58c8c · outbound

This paper cites an unresolved cited work.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.369125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.369125Z digest=sha256:312a0c798ae755a5982c64176ce06d5c0d004e65e8e4d2a25880dacd6a57dcc5

Observation 435783a6-c85b-4a06-95aa-502e3d4c7815 · outbound

This paper cites Expel: Llm agents are ex- periential learners.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Expel: Llm agents are ex- periential learners

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.439101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.439101Z digest=sha256:55fa877cbb409eacf4127e0cc56c32d41e077e50c6e7f0dbb005f1c8fb66d055

Observation 9ba649ac-fb75-41ad-98e5-0ec9e7e9ca01 · outbound

This paper cites Efficient {Direct-Connect} topologies for collective communications.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient {Direct-Connect} topologies for collective communications

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.509184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.509184Z digest=sha256:06e7d9a7aba306deefde5ca35c9cc644a3a47e9c64b2f97028ed6d6ec83a4235

Observation da20ac6b-e61e-4283-b465-2ecc048e4ac1 · outbound

This paper cites Tally: Non-intrusive performance isolation for concur- rent deep learning workloads.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Tally: Non-intrusive performance isolation for concur- rent deep learning workloads

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.594996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.594996Z digest=sha256:45c0332e1d3f4002bbfaf447cd52fa5eb80e238e96f9656290a13584685f1d32

Observation 49ea237c-f11d-4484-a700-a0a96ebfd31c · outbound

This paper cites Sglang: Efficient execution of structured language model pro- grams.Advances in neural information processing sys- tems, 37:62557–62583, 2024.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Sglang: Efficient execution of structured language model pro- grams.Advances in neural information processing sys- tems, 37:62557–62583, 2024

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.665111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.665111Z digest=sha256:9b10986b10c2c2bffba4db40f8271b19da0dfdbec75fa1e68e9e150c39b7d4a5

Observation ab65e9fa-816f-4f3a-afe8-2b559f56f520 · outbound

This paper cites Kernelet: High- throughput gpu kernel executions with dynamic slicing and scheduling.IEEE Transactions on Parallel and Distributed Systems, 25(6):1522–1532, 2014.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kernelet: High- throughput gpu kernel executions with dynamic slicing and scheduling.IEEE Transactions on Parallel and Distributed Systems, 25(6):1522–1532, 2014

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T03:40:12.683698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.683698Z digest=sha256:7909ec2308d5a1390ebcc0ecdceba4cca3d2ea1b9c0067317f629c07653f6bd4

Observation dbee7ac1-5724-4102-8cf5-d0e06cce693b · outbound

This paper cites Megascale-infer: Efficient mixture-of-experts model serving with disag- gregated expert parallelism.

Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Megascale-infer: Efficient mixture-of-experts model serving with disag- gregated expert parallelism

Reference 86

Resolution
malformed identifier
no resolver link, observed 2026-08-01T03:40:12.723229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:40:12.723229Z digest=sha256:a1b6444cc61d84b94f43d3d72d8e83d3cf6c9535cb7e4209e0dd2d5449138f6c

Pith citing papers

No inbound Pith citation observations are available.