Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:40:12.723229Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2607.23115.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:40:12.723229Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
86 of 86 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e81fbd40-63ff-4fc2-bd55-05114131724b · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4da32b21-7b4a-4b7e-b44f-f20dc4ec28b5 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5039eec-7bfa-4d8e-85b5-96275f47ad23 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee08d559-9968-4f12-9b7c-842efcc8f3e2 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75dcedb9-39d9-49cb-8803-3983dd7e0a82 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf32578-7480-4dd9-8ee9-a728a93cabdc · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398ec4d2-4f66-46a8-8a2b-aebc9527be8d · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d493f9-cf75-428d-9f6b-79d4780f6d6c · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f91212a2-8ee2-4d32-b5d3-de5ef048aaf6 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf139c7-f5ae-4003-a1d8-05b1a98d601e · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Scissionlite: Accelerating distributed deep learning with lightweight data compression for iiot.IEEE Transac- tions on Industrial Informatics, 20(10):11950–11960, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd4fa3d-553e-4145-a395-0831c5726fd3 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Tooth: Toward optimal balance of video {QoE} and redundancy cost by {Fine- Grained}{FEC} in cloud gaming streaming
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a98f0dba-9b93-43cf-930d-3a588ffe6c00 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Crux: Gpu-efficient communication scheduling for deep learning training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aba9b77-b011-46a8-995d-5a2d2eb23de6 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Eva: Cost- efficient cloud-based cluster scheduling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5af2baa5-b736-4b88-9b16-8515905d1d06 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kernel oper- ations on the gpu, with autodiff, without memory over- flows.Journal of Machine Learning Research, 22(74):1– 6, 2021
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bbd9110-e6f2-46af-85ee-526807ff3686 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Remote procedure call as a managed system service
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 591726b1-51c1-47dc-b966-fdb425e9b7e1 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Multiplexing dynamic deep learn- ing workloads with slo-awareness in gpu clusters
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45221d94-c5ea-4cb3-a447-c93037871049 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {GRACE}:{Loss- Resilient}{Real-Time} video through neural codecs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7379fdeb-82e1-4252-8bb5-d5414e62298a · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Flashattention: Fast and memory- efficient exact attention with io-awareness.Advances in neural information processing systems, 35:16344– 16359, 2022
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4812f51-687c-4fc5-a585-03122aa6c47f · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Oneadapt: Fast adapta- tion for deep learning applications via backpropagation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a1258d-f65f-47ea-86ae-05737e63e0c4 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7fc549-6faf-42cb-82e9-bc59ca50a6b4 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 379c6fbb-fec8-47e1-9d20-3021728b2615 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Dgsf: Disaggregated gpus for serverless functions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 818656c9-a51e-4238-b955-4067ca201da4 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Rdma over ethernet for distributed training at meta scale
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d63b93d0-c65c-4979-947e-2eec9e855e89 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee927be3-9751-4415-8937-0d7a837b8199 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A gpgpu transparent virtualization component for high performance computing clouds
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab9ae254-bd5c-408d-9972-c3ec21c2066f · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e82ee4d-322c-478b-b98b-eddd16370dcd · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kace: Kernel-aware colocation for effi- cient gpu spatial sharing
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4ca0d1e-f0ff-4e5f-a691-c506cc9959ad · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834d32cb-73f6-4e79-96b7-042367f4f229 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Multi-agent collaborative infer- ence via dnn decoupling: Intermediate feature compres- sion and edge learning.IEEE Transactions on Mobile Computing, 22(10):6041–6055, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4fff8d-2f02-44bf-a2e0-5806d25bd8c6 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af4123f-65db-4204-9fe1-e3ee783c1dfa · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), pages 87–101, 2023
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 895f8f53-7860-40f3-8e5f-32e8a2ff6f53 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Charlie Hu, Xiaojun Lin, and Nan Deng
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a12eb9f5-097c-4c74-ade9-57c9a012db8c · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A house united within itself: Slo-awareness for on-premises containerized ml inference clusters via faro
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 817605af-c388-4aec-bca1-2b9a568f9d27 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Mor- ley Mao
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d512668-de1e-402f-84e4-697d9c7ef85a · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Deepum: Tensor migration and prefetching in unified memory
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ab778c-3b3f-4b23-b486-80fd42e4dd23 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A neural-network- based realization of in-network computation for the in- ternet of things
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee45b07-34d1-42ca-b61d-e5cee2fea889 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {SuperServe}:{Fine-Grained} inference serving for unpredictable workloads
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e34d1c-f630-4ea1-b661-51f6b00f7317 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey on in-network computing: Programmable data plane and technology specific applications.IEEE Communications Surveys & Tutorials, 25(1):701–761, 2023
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d438c829-33e9-49ed-9b5d-8fc735a80211 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Navigator: Dynamic multi-kernel scheduling to improve gpu per- formance
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57aef14b-de47-4fa5-a3d9-552485dd5783 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient memory manage- ment for large language model serving with pagedatten- tion
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2784287b-4418-4329-bb95-87bea9ede7f6 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Forecasting gpu performance for deep learning train- ing and inference
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93854fcd-8bc9-495f-8ba3-0e35228838d0 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey on large language model acceleration based on kv cache management, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59f34cb6-8c73-46c1-bac8-e7ddd46290ff · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs THC: Accelerating distributed deep learning using ten- sor homomorphic compression
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8832f8da-a903-4fa7-a36e-7e876239de19 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f757e3e-db2e-4b27-86fc-7c7bf30891d0 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Incbricks: To- ward in-network computation with an in-network cache
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cc3512f-8635-4f4a-a61d-ab588b7d3118 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Cachegen: Kv cache compression and streaming for fast large lan- guage model serving
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74a2a691-484f-422e-ad78-ae137fa5c5cf · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A convnet for the 2020s
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d225d4c4-314e-4207-a6cf-e764f6337ada · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5147f1f6-46fc-4e50-8ab4-02134dfaec9b · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs A survey of storage systems in the rdma era.IEEE Transac- tions on Parallel and Distributed Systems, 33(12):4395– 4409, 2022
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803fc393-b9d1-4a27-9bda-af2fae79b699 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Skyserve: Serving ai mod- els across regions and clouds with spot instances
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7e2269-9f65-4a92-9f83-c07e307cbf42 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient scheduling policies for Microsecond-Scale tasks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf9d1f6-3277-4920-99de-e1685840341c · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs To- ward performance-portable petsc for gpu-based exascale systems.Parallel Computing, 108:102831, 2021
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5220bb3a-f75b-48e1-ad6c-dcef1b447bc0 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Porting warpx to gpu- accelerated platforms.Parallel Computing, 108:102833, 2021
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba9cc23b-1df0-4c96-a5f0-166389d8f973 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Jellyfish: Timely inference serving for dynamic edge networks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bf2f227-ab26-4dc9-9b0a-8a3d888cce04 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Bringing umap closer to the speed of light with gpu acceleration
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e333a6ea-f011-4074-9c60-244140e14361 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Nvidia nvswitch: The world’s highest- bandwidth on-node switch
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e97e016e-0e66-4ac1-a005-0ec3c96f368a · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Gemel: Model merging for memory-efficient,real-time video analytics at the edge
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba9ad0f5-3866-4e5e-8ef5-184567811aed · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e56c7da-083f-4b4d-9639-22c3682c3071 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {CASSINI}:{Network-Aware} job scheduling in machine learning clusters
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84179bc-e235-4d4c-ba79-78842ad7de83 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2900bced-7c98-4149-8e37-75d00f579e1d · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Enabling large dynamic neural net- work training with learning-based memory manage- ment
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 748e008d-6c97-475a-b8f2-1ed5e89acc24 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs {Cloud-LoRa}: Enabling cloud radio access {LoRa} networks using reinforce- ment learning based {Bandwidth-Adaptive} compres- sion
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b275ef95-bf4f-448f-9b27-da8cd39c9bc1 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Exploiting simultaneous communications to accelerate data parallel distributed deep learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 693cf1f8-96ac-43b0-b349-607ef9f9dd10 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Orion: Interference-aware, fine-grained gpu sharing for ml ap- plications
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e32dc8d-be47-41bc-89af-2f2e08ef8975 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs gremote: Cloud ren- dering on gpu resource pool based on api-forwarding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 188d3cf4-d525-4a75-9005-8c8caff3d0ab · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Lammps-a flexible simu- lation tool for particle-based materials modeling at the atomic, meso, and continuum scales.Computer physics communications, 271:108171, 2022
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9075125-c546-4c9d-8edd-367ac0d72b9c · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Plssvm—parallel least squares support vector machine
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5771c190-475e-4cae-95be-a852315c3cea · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Aqua: Network-accelerated memory offloading for llms in scale-up gpu domains
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ecc8377-7758-414f-9506-64ac79e2ed9a · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Coflow scheduling for llm training
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158cf9d6-e008-4c7c-a59e-73473353b2b2 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Characterizing Network Requirements for GPU API Remoting in AI Applications
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7884337-5363-4d05-b8a1-1ae263e751ab · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b25a7077-5fa2-4b4e-b832-c3e0c4d9b323 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Beware of fragmentation: Scheduling {GPU-Sharing} workloads with fragmentation gradient descent
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9fd065-ddb7-48c9-a4c8-7b2ef3390bcb · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Transparent {GPU} sharing in container clouds for deep learning workloads
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e46584d6-1c14-4058-bc99-72391d454c51 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Yan, and Junchen Jiang
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c59993-3ce9-4598-a883-0405c0d8b2a1 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient tensor offloading for large deep-learning model training based on compute express link
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0699efe-20a1-4124-a9e0-7c8f547898d5 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Infless: a native serverless system for low-latency, high- throughput inference
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e87e1bfe-f4ef-4c12-85e7-b6ea4ff2bef8 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Deep compressive offloading: Speeding up neural net- work inference by trading edge computation for network latency
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9dce3a7-97b3-4d3a-8f78-165427374f2d · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Horus: granular in-network task sched- uler for cloud datacenters
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0525d198-f573-46b9-bfc1-ec895eb25d05 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32543cfe-c61d-4c10-80df-6a54ead58c8c · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435783a6-c85b-4a06-95aa-502e3d4c7815 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Expel: Llm agents are ex- periential learners
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba649ac-fb75-41ad-98e5-0ec9e7e9ca01 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Efficient {Direct-Connect} topologies for collective communications
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da20ac6b-e61e-4283-b465-2ecc048e4ac1 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Tally: Non-intrusive performance isolation for concur- rent deep learning workloads
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ea237c-f11d-4484-a700-a0a96ebfd31c · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Sglang: Efficient execution of structured language model pro- grams.Advances in neural information processing sys- tems, 37:62557–62583, 2024
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab65e9fa-816f-4f3a-afe8-2b559f56f520 · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Kernelet: High- throughput gpu kernel executions with dynamic slicing and scheduling.IEEE Transactions on Parallel and Distributed Systems, 25(6):1522–1532, 2014
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbee7ac1-5724-4102-8cf5-d0e06cce693b · outbound
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs Megascale-infer: Efficient mixture-of-experts model serving with disag- gregated expert parallelism
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.