Pith. sign in

Paper Citation Record · LEDGER

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines

As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2412.14335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14335 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:23:19.004428Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:27:43.845408Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:30:33.076250Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 890877d4-5f77-49ac-b521-e89d0ff47993 · outbound

This paper cites A bridging model for parallel computation,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines A bridging model for parallel computation,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.468135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.468135Z digest=sha256:1765085294030d3dc933de28ab9a1a4c5ed1f73f5f27ee7512d894488c408306

Observation d1a0b163-db15-42ec-8fb3-c4784c7ce237 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Pytorch fsdp: Experiences on scaling fully sharded data parallel,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.500285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.500285Z digest=sha256:0f38d044d6db0aeed0e2c6ed506e269a9246e1c55a78a84149a511216e8e2dea

Observation 3c7b4e4f-8200-4234-ad70-3c5a017746e5 · outbound

This paper cites NanoFlow: Towards Optimal Large Language Model Serving Throughput.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines NanoFlow: Towards Optimal Large Language Model Serving Throughput

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.596808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.596808Z digest=sha256:f14654a5a3a7f524b88b5b5d023dabe4701f49e76f7c7a22624c210b75997b75

Observation c32deaf5-bd94-4184-bdab-1ec72303a5b3 · outbound

This paper cites The Llama 3 Herd of Models.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.637910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.637910Z digest=sha256:3dbfb425ba9a9d7d1839873a7870c5dafcd847d6075c56d7517ed9fbcedbe2eb

Observation 1ae43603-382f-4fc8-968c-5d3e08f8b168 · outbound

This paper cites {ARK}:{GPU-driven} code execution for distributed deep learning,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines {ARK}:{GPU-driven} code execution for distributed deep learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.799105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.651544Z digest=sha256:0fdae855c2a6119cd0615c06ffda35d45435682cb76f6b25f9ecf103560ef450

Observation e6ebe5e3-331a-4d90-98ad-863da059bdbe · outbound

This paper cites 11.1 amd instincttm mi300 series modular chiplet package – hpc and ai accelerator for exa-class systems,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines 11.1 amd instincttm mi300 series modular chiplet package – hpc and ai accelerator for exa-class systems,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.783250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.658910Z digest=sha256:8baedfa986ad7df60d3ec549e70630e298834a2050ec43ff8e90260d6d5f0530

Observation 6a1f410e-90f4-4d51-b122-aa53bde69838 · outbound

This paper cites AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.772487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.670888Z digest=sha256:79c42d871e73c00da4613ec614e4c420689affc57cc37109585ba7a1f10ad609

Observation 1dd6be83-e635-4c87-9d31-354fad14baf0 · outbound

This paper cites The AMD CDNA™ 3 architecture,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The AMD CDNA™ 3 architecture,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.761505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.681382Z digest=sha256:e2ece8547868268ff4d9c5ab10df64e636782ed7beb52a5d06628ada5cac9995

Observation c75bc5a0-a566-4c90-846e-9ad51407dc64 · outbound

This paper cites HSA Runtime API and runtime for ROCm,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines HSA Runtime API and runtime for ROCm,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.750820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.690115Z digest=sha256:5cf5376d93b7c132324f837a9090ed207494d56f12b3750d601b5cdc3ac21f98

Observation ba5d4042-d53e-4154-9a86-8291b703c223 · outbound

This paper cites HIP: C++ Heterogeneous-Compute Interface for Portability,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines HIP: C++ Heterogeneous-Compute Interface for Portability,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.739086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.698583Z digest=sha256:50953c0a165c9d9d033f09e2903b8ac083a42e31d3a1620e028bc36f05d7717e

Observation 243312cf-886f-423c-86bb-63307a79a35f · outbound

This paper cites Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.727367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.703200Z digest=sha256:adcd5c59ef1462c1f14f124343098edc969a29aa51d6e285802f792b89600467

Observation 35802812-fcae-49fc-96f9-f86b8d45709e · outbound

This paper cites AMD ROCm™ Software,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD ROCm™ Software,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.716251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.707350Z digest=sha256:4d817e900419930cd0b46e7a3199835ce908657ff03fd6fb4123db6521ea1f74

Observation f858a095-6ea6-4c4f-8ac3-3a7610b2952d · outbound

This paper cites ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.703069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.711930Z digest=sha256:fbe70ff0013bf07bb5d66b36d037aa4334e34a0530f9079959003f5599c32f7b

Observation 66cd9b53-7e39-4089-b23f-8874fe049d04 · outbound

This paper cites ROCm Communication Collectives Library (RCCL).

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm Communication Collectives Library (RCCL)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.690679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.716063Z digest=sha256:c157c4d4f97988bc959accc91d2a7501b4a7afeec625fc1252432ef8f1a3d87b

Observation 6927e377-6619-403c-9cde-3c0acd2c1728 · outbound

This paper cites ROCm: HIPStream,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm: HIPStream,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.679660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.720394Z digest=sha256:1a344cc8c107d8c0c54ad78ee9b572eeb3cf6fdfdb07cd236d7f2579080fe89e

Observation 5cdb2bd9-e3db-4143-9029-4dceaab0961b · outbound

This paper cites rocprof — ROC Profiler Documentation,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines rocprof — ROC Profiler Documentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.668818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.724462Z digest=sha256:431835b15c5ed2bbc00b00d64ef0fca5abc47dae2187b6be8471f767cc006ebf

Observation 2c930b0c-a81c-455d-aa56-f678e87621f8 · outbound

This paper cites AMD Instinct™ MI300X Accelerator Perfor- mance Validation Guide,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD Instinct™ MI300X Accelerator Perfor- mance Validation Guide,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.658083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.728339Z digest=sha256:bcee85feaf884b50ebf1debd04b1b627958a03f62b36821d1fa7131c084c6b39

Observation a8f051b4-d58a-4f55-9839-7cdb3ccaf05a · outbound

This paper cites ROCm: ROCR-Runtime,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm: ROCR-Runtime,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.647582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.732671Z digest=sha256:aaa7fbffdddc0b8e6fd921a68ab8e93d04bc6243d28428f35c24d44e8bf5c93b

Observation 19addbb6-178a-457f-b8f0-8dbceae946c7 · outbound

This paper cites T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.637138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.736987Z digest=sha256:a29b646b0eb216c36ff40154058e1382705e5d8e18afb26f523448ce284876dc

Observation f1ba0c67-19ca-41da-a15f-04e42cd467ac · outbound

This paper cites DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.741039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.741039Z digest=sha256:a6767c7a79e3f396b2920cdddd821bf4685364817e856052dd36d267573edb3a

Observation 918f3ca6-b710-4094-97a6-603550ca0fb0 · outbound

This paper cites TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.626380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.745563Z digest=sha256:cdbc5f5acc96c55c51ffd21587e57e8d9fdc13a5fbc7afd94307e53d21cccc35

Observation 79aa09f4-c0dc-4557-bc57-6e0ea243859b · outbound

This paper cites MSC- CLang: Microsoft Collective Communication Language,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines MSC- CLang: Microsoft Collective Communication Language,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.615472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.749301Z digest=sha256:d9b8e54b1c0f40f2e3a2db09bf9ee7680fa296320865de18009e1b21bbcbf5cb

Observation b6541107-eac4-40d7-9ec4-7267ff685389 · outbound

This paper cites Available: https://github.com/NVIDIA/nccl.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Available: https://github.com/NVIDIA/nccl

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.604358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.752792Z digest=sha256:e52e0f0fc834e7c9f3095299b8fe8045c78389fa754bd33ef5c262847ce20e6d

Observation d7372df3-5cb9-4ae8-828a-430e2ab53fb6 · outbound

This paper cites TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.593400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.756341Z digest=sha256:0fde067e81b191f73d864c44bb994a60b132ab1c9baa914cf865540e0d1f32e3

Observation 5c965cdf-de6a-4e2e-a15c-201a20225f60 · outbound

This paper cites MSCCL++: A GPU-driven communication stack for scalable AI applications.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines MSCCL++: A GPU-driven communication stack for scalable AI applications

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.581937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.795376Z digest=sha256:e3be51ed7fba97f99d1db0e8d320400908a139c7d1136865026b2ce32e8499e8

Observation 33a9f01b-36cd-481c-abf5-2b04b1c5de54 · outbound

This paper cites Deep Learning Recommendation Model for Personalization and Recommendation Systems.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Deep Learning Recommendation Model for Personalization and Recommendation Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.838572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.838572Z digest=sha256:1b36d18af73acca1370c4db7fffa26ec9b899bbe71260a30aa85e72dc72acc8e

Observation c3c9734d-8040-4f14-ac61-bcea09cbcb98 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.570814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.888958Z digest=sha256:b2ccff078a641209a846c25af8ff6232318238534f480b2340227a355c0fdd09

Observation 93de37ab-023d-4849-9713-109d1645eaee · outbound

This paper cites Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.917412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.917412Z digest=sha256:727017165cbc96d3b264f0bf61ff3b5629ef15006d0afa573d33e3ad90224a56

Observation febf812d-6ccd-40b8-b15f-3c2bfb32a186 · outbound

This paper cites The Case for GPGPU Spatial Multitasking,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The Case for GPGPU Spatial Multitasking,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.559159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.936810Z digest=sha256:ae9bf660d58bb58c526c957e37cca5726cafca18d4da6e7763df4fe6f98141c2

Observation f1ed9c9e-0e3f-42c4-9330-7a15f54fc25a · outbound

This paper cites Improving GPGPU energy-efficiency through concurrent kernel execution and DVFS,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Improving GPGPU energy-efficiency through concurrent kernel execution and DVFS,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.546408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.941351Z digest=sha256:0a705fe748f707cd6fde421dcc8f5d9dd8f8d3cacc18068ff694984f057f0baa

Observation 0de2b38d-dcb5-4768-81ea-ba7fe669f6cb · outbound

This paper cites Improving GPGPU Concurrency with Elastic Kernels,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Improving GPGPU Concurrency with Elastic Kernels,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.533927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.945764Z digest=sha256:b76cd529c41ebf6b0e3763dfc76e06376aca3f6048fb15bec257bb2ad39a4212

Observation 6ba035e1-6067-4d66-82c6-0c07a7857b84 · outbound

This paper cites Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.521377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.950054Z digest=sha256:479aab42176e59a2a0dd0eb4ad675c146006b0f58f8f203b7d5e991f406b6e41

Observation a3a604e8-475b-4bff-abd0-253409ab8303 · outbound

This paper cites Global Optimizations & Lightweight Dynamic Logic for Concurrency.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Global Optimizations & Lightweight Dynamic Logic for Concurrency

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:23:19.250779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.954391Z digest=sha256:98c996abf9c43df66fbcca1b8d37f618c1499216ab38cc5455833fdd8240f98c

Observation fb53f232-3849-4ca5-ae25-0ac468d29a7e · outbound

This paper cites An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collec- tives,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collec- tives,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.508949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.966166Z digest=sha256:a3464394c1cd6fba7f2e10553a493ba813cac7c041336a0160404c823600ec3d

Observation 1925a9da-25db-43c4-948d-e6befcaffe78 · outbound

This paper cites Introducing Async Tensor Parallelism in PyTorch,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Introducing Async Tensor Parallelism in PyTorch,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.496119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T12:23:18.970195Z digest=sha256:814e27b43e2131d725f19692b4b9914360bbebfb43fef9f395018b86c0e0443d

Observation 0a16a08d-0815-4c5e-98c6-19f098eae38b · outbound

This paper cites Optimizing Distributed ML Communication with Fused Computation-Collective Operations.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Optimizing Distributed ML Communication with Fused Computation-Collective Operations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.976636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.976636Z digest=sha256:10783c604e4a575b9098dc5b2f4e05dabcab4e6a515f29c43ad1b27f511264fd

Observation 25d52610-b9b8-4db3-917c-45ac4402b6f3 · outbound

This paper cites Breaking the Computation and Communication Abstraction Barrier in Distributed Machine Learning Workloads,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Breaking the Computation and Communication Abstraction Barrier in Distributed Machine Learning Workloads,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.988507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.988507Z digest=sha256:22dfb5b3df80727f7b31081e35fdd48cfa27b363c76bdb3aa1a568b5b3dfaf92

Observation 54b67e22-4c11-4532-a4be-47ce55e0e45e · outbound

This paper cites Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:19.004428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:19.004428Z digest=sha256:c3542308fdbf2da7c9897234bad13a0a94c88c2acd0d8f70969b922e678a500c

Observation babe03a7-517b-4b39-9cf5-f1d3e8580eb1 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.559962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.559962Z digest=sha256:70a58f6c8410c52a2449a2c78b5df2b4dd19bb02ef1cab65b5c05ecc285aa1ab

Pith citing papers

Observation a6d06ad2-aa27-4edc-8a66-4da6323f8d25 · inbound

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication cites this paper.

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication Optimizing ML Concurrent Computation and Communication with GPU DMA Engines

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:33.078708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T00:27:43.845408Z digest=sha256:c7aa4cead83f7c57e5b843412b77c04d8ffe62f92a0c7d2f5a41adafac913734