Pith. sign in

Paper Citation Record · LEDGER

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap

As of 9 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2512.10236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.10236 v3

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:17:35.344583Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved67
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b5929eb-a249-4f62-9e58-090973696279 · outbound

This paper cites Conccl: Optimizing ML concurrent computation and communication with GPU DMA engines,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Conccl: Optimizing ML concurrent computation and communication with GPU DMA engines,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:25.471057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:25.471057Z digest=sha256:95bd15b6261dded14a9b65f3a1fa6466299bcb39707659cb1d0fcabdab0b96c4

Observation f07a0ad4-d66e-420c-9af1-4712e0a6223e · outbound

This paper cites [Distributed GEMM: A novel CUTLASS-based implementation of Tensor Parallelism for NVLink-enabled systems,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap [Distributed GEMM: A novel CUTLASS-based implementation of Tensor Parallelism for NVLink-enabled systems,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:25.502992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:25.502992Z digest=sha256:f634e72dc153f2ddd31761c17d2f6f2bb5dd924f4a58904ac40e0e35916ee2b9

Observation 18e5e7a7-f1bd-46ac-afca-0966b0e75a2f · outbound

This paper cites (2023) Amd instinct™ mi300x accelerators.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap (2023) Amd instinct™ mi300x accelerators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:25.528116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:25.528116Z digest=sha256:dd2f2298ab5149797e93ee0efb1912ef69ba36a64461e1ae5bfa2e1bdec89cdb

Observation ab604659-6c91-4c41-8850-808e7a442ee6 · outbound

This paper cites HIP: C++ Heterogeneous-Compute Interface for Portability,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap HIP: C++ Heterogeneous-Compute Interface for Portability,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:25.636722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:25.636722Z digest=sha256:3e7d74c477d5d2a3802a01fbb2d4c54bb601b7e1fdbded94d2b2bca527430ab5

Observation 1c2d3159-5199-4556-a994-599e2fd816df · outbound

This paper cites ROCm Communication Collectives Library (RCCL),.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap ROCm Communication Collectives Library (RCCL),

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:25.804954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:25.804954Z digest=sha256:ea199a571edc724dc9270a349e04b377ec8c1af52a337245919e24eb856d9b5b

Observation 6020bb3a-af2c-42c6-ba41-1e9db5ad97e7 · outbound

This paper cites ROCm: HIPStream,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap ROCm: HIPStream,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:25.884659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:25.884659Z digest=sha256:d6aee3c0a7c34d72c5dee1cb52c6d37e8ebb640b002c65b0c2816e6e7d222f8c

Observation 63efbc59-c3e8-4f35-bbe1-b0663b9d1deb · outbound

This paper cites ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:26.005505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:26.005505Z digest=sha256:99db58293ed9d6b6cb2320f56f818e75bc2321d1d9e35f701f5c135d488fb1ed

Observation fba75d10-4eb1-4289-8956-92994a273211 · outbound

This paper cites (2025) Hip graphs.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap (2025) Hip graphs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:26.264006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:26.264006Z digest=sha256:c40c917c5f72aae677eff4c3613e49c8d222bf610c378fee59561a89ae4124e7

Observation cad949b9-fbba-4e12-834e-8b5bfee9bb83 · outbound

This paper cites (2025) hipblaslt.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap (2025) hipblaslt

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:26.379167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:26.379167Z digest=sha256:33bd5a72c4c64a2b902d187139d2b17357c1bec8a851cd55afbb89cb531ddc6f

Observation 36560a66-9cab-4c92-b8b5-e8ead0d488de · outbound

This paper cites FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:26.613570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:26.613570Z digest=sha256:ebffeafd4eafd927e4b2951fcc35f8f96c79429ae221fb23097a3ab297ac686e

Observation 4fe50f2a-2b6e-4187-b873-a50a69736e48 · outbound

This paper cites Centauri: Enabling efficient scheduling for communication- computation overlap in large model training via communication partitioning,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Centauri: Enabling efficient scheduling for communication- computation overlap in large model training via communication partitioning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:26.799537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:26.799537Z digest=sha256:cf56676a69f1fbbcc7103473f71c31910632e459a0ed74e15ad2250af2d5a1a4

Observation 007169a4-fc2a-41d4-8285-83331db20969 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Evaluating Large Language Models Trained on Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:27.038056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:27.038056Z digest=sha256:97d9bd9c12cc73ccc17e85bf492fe28e117c4db1ab4fa09fbbeec4556f9d4ba2

Observation c0cc879a-857b-4c85-acb6-4e1133c75e4d · outbound

This paper cites Revisiting scaling laws for language models: The role of data quality and training strategies,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Revisiting scaling laws for language models: The role of data quality and training strategies,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:27.194833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:27.194833Z digest=sha256:82570ee09566d8869b014d11a86d8236e1b08d1cd3df26cfd8b21e618bbba3d2

Observation 520d00d8-4d08-4ac7-8eec-e0e1a80c8a79 · outbound

This paper cites Concerto: Automatic communication optimization and scheduling for large-scale deep learning,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Concerto: Automatic communication optimization and scheduling for large-scale deep learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:27.296653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:27.296653Z digest=sha256:6c8450cbfae5932fe3136f90e397baa044db91d6c8344afd1f749875d0188f57

Observation ecfe41e1-9461-404b-9279-34c042912945 · outbound

This paper cites Scaling llama 3 training with efficient parallelism strategies,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Scaling llama 3 training with efficient parallelism strategies,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:27.429870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:27.429870Z digest=sha256:b77dc6aca47a3d357df7e0af714708f6bf6f09c9e5b29f3d84d0db88b99ac590

Observation 9811d8b7-41df-4da8-a041-b8ad608cf761 · outbound

This paper cites Transformations to parallel codes for communication-computation overlap,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Transformations to parallel codes for communication-computation overlap,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:27.653172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:27.653172Z digest=sha256:67d3a8cd5b5e2a31c006a3709ef63e395741909aab64543f20106994afb2ac15

Observation d3436e9a-4498-4e81-bcf5-0032a61cec4c · outbound

This paper cites Mpi-aware compiler optimizations for improving communication-computation overlap,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Mpi-aware compiler optimizations for improving communication-computation overlap,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:27.846495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:27.846495Z digest=sha256:7fac6bc79fb514e956b3fa29463e54f06801f9ecfe7d86b19b8cca8fdab43f94

Observation 4323b855-f1b8-4f81-b1e4-85ffb603c56c · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:27.946846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:27.946846Z digest=sha256:eb42501a4e8d9e06aff3c9311fa4dc8dba7fed450c98b55b42742d605dd585b6

Observation c68991d0-8eac-46a7-baed-56869f14f2e3 · outbound

This paper cites Deepseek-v3 technical report,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Deepseek-v3 technical report,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.126438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.126438Z digest=sha256:b3045ee113eef70e1f345344f63836122ea51520d10a67f74422935283aa6d70

Observation 666f6fc9-1976-4e72-998a-8143d8fdd541 · outbound

This paper cites The llama 3 herd of models,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap The llama 3 herd of models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.354666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.354666Z digest=sha256:055b540ffda21990efded172aa5af792ef7204cd781b5749bdb7e52f9a429faa

Observation 70a54b9f-f67f-4119-b29c-5b9ef0433d25 · outbound

This paper cites Compiler-assisted overlapping of communication and computation in mpi applications,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Compiler-assisted overlapping of communication and computation in mpi applications,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.544971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.544971Z digest=sha256:6e94df1a1aad81dd06fd805ab724c3a8a4f9d6ac5c0e51390ccd5b88c632f422

Observation 6d022239-46ad-47cc-91a5-c98b24a20687 · outbound

This paper cites Bandwidth characterization of deepspeed on distributed large language model training,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Bandwidth characterization of deepspeed on distributed large language model training,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.652192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.652192Z digest=sha256:d20d4b7216468af098ae9e8a11e2916407359bd7fe02e3b5f7f96775d103eebb

Observation 5ee7b059-9eea-40a1-aa49-c024e50c77d5 · outbound

This paper cites Efficient and adaptable overlapping for computation and communication via signaling and reordering,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Efficient and adaptable overlapping for computation and communication via signaling and reordering,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.773221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.773221Z digest=sha256:e71ffd07f9ad9c759a96e9374473e8b17bcb61605d26dbca7ac5825c8de64446

Observation 2a11c22a-edaa-4416-b3d4-10e5d72b5d13 · outbound

This paper cites [Distributed w/ TorchTitan] Introducing Async Tensor Parallelism in PyTorch,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap [Distributed w/ TorchTitan] Introducing Async Tensor Parallelism in PyTorch,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.897418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.897418Z digest=sha256:e750419c59c546a8804de7a56b94f221f7f0c3b62e5046e36c786d8003490595

Observation 1ebf80fd-0592-4f99-a222-0cce9d2d799c · outbound

This paper cites Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:29.056475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:29.056475Z digest=sha256:e10580d9669a7af7aca3285420880700b8c0473368ae5840da2d53d0c9bd992e

Observation d65ced37-f761-4438-aef1-b93ed96b4257 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Gpipe: Efficient training of giant neural networks using pipeline parallelism,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:29.440449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:29.440449Z digest=sha256:2e2ada707579af53f397cad5f5f569344750ab3c5af67e040575f8d4c0c3568c

Observation b7e1d71e-f259-4e1a-927d-0c2c9f3f1aeb · outbound

This paper cites A loop transformation algorithm for communication overlapping,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap A loop transformation algorithm for communication overlapping,

Reference 27

Resolution
verified exact
doi, observed 2026-08-03T17:18:24.297106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-03T17:17:29.614433Z digest=sha256:e91dcf4f1e9b962d0a2a490db5bd4d6da2e24190143d3ba4dd2279adae1bce81

Observation a190159f-1426-42ce-8af7-8d011f88fb58 · outbound

This paper cites Breaking the computation and communication abstraction barrier in distributed machine learning workloads,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Breaking the computation and communication abstraction barrier in distributed machine learning workloads,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:29.852284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:29.852284Z digest=sha256:f0c5f3232f46630f8c4089cae1d70c2cc77560b62c83d744f7dbf1a9a45443dd

Observation b1345a36-4767-4ab8-83c5-5bd73eec4c9c · outbound

This paper cites Available: https://arxiv.org/abs/2507.04786.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Available: https://arxiv.org/abs/2507.04786

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:29.279492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:29.279492Z digest=sha256:673d196e46d43f225b9c50fcd60c69cdb969b6a3b04924ec45b52e006aacd4ea

Observation 4c219591-1d15-4ce0-8b24-73f4195677c4 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap A Survey on Large Language Models for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:30.192100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:30.192100Z digest=sha256:251e385e8700ff6bf8dd0be41eac6c8b93530b71ebc100c84c5a97c98d20fd97

Observation 35986078-4f95-4553-9cae-73c0919d8f9a · outbound

This paper cites Reducing activation recomputation in large transformer models,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Reducing activation recomputation in large transformer models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:30.446070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:30.446070Z digest=sha256:f43f5f0d6a4a9609cdde5752a9f74eea32a0a632f7f2b2c8e003d1b7a149d41a

Observation 9c403f43-9525-42eb-9690-d3ca58d9cde9 · outbound

This paper cites Lit silicon: A case where thermal imbalance couples concurrent execution in multiple gpus,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Lit silicon: A case where thermal imbalance couples concurrent execution in multiple gpus,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:30.609498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:30.609498Z digest=sha256:937cb9e2082ed802992cd98e2f6b08550758d99d609587572f6b1a778e74b173

Observation dd51c416-29b7-44c5-985f-da686e4fc045 · outbound

This paper cites Mixtral of Experts.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Mixtral of Experts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:30.019371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:30.019371Z digest=sha256:8e2fd06af4ac05b205a728a2e87c27c5a926ab882c116528681749f8bc237a4f

Observation ba0b4df0-aedb-4035-9f02-df2167f12778 · outbound

This paper cites Pytorch 11 distributed: Experiences on accelerating data parallel training,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Pytorch 11 distributed: Experiences on accelerating data parallel training,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:31.048465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:31.048465Z digest=sha256:382dfe6ee3b66e100a799f9fba5a51dcb3115dd8ba3d2ffb4526f6bc50fa60f5

Observation cbf99ac6-1ccd-456d-a7a4-8afb7573b594 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap A Survey on Large Language Models for Code Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:30.342823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:30.342823Z digest=sha256:d2e0b3d4ecb864c0eeb827a9361b5aabcc4122866790d92684467962383b0b4f

Observation 101087c2-49b8-4ff1-b935-1f8046432efc · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:31.451618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:31.451618Z digest=sha256:da193f73d32ba5893fe0f8ddd4186d832a52972920f62c415150715c02b6cc4b

Observation a0455977-1bff-42a9-8448-1d2992354bd7 · outbound

This paper cites MLPerf Inference Results v5.0,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap MLPerf Inference Results v5.0,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:31.595099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:31.595099Z digest=sha256:0dd47071a101eab1c5e55ddd9d76ce5614e9f1e4bac45824409ccd255ba06c2d

Observation 1f74fec2-21cf-4e63-a31c-a1e027f276db · outbound

This paper cites Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:30.775766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:30.775766Z digest=sha256:03a3e87944fbea30e24186436b7010ddb47308e422c6fb9a66c58810ab35c53b

Observation 6730e8b1-c68d-4481-bcf3-95ac4da96a3d · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Gshard: Scaling giant models with conditional computation and automatic sharding,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:30.901701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:30.901701Z digest=sha256:bbe227c88ef1c8bf65a3d6cd71247050a0b9e2343fa22f7429b81151b7c98b0e

Observation 35790693-991d-4143-ad02-0f3e1e4ed316 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:32.059112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:32.059112Z digest=sha256:3635bf3cc1c91f8b405fbbef43fd1f86c59bf8e219851bcedf00fc792ac16520

Observation 78df471c-704c-43fc-9cff-2802f02b0c11 · outbound

This paper cites T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:32.345045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:32.345045Z digest=sha256:11fe5428448b134e2623231b89a92032ebc6a8399c0141a32fbc18baa9afee13

Observation 2c38d75e-b571-4dd8-8b1b-85a3f11921a8 · outbound

This paper cites Exact dependence analysis for increased communication overlap,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Exact dependence analysis for increased communication overlap,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:32.583547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:32.583547Z digest=sha256:a7bebce6b7c1f95e3198c8521311bcb910eef3dc0a0a2bc3dba0546abaaeb7c4

Observation 7e2b64c6-aba8-4f1e-9b47-ebbc64f35590 · outbound

This paper cites Automatic gen- eration of software pipelines for heterogeneous parallel systems,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Automatic gen- eration of software pipelines for heterogeneous parallel systems,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:32.835860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:32.835860Z digest=sha256:40aeb01694193fbffbe3d6aff3b1f99a2a3bcd9ab329be8480acc9725bbb6884

Observation 51b9d562-2af8-4d7f-a733-486723796a23 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:31.717025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:31.717025Z digest=sha256:70446e9547b7d9e88bb8ef19e90a2fd375e72a0c32e2fa36c212f6ff3dc1eb4a

Observation 29046abe-f2a1-4ce9-9571-a9d606a4e07d · outbound

This paper cites Stream-k: Work-centric parallel decomposition for dense matrix- matrix multiplication on the GPU,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Stream-k: Work-centric parallel decomposition for dense matrix- matrix multiplication on the GPU,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:31.902770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:31.902770Z digest=sha256:ba4c22f350ba056c8098aae1b437994cf6f75d54aadd2eb7ad7b0cd2802ea461

Observation 4759e32d-4889-486c-9acc-2fc064a47277 · outbound

This paper cites Enabling compute-communication overlap in distributed deep learning training platforms,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Enabling compute-communication overlap in distributed deep learning training platforms,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.295676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.295676Z digest=sha256:c92f99a742100fde3ea545d00f4a394c1abd31e734da1144aad0b8ea1ba0639b

Observation 0a51ad9f-d781-4808-a33e-993e35c3bac1 · outbound

This paper cites Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:32.211232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:32.211232Z digest=sha256:a57557bfc59f5b963c6ecb8c70a3aac6744cf6168bd676ff79819f75072b9d39

Observation 024e74a0-517f-4ac1-b941-a37a295e142b · outbound

This paper cites Slechta, N.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Slechta, N

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.649111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.649111Z digest=sha256:654988505331d75375696804e284f066250cb547e7561ef4589421103beb0758

Observation b8c98e3c-20bd-4dcf-b2e5-312232c0e7db · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Triton: an intermediate language and compiler for tiled neural network computations,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.749522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.749522Z digest=sha256:4a5d18d41b8f7a4203c7c0719d7d3d2005f06126be2dae100e67c68ef00746ae

Observation 1225fe07-a3f5-48c3-94a6-deb695ae2959 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap LLaMA: Open and Efficient Foundation Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.833361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.833361Z digest=sha256:8b974db55adc2ee6c1ec3a58c45091fc120925056952f87e6b96c70e880a9411

Observation 1f1f6275-507a-4394-a210-3b850abe2b1a · outbound

This paper cites Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.011909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.011909Z digest=sha256:c4048080128dff96750c607406da26839753dd7acd145243626bdcb359e6669b

Observation fe6e6e1c-1e64-4d4a-96e1-725d3449dad6 · outbound

This paper cites Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation AI scale,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation AI scale,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.183239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.183239Z digest=sha256:1869a9d6b79f6576cb3710ff2da0cca96a7f6e40ab73056c16a0bb19f796ee5f

Observation 85adbf3d-dedf-4211-80e7-af2315733160 · outbound

This paper cites Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-03T17:18:24.045766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-03T17:17:34.091313Z digest=sha256:11e02bcb2409d476e460825ee75e0eb35a73753599784f70f1698f2df5275168

Observation ca9c0139-2eda-49ab-8e73-ef219fd1186f · outbound

This paper cites A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.472027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.472027Z digest=sha256:12d24c8f83e1f023a16813e7a90ae437ead9fbc14b9330e96499877881fe7ca3

Observation 7f2c07d2-85d7-46f3-971c-d3445892bc89 · outbound

This paper cites Pytorch symmetricmemory: Harnessing nvlink programmability with ease,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Pytorch symmetricmemory: Harnessing nvlink programmability with ease,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:34.367896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:34.367896Z digest=sha256:98b916b35b63aad221028291c2aa4fb5e5e1a2d7cde6c09747642c0edd000e7f

Observation 69d62d58-031f-4986-8647-014bfaa6c838 · outbound

This paper cites Petuum: A new platform for distributed machine learning on big data,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Petuum: A new platform for distributed machine learning on big data,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:34.437355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:34.437355Z digest=sha256:f5e78816651b4c9d15677a48c54bcb90fbcbe0c01b28d4c9993e5089f7266b78

Observation 855e07e4-b108-4c24-86cd-a60d72f0414a · outbound

This paper cites Context Parallelism for Scalable Million-Token Inference.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Context Parallelism for Scalable Million-Token Inference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:34.555509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:34.555509Z digest=sha256:70156d89dddbd99b9ee442050387770ebdc9e5b7b15bdc9b7cf1611562ecf9a8

Observation 063917ac-cb0f-459a-b90d-9a7dd90d4a57 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.914217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.914217Z digest=sha256:412396b86a808c96d720c4a49b24a6463c2ce06320bec6e52e00c90bc004e806

Observation 94fe4ef9-6f68-4be2-9276-33b6f676953e · outbound

This paper cites Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:33.981248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:33.981248Z digest=sha256:ec888d4b13b821e428ef0d02e0b84388c9f002736553f6a372d66804b9273842

Observation 5f4c9bf8-c59f-4d2c-819f-3ee88f462063 · outbound

This paper cites Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:35.012041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:35.012041Z digest=sha256:853546b331f0445e3d77712fe62a7286531dcd8f933849e2cfafd9664730adb4

Observation 4144954e-45a9-4361-95a8-7eea2a9c58ac · outbound

This paper cites Overlap communication with dependent computation via decomposition in large deep learning models,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Overlap communication with dependent computation via decomposition in large deep learning models,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:34.315521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:34.315521Z digest=sha256:41cd4c198f538d8728f28a55c652e71e8f1c3b91a1b050c24eec50746fffd158

Observation b7dfa958-f4b1-49c2-bcd7-168b8b0d4c76 · outbound

This paper cites Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:34.668108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:34.668108Z digest=sha256:99a86962a4ac3129cb6ca3c2affcb693780409f6b7ef7f805998fd998b808d84

Observation eba6e6b8-f352-423f-9a99-c917c6ea7b77 · outbound

This paper cites Pytorch FSDP: experiences on scaling fully sharded data parallel,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Pytorch FSDP: experiences on scaling fully sharded data parallel,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:34.804923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:34.804923Z digest=sha256:68883be2efdf63cbb1e148d7dd30e741996c161abe7a263484dc4ec4e7fb3a55

Observation d6addf3b-2ebb-4f6b-ac3f-912ead8416bd · outbound

This paper cites Tilelink: Generating efficient compute-communication overlapping kernels using tile-centric primitives,.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Tilelink: Generating efficient compute-communication overlapping kernels using tile-centric primitives,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:35.176285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:35.176285Z digest=sha256:4a2657ba510274128cc78f76d70644b41b5882e668bb773c7db2ce6a972ab5b8

Observation e6fe2106-d242-46e5-a249-fe60a74d8e6c · outbound

This paper cites Available: https://openreview.net/forum?id=ccjvBkTRRe 13.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Available: https://openreview.net/forum?id=ccjvBkTRRe 13

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:35.344583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:35.344583Z digest=sha256:a09d3e93bf350e71419ff25e7297099f46d50093e7f7dddc3e2d84fb8b14910a

Observation 6cbc2a96-370e-4b9c-a697-2893791ccdb7 · outbound

This paper cites Available: http://papers.nips.cc/paper files/paper/2022/ hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Available: http://papers.nips.cc/paper files/paper/2022/ hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.040628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.040628Z digest=sha256:65c2d0e6c16d10e1f512d6ac1209e5ff0e0c92078e7b91ed79c8051eae02fe48

Observation 1b136b9c-064f-4acf-9c87-707cddfa654d · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:31.325255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:31.325255Z digest=sha256:a59a36300e5cdc23166b98aaf600206b5fc7a1b073a853ed2ef80ab57c9ac14e

Observation 9aa5a523-c969-4931-9369-4adea9d14bec · outbound

This paper cites The Llama 3 Herd of Models.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.465120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.465120Z digest=sha256:8d0d9e6fab8e4c0dd62125427b47928811ca39793458324adc2a66439b84fc71

Observation f92680db-59e5-4637-81ea-749be78a85ed · outbound

This paper cites DeepSeek-V3 Technical Report.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap DeepSeek-V3 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:28.237971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:28.237971Z digest=sha256:df55e904b19a819b1c4050545fe98870d06a8cd91bea81c3335217e16f313853

Pith citing papers

No inbound Pith citation observations are available.