Pith. sign in

REVIEW 3 major objections 5 minor 69 references

By sharding communication one level deeper than existing shard-based overlap, this paper shows that data-dependent GPU communication can be overlapped with computation in an all-to-all pattern, delivering up to 1.6x speedup and a heuristic

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 17:12 UTC pith:LXHGMB5Q

load-bearing objection Finer-grain decomposition for compute–communication overlap is a solid experimental idea with real measured speedups on MI300X, but the schedule-picking heuristic is under-validated and sloppily specified. the 3 major comments →

arxiv 2512.10236 v3 pith:LXHGMB5Q submitted 2025-12-11 cs.DC cs.ARcs.LG

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap

classification cs.DC cs.ARcs.LG
keywords compute-communication overlapfiner-grain shardingall-to-all communicationGPU DMA enginesGEMM decomposition inefficiencyschedule heuristicsarithmetic intensitydistributed ML training
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Distributed ML training and inference often stall on communication that sits on the critical path, leaving up to 1.7x ideal performance unused. The usual remedy, overlapping computation and communication at the granularity of shards, is too coarse: it uses peer-to-peer transfers that leave direct-connection GPU networks idle. This paper proposes FiCCO, which splits communication one level deeper, turning each transfer into an all-to-all pattern. On an eight-GPU system, FiCCO opens a design space of schedules, and a simple heuristic based on a GEMM's arithmetic intensity and memory traffic picks the right schedule in 81% of unseen cases, delivering up to 1.6x speedups.

Core claim

The paper's central claim is that finer-grain decomposition of data-dependent communication, one level below shard granularity, converts peer-to-peer transfers into all-to-all transfers that keep direct-connection GPU interconnects busy. The resulting overheads, decomposition and contention inefficiencies, can be characterized using two static GEMM properties: operations per byte (OTB) and memory traffic (MT). From this characterization the paper derives four concrete schedules and a heuristic that selects among them, and offloading communication to GPU DMA engines reduces contention further. This combination is what delivers the reported speedups.

What carries the argument

FiCCO (Finer-grain Compute-Communication Overlap): communication is re-sharded by the number of GPUs, so in an eight-GPU system each transfer is one-eighth the size of a shard-level transfer. This turns communication into an all-to-all pattern and enables a design space of schedules distinguished by computation uniformity (uniform vs. heterogeneous), computation granularity (fused vs. unfused GEMM kernels), and communication shape (1D vs. 2D). The selection heuristic uses a GEMM's OTB and MT, forms their product, compares it against a machine-level threshold with a 5x multiplier, and picks among the schedules.

Load-bearing premise

The heuristic assumes that a GEMM's static arithmetic intensity (OTB) and memory traffic (MT), combined as a product and compared with a hand-chosen 5x machine-level threshold, reliably predicts which FiCCO schedule minimizes total inefficiency on unseen workloads and hardware.

What would settle it

Run FiCCO on a ring or torus topology, or with GEMM shapes outside the 16 synthetic scenarios, and measure whether the heuristic still picks the winning schedule in roughly 81% of cases; if accuracy drops well below that, the static OTB*MT signature is not portable across topologies.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • FiCCO attains up to 1.6x speedup over serial execution across realistic GEMM and all-gather scenarios.
  • The proposed heuristic picks the optimal schedule in 81% of unseen synthetic scenarios, with mispredictions losing about 14% of speedup.
  • DMA-based communication offload reduces contention inefficiency versus GPU-core-driven communication in all measured cases.
  • Shard-based overlap can be slower than serial execution on direct-connection topologies, while FiCCO avoids that degradation.
  • The design-space framing gives frameworks and runtimes a concrete mechanism for choosing overlap schedules based on operation shapes.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The heuristic's 5x threshold and OTB*MT product are likely calibrated to the specific GPU and topology studied; on other hardware the demarcation may shift, so the heuristic may need recalibration rather than re-derivation.
  • If DMA engines gain support for 2D copies, the emulated 2D schedules should be re-measured; real 2D DMA may close part of the remaining gap to ideal speedup.
  • The same one-level-deeper decomposition idea could extend to reduce-scatter-based parallelism, such as tensor parallelism with gradients, once DMA engines support arithmetic operations, which the paper explicitly leaves out.
  • The correlation between static GEMM properties and inefficiency signatures suggests that a similar OTB/MT-based selector could be applied to other overlap schemes, not just the four schedules studied here.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes FiCCO, a finer-granularity compute-communication overlap scheme that decomposes communication one level deeper than shard-based overlap (e.g., transfer sizes one-eighth of shard-based in an 8-GPU system). This decomposition is argued to unlock all-to-all communication, better utilize direct-connection topologies such as AMD MI300X, and enable a richer space of execution schedules. The paper characterizes two inefficiencies—decomposition inefficiency loss (DIL) and contention inefficiency loss (CIL)—and maps them to static GEMM features (OTB and MT). It then defines four FiCCO schedules, proposes a heuristic based on OTB×MT and a 5× machine-dependent threshold to select among them, and offloads communication to GPU DMA engines. Experiments on 16 real GEMM scenarios report up to 1.6× speedup over serial execution, and an additional 16 synthetic scenarios are used to claim 81% heuristic accuracy.

Significance. If the heuristic generalizes, the paper offers a practical, software-only method for improving dependent compute-communication overlap on full-mesh GPU systems, a setting where existing shard-based overlap degrades. The strengths are the concrete measurements on a real MI300X system, the use of DMA offload to reduce contention, and the attempt to ground schedule selection in static operator properties. The core speedup result is plausible and independent of the heuristic. However, the heuristic's central claim of 81% accuracy on unseen scenarios rests on a small, author-generated validation set, a hand-picked threshold, and a unit-inconsistent decision variable. The 2D communication schedule is also emulated rather than measured. These issues limit the current evidence for transferability, which is the main practical contribution.

major comments (3)
  1. [V-C and VI-C] The heuristic's decision boundary is not well-founded. Section V-C defines combined OTB×MT for the GEMM and compares it to 'machine-level' OTB×MT defined as 'peak compute FLOPs' (FLOPs/s). This is a unit mismatch: OTB (FLOPs/byte) times MT (bytes) yields FLOPs, so the comparison implicitly imposes a 1-second timescale. The 5× threshold is hand-chosen and appears tuned to the 15 real scenarios in Table I, on which the heuristic then achieves 100% accuracy. The only out-of-sample evaluation is 16 synthetic scenarios with no cross-validation, no sensitivity analysis for the 5× value, and no confidence intervals. With one free parameter and 16 test points, 13/16 correct is weak evidence for the '81% of unseen scenarios' claim. Please provide a principled derivation of the threshold, report accuracy as a function of the threshold, or validate on independently collected workloads/hardware.
  2. [VI-B and Figure 12b] The 2D schedule is emulated using 1D memory copies of the same size because '2D memory copies with DMAs are not supported today.' The paper still reports 'with emulated 2D schedules we attain as high as 1.7× speedup' and includes uniform-fused-2D in the headline results and heuristic evaluation. A 1D copy does not reproduce the buffer layout, gather/scatter behavior, or link-level traffic pattern of a true 2D communication shape. This means the 2D arm of the design space is not actually evaluated, and the 1.7× number is a best-case estimate, not a measured result. Please either implement true 2D transfers, use a simulator validated against the 1D results, or clearly label these as estimates throughout the abstract, Section V, and Section VI.
  3. [IV-B and VI] The paper reports speedups and DIL/CIL values as point averages of 5 runs (after 10 warmups) but gives no error bars, standard deviations, or statistical significance tests. This matters because the heuristic is selecting among schedules whose speedups can be close (e.g., in Figure 12b the difference between schedules is often small relative to the reported 6% operator variation). Without variance information, the reader cannot tell whether the 81% accuracy or the 1.6× speedup is robust, or whether the heuristic's choices are within measurement noise. Please report error bars or variance for the headline numbers and for the per-schedule speedups in Figure 12b.
minor comments (5)
  1. [VI-C] The 16 synthetic scenarios are described only as 'wide ranging OTB and MT combinations.' Please specify how they were generated (parameter ranges, GEMM shapes, whether they include the M<K case), and list them so the out-of-sample claim is reproducible.
  2. [IV-C] The phrase 'op-to-byte' is used inconsistently; consider using 'OTB' consistently after first definition. Also, in Figure 7 the data labels are small and hard to read; enlarge them or report the values in a table.
  3. [VI-A] The text says 'we observe 7× communication slowdown' for shard-based overlap on MI300X, but Figure 13 shows speedup, not communication slowdown. Clarify where the 7× comes from or add a direct measurement.
  4. [IV-B.2] The omitted scenarios (e.g., tensor parallelism with reduce-scatter) are excluded because DMA engines do not support math today. This is a reasonable limitation, but it should be stated earlier and in the abstract or conclusion, since it bounds the applicability of FiCCO.
  5. [VII] The related work discussion is brief for a design-space paper. In particular, the comparison to Triton-Distributed is reported as out-of-memory; consider adding a small-scale comparison or at least a qualitative discussion of how FiCCO differs in kernel-authoring burden.

Circularity Check

0 steps flagged

No significant circularity: FiCCO speedups are measured and the heuristic, though in-sample validated, is not defined by the predicted schedule labels.

full rationale

The paper's central speedup claim is an empirical result: FiCCO is implemented and measured on MI300X against serial execution, shard-overlap, RCCL-based FiCCO, and an ideal roofline, with the up-to-1.6x speedup coming from measured execution times (Sections VI-B and VI-D). The DIL/CIL characterization is also measured, and the heuristic in Section V-C selects schedules from static GEMM features (M vs K, OTB, MT) using a decision rule; the rule is not defined in terms of the measured winning schedule, so no predicted quantity is equal to an input by construction. The 5x machine-level threshold is hand-picked rather than derived, and evaluating the heuristic on the same 15 real scenarios is in-sample validation—a generalization limitation, not a definitional circularity. The 16-scenario synthetic test provides a separate, though small, out-of-sample check. Self-citations such as ConCCL [1], T3 [42], and Tale of Two Cs [41] provide background on DMA offload and motivation, but the paper implements DMA via hipMemcpyDtoDAsync and measures its effect directly, so the citations are not load-bearing. No step in the derivation chain reduces a predicted result to its own inputs.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The paper is an empirical systems study, not a derivation. The central speedup is measured, but the heuristic's boundary (5x) is a hand-fitted constant; the representativeness of the 15 GEMMs and the 8-GPU MI300X topology are assumed; the synthetic validation set is generated by the authors.

free parameters (2)
  • 5x threshold in FiCCO heuristic = 5
    In Section V-C, scenarios with combined OTB*MT > 5× machine-level combined OTB*MT are assigned hetero-unfused-1D. The 5× multiplier is chosen to match the 15 measured scenarios; no independent derivation is provided, making it a fitted boundary.
  • OTB*MT product as combined static metric
    The heuristic combines GEMM op-to-byte and memory traffic into a single product (TS = OTB*MT in Figure 12a) without a derivation showing this product is the correct joint predictor of DIL/CIL; it is an ad hoc modeling choice.
axioms (4)
  • domain assumption 8-GPU MI300X full-mesh topology is representative of modern GPU systems
    All experiments use one 8-GPU MI300X platform with Infinity Fabric; the claimed advantage of all-to-all over P2P depends on this topology.
  • domain assumption DMA-based async memory copies can overlap with GEMM kernels and reduce interference
    Section IV-B1 uses hipMemcpyDtoDAsync to offload communication; the benefit of FiCCO depends on this capability inherited from prior ConCCL work [1].
  • domain assumption Static GEMM dimensions (M,N,K) determine OTB and MT, which predict DIL and CIL
    Section IV-C and IV-D establish correlations; the heuristic (Section V-C) relies on this as a predictive axiom.
  • domain assumption The 16 synthetic scenarios are representative of unseen real deployments
    Section VI-C validates the heuristic on synthetic scenarios; the claim of 81% accuracy assumes these synthetic cases resemble real workloads.

pith-pipeline@v1.3.0-alltime-deepseek · 24066 in / 11229 out tokens · 104723 ms · 2026-08-03T17:12:02.375967+00:00 · methodology

0 comments
read the original abstract

Modern ML workloads demand distributing training and inference across multiple GPUs. However, these parallelization techniques often suffer from exposed critical-path communication, leaving a potential 1.7x speedup on the table through compute-communication overlap. Prior overlapping methods harness the fact that ML model state and inputs are already sharded into the number of GPUs, and overlap the compute and communication at shard granularity. However, such coarse-grained overlap suffers from limited network topology support, and suboptimal dataflows. In this work, we instead make a case for finer-grain compute-communication overlap which we term FiCCO. FiCCO operates one level deeper than traditional sharding, and unlocks overlap for a wider set of network topologies and enables finer-grain dataflow. We show that FiCCO opens up a wider design space of execution schedules than possible at shard-level alone. To walk the design space of schedules, we study and characterize the performance inefficiencies on doing overlap and overlay the schedules with the associated inefficiency signatures. Our characterization reveals decomposition and contention based slowdowns to be the major performance limiters, and we correlate the slowdown factors with the static compute/communication operator sizes. This helps us design heuristics (that frameworks and runtimes can harness) to select bespoke FiCCO schedules based on the nature of underlying ML operations. Finally, to further minimize contention inefficiencies inherent with operation overlap, we offload communication to GPU DMA engines. We evaluate several scenarios from realistic ML deployments and demonstrate that our proposed heuristics driven bespoke schedules deliver up to 1.6x speedup. Further, our heuristics provide accurate guidance to pick the optimal schedule in 81% of unseen scenarios.

Figures

Figures reproduced from arXiv: 2512.10236 by Lizy K. John, Mahzabeen Islam, Shagnik Pal, Shaizeen Aga, Suchita Pati.

Figure 1
Figure 1. Figure 1: Speedup with finer-grain decomposition of data [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: (a) Sample inputs and ML model. (b) Data-parallelism (Fully-shared data-parallel - FSDP) and context parallelism [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: (a) Start state for weights (W) and inputs (I) on a four GPU system. (b) Baseline serial execution of communication [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Shared-based overlap versus FiCCO in action. [PITH_FULL_IMAGE:figures/full_fig_p003_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: Inefficiencies with operator decomposition and overlap. [PITH_FULL_IMAGE:figures/full_fig_p004_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Decomposition inefficiency loss (DIL) for GEMM with row (M) or column (K) sharding (8-way and 64-way). [PITH_FULL_IMAGE:figures/full_fig_p006_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Contention inefficiency loss (CIL) for GEMM (left) and all-gather communication (right). CIL for GEMM is reported [PITH_FULL_IMAGE:figures/full_fig_p006_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Decomposition inefficiency loss for DMA all-gather. [PITH_FULL_IMAGE:figures/full_fig_p006_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Proportion of DIL versus CIL for GEMMs and all-gather. [PITH_FULL_IMAGE:figures/full_fig_p007_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: (a) FiCCO design space resulting in eight possible schedules. (b) FiCCO schedules under consideration. [PITH_FULL_IMAGE:figures/full_fig_p007_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: (a) FiCCO heuristics – GEMM op-to-byte/OTB, GEMM memory traffic/MT). (b) FiCCO schedules performance. Decoding the provided schedules in a comparative fash￾ion, we observe that all schedules communicate the same effective buffer size and that uniform-fused-2D com￾municates 2D buffers. Next, to ensure uniformity in GEMM sizes, all uniform schedules incur gather of local and remote received buffers. Next, s… view at source ↗
Figure 13
Figure 13. Figure 13: Deficiencies of shard-based overlap. For ideal, we observe a bell curve in relation to relative GEMM/communication time (along x-axis) as expected. That is, the more balanced GEMM and communication times are the higher the benefit of perfectly overlapping them. 8 [PITH_FULL_IMAGE:figures/full_fig_p008_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Comparing FiCCO to other techniques. For shard-based overlap however, with peer-to-peer com￾munication which fails to utilize available network links in AMD Instinct™ MI300X (we observe 7× communication slowdown), we observe negative correlation between speedup and GEMM/communication time ratio. This is so as going from left-to-right longer GEMM executions hide communica￾tion inefficiencies in shard-overl… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 2 canonical work pages · 1 internal anchor

  1. [1]

    Conccl: Optimizing ML concurrent computation and communication with GPU DMA engines,

    A. Agrawal, S. Aga, S. Pati, and M. Islam, “Conccl: Optimizing ML concurrent computation and communication with GPU DMA engines,” inIEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2025, Ghent, Belgium, May 11-13, 2025. IEEE, 2025, pp. 1–11. [Online]. Available: https: //doi.org/10.1109/ISPASS64960.2025.00018

  2. [2]

    [Distributed GEMM: A novel CUTLASS-based implementation of Tensor Parallelism for NVLink-enabled systems,

    Ali Hassani, Michael Isaev, Nic McDonald, Jie Ren, Vijay Thakkar, Haicheng Wu, and Humphrey Shi, “[Distributed GEMM: A novel CUTLASS-based implementation of Tensor Parallelism for NVLink-enabled systems,” https://blog.shi-labs. com/distributed-gemm-88be6a481e2b, December 2024

  3. [3]

    (2023) Amd instinct™ mi300x accelerators

    AMD. (2023) Amd instinct™ mi300x accelerators. [On- line]. Available: https://www.amd.com/en/products/accelerators/instinct/ mi300/mi300x.html

  4. [4]

    HIP: C++ Heterogeneous-Compute Interface for Portability,

    AMD, “HIP: C++ Heterogeneous-Compute Interface for Portability,” https://github.com/ROCm/HIP, 2024

  5. [5]

    ROCm Communication Collectives Library (RCCL),

    ——, “ROCm Communication Collectives Library (RCCL),” https: //github.com/ROCm/rccl, 2024

  6. [6]

    ROCm: HIPStream,

    ——, “ROCm: HIPStream,” https://rocm.docs.amd.com/projects/HIP/ en/latest/reference/hip runtime api/modules/stream management.html, 2024

  7. [7]

    ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,

    ——, “ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,” https://github.com/ROCm/rocBLAS, 2024

  8. [8]

    (2025) Hip graphs

    AMD. (2025) Hip graphs. [Online]. Avail- able: https://rocm.docs.amd.com/projects/HIP/en/docs-develop/how-to/ hip runtime api/hipgraph.html

  9. [9]

    (2025) hipblaslt

    ——. (2025) hipblaslt. [Online]. Available: https://github.com/ROCm/ rocm-libraries

  10. [10]

    FLUX: fast software-based communication overlap on gpus through kernel fusion,

    L. Chang, W. Bao, Q. Hou, C. Jiang, N. Zheng, Y . Zhong, X. Zhang, Z. Song, Z. Jiang, H. Lin, X. Jin, and X. Liu, “FLUX: fast software-based communication overlap on gpus through kernel fusion,”CoRR, vol. abs/2406.06858, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2406.06858

  11. [11]

    Centauri: Enabling efficient scheduling for communication- computation overlap in large model training via communication partitioning,

    C. Chen, X. Li, Q. Zhu, J. Duan, P. Sun, X. Zhang, and C. Yang, “Centauri: Enabling efficient scheduling for communication- computation overlap in large model training via communication partitioning,” inProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’24). New York, NY ,...

  12. [12]

    Evaluating large language models trained on code,

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Her...

  13. [13]

    Revisiting scaling laws for language models: The role of data quality and training strategies,

    Z. Chen, S. Wang, T. Xiao, Y . Wang, S. Chen, X. Cai, J. He, and J. Wang, “Revisiting scaling laws for language models: The role of data quality and training strategies,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austr...

  14. [14]

    Concerto: Automatic communication optimization and scheduling for large-scale deep learning,

    S. Cheng, S. Lin, L. Diao, H. Wu, S. Wang, C. Si, Z. Liu, X. Zhao, J. Du, W. Lin, and Y . You, “Concerto: Automatic communication optimization and scheduling for large-scale deep learning,” inProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS 2025, Rotterdam, The ...

  15. [15]

    Scaling llama 3 training with efficient parallelism strategies,

    W. Chu, X. Xie, J. Yu, J. Wang, A. Phanishayee, C. Tang, Y . Hao, J. Huang, M. Ozdal, J. Wang, V . Goswami, N. Goyal, A. Kadian, A. Gu, C. Cai, F. Tian, X. Wang, M. Si, P. Balaji, C.-H. Chu, and J. Park, “Scaling llama 3 training with efficient parallelism strategies,” inProceedings of the 52nd Annual International Symposium on Computer Architecture, ser....

  16. [16]

    Transformations to parallel codes for communication-computation overlap,

    A. Danalis, K.-Y . Kim, L. Pollock, and M. Swany, “Transformations to parallel codes for communication-computation overlap,” inSC ’05: Proceedings of the 2005 ACM/IEEE Conference on Supercomputing, 2005, pp. 58–58

  17. [17]

    Mpi-aware compiler optimizations for improving communication-computation overlap,

    A. Danalis, L. Pollock, M. Swany, and J. Cavazos, “Mpi-aware compiler optimizations for improving communication-computation overlap,” inProceedings of the 23rd International Conference on Supercomputing, ser. ICS ’09. New York, NY , USA: Association for Computing Machinery, 2009, p. 316–325. [Online]. Available: https://doi.org/10.1145/1542275.1542321

  18. [18]

    Flashattention: Fast and memory-efficient exact attention with io-awareness,

    T. Dao, D. Y . Fu, S. Ermon, A. Rudra, and C. R ´e, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” inAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D....

  19. [19]

    Deepseek-v3 technical report,

    DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, J...

  20. [20]

    The llama 3 herd of models,

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al- Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Ma...

  21. [21]

    Compiler-assisted overlapping of communication and computation in mpi applications,

    J. Guo, Q. Yi, J. Meng, J. Zhang, and P. Balaji, “Compiler-assisted overlapping of communication and computation in mpi applications,” in 2016 IEEE International Conference on Cluster Computing (CLUSTER), 2016, pp. 60–69

  22. [22]

    Bandwidth characterization of deepspeed on distributed large language model training,

    B. Hanindhito, B. Patel, and L. K. John, “Bandwidth characterization of deepspeed on distributed large language model training,” in IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2024, Indianapolis, IN, USA, May 5- 7, 2024. IEEE, 2024, pp. 241–256. [Online]. Available: https: //doi.org/10.1109/ISPASS61541.2024.00031

  23. [23]

    Efficient and adaptable overlapping for computation and communication via signaling and reordering,

    K. Hong, X. Li, M. Liu, Q. Mao, T. Wu, Z. Huang, L. Chen, Z. Wang, Y . Zhang, Z. Zhu, G. Dai, and Y . Wang, “Efficient and adaptable overlapping for computation and communication via signaling and reordering,” 2025

  24. [24]

    [Distributed w/ TorchTitan] Introducing Async Tensor Parallelism in PyTorch,

    Horace He, Less Wright, Luca Wehrstedt, Tianyu Liu, Wanchao Liang, “[Distributed w/ TorchTitan] Introducing Async Tensor Parallelism in PyTorch,” https://discuss.pytorch.org/t/ distributed-w-torchtitan-introducing-async-tensor-parallelism-in-pytorch/ 209487, September 2024

  25. [25]

    Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms,

    Z. Hu, S. Shen, T. Bonato, S. Jeaugey, C. Alexander, E. Spada, J. Dinan, J. Hammond, and T. Hoefler, “Demystifying nccl: An in-depth analysis of gpu communication protocols and algorithms,”

  26. [26]

    Gpipe: Efficient training of giant neural networks using pipeline parallelism,

    Y . Huang, Y . Cheng, A. Bapna, O. Firat, D. Chen, M. X. Chen, H. Lee, J. Ngiam, Q. V . Le, Y . Wu, and Z. Chen, “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” inAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouv...

  27. [27]

    A loop transformation algorithm for communication overlapping,

    K. Ishizaki, H. Komatsu, and T. Nakatani, “A loop transformation algorithm for communication overlapping,”Int. J. Parallel Program., vol. 28, no. 2, p. 135–154, Apr. 2000. [Online]. Available: https://doi.org/10.1023/A:1007554715418

  28. [28]

    Breaking the computation and communication abstraction barrier in distributed machine learning workloads,

    A. Jangda, J. Huang, G. Liu, A. H. N. Sabet, S. Maleki, Y . Miao, M. Musuvathi, T. Mytkowicz, and O. Saarikivi, “Breaking the computation and communication abstraction barrier in distributed machine learning workloads,” inProceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, ser. ASP...

  29. [29]

    Available: https://arxiv.org/abs/2507.04786

    [Online]. Available: https://arxiv.org/abs/2507.04786

  30. [30]

    A survey on large language models for code generation,

    J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim, “A survey on large language models for code generation,”CoRR, vol. abs/2406.00515,

  31. [31]

    Reducing activation recomputation in large transformer models,

    V . A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” inProceedings of the Sixth Conference on Machine Learning and Systems, MLSys 2023, Miami, FL, USA, June 4-8, 2023, D. Song, M. Carbin, and T. Chen, Eds. mlsys.org, 2023. [On- line]. Available: https:...

  32. [32]

    Lit silicon: A case where thermal imbalance couples concurrent execution in multiple gpus,

    M. Kurzynski, S. Aga, and D. Wu, “Lit silicon: A case where thermal imbalance couples concurrent execution in multiple gpus,”

  33. [33]

    Mixtral of experts,

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de Las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mixtral of experts,”CoRR, vol. abs...

  34. [34]

    Pytorch 11 distributed: Experiences on accelerating data parallel training,

    S. Li, Y . Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania, and S. Chintala, “Pytorch 11 distributed: Experiences on accelerating data parallel training,”Proc. VLDB Endow., vol. 13, no. 12, pp. 3005–3018, 2020. [Online]. Available: http://www.vldb.org/pvldb/vol13/p3005-li.pdf

  35. [35]

    Available: https://doi.org/10.48550/arXiv.2406.00515

    [Online]. Available: https://doi.org/10.48550/arXiv.2406.00515

  36. [36]

    Ring attention with blockwise transformers for near-infinite context,

    ——, “Ring attention with blockwise transformers for near-infinite context,” 2023. [Online]. Available: https://arxiv.org/abs/2310.01889

  37. [37]

    MLPerf Inference Results v5.0,

    MLCommons, “MLPerf Inference Results v5.0,” https://github.com/ mlcommons/inference results v5.0, 2025, accessed: 2025-11-16

  38. [38]

    Available: https://arxiv.org/abs/2511.09861

    [Online]. Available: https://arxiv.org/abs/2511.09861

  39. [39]

    Gshard: Scaling giant models with conditional computation and automatic sharding,

    D. Lepikhin, H. Lee, Y . Xu, D. Chen, O. Firat, Y . Huang, M. Krikun, N. Shazeer, and Z. Chen, “Gshard: Scaling giant models with conditional computation and automatic sharding,” in9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. [Online]. Available: https://openreview.net/fo...

  40. [40]

    Pytorch: An imperative style, high-performance deep learning library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” 2019. [Online]. Available: https://arxiv.org...

  41. [42]

    T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,

    ——, “T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,” inProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ser. ASPLOS ’24. New York, NY , USA: Association for Computing Machinery, 2024, p. 1146–1164. [Online]. Available: https://...

  42. [43]

    Exact dependence analysis for increased communication overlap,

    S. Pellegrini, T. Hoefler, and T. Fahringer, “Exact dependence analysis for increased communication overlap,” inRecent Advances in the Message Passing Interface, J. L. Tr ¨aff, S. Benkner, and J. J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, pp. 89–99

  43. [44]

    Automatic gen- eration of software pipelines for heterogeneous parallel systems,

    J. A. Pienaar, S. Chakradhar, and A. Raghunathan, “Automatic gen- eration of software pipelines for heterogeneous parallel systems,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, ser. SC ’12. Washington, DC, USA: IEEE Computer Society Press, 2012

  44. [45]

    Efficient large-scale language model training on gpu clusters using megatron-lm,

    D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V . Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia, “Efficient large-scale language model training on gpu clusters using megatron-lm,” inProceedings of the International Conference for High Performance Computing, Networking, Storage and Ana...

  45. [46]

    Stream-k: Work-centric parallel decomposition for dense matrix- matrix multiplication on the GPU,

    M. Osama, D. Merrill, C. Cecka, M. Garland, and J. D. Owens, “Stream-k: Work-centric parallel decomposition for dense matrix- matrix multiplication on the GPU,” inProceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, PPoPP 2023, Montreal, QC, Canada, 25 February 2023 - 1 March 2023, M. M. Dehnavi, M. Kulk...

  46. [47]

    Enabling compute-communication overlap in distributed deep learning training platforms,

    S. Rashidi, M. Denton, S. Sridharan, S. Srinivasan, A. Suresh, J. Nie, and T. Krishna, “Enabling compute-communication overlap in distributed deep learning training platforms,” in48th ACM/IEEE Annual International Symposium on Computer Architecture, ISCA 2021, Virtual Event / Valencia, Spain, June 14-18, 2021. IEEE, 2021, pp. 540–553. [Online]. Available:...

  47. [48]

    Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,

    S. Pati, S. Aga, M. Islam, N. Jayasena, and M. D. Sinclair, “Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,” inIEEE International Symposium on Workload Characterization, IISWC 2023, Ghent, Belgium, October 1-3, 2023. IEEE, 2023, pp. 140–153. [Online]. Available: https://doi.org/10.1109/IISWC59245.2023.00026

  48. [49]

    Slechta, N

    B. Slechta, N. Comly, A. Eassa, J. DeLaere, and S. Raj. (2024, Aug) Nvidia nvlink and nvidia nvswitch supercharge large language model inference. NVIDIA Technical Blog. [Online]. Available: https://developer.nvidia.com/blog/ nvidia-nvlink-and-nvidia-nvswitch-supercharge-large-language-model-inference/ #nvswitch is critical for fast multi-gpu llm inference

  49. [50]

    Triton: an intermediate language and compiler for tiled neural network computations,

    P. Tillet, H. Kung, and D. D. Cox, “Triton: an intermediate language and compiler for tiled neural network computations,” in Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, MAPL@PLDI 2019, Phoenix, AZ, USA, June 22, 2019, T. Mattson, A. Muzahid, and A. Solar-Lezama, Eds. ACM, 2019, pp. 10–19. [Onlin...

  50. [51]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023. [Online]. Available: https: //arxiv.org/abs/2302.13971

  51. [52]

    Optimizing distributed ML communication with fused computation-collective operations,

    K. Punniyamurthy, K. Hamidouche, and B. M. Beckmann, “Optimizing distributed ML communication with fused computation-collective operations,” inProceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC 2024, Atlanta, GA, USA, November 17-22, 2024. IEEE, 2024, p. 88. [Online]. Available: https://dl.acm...

  52. [53]

    Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation AI scale,

    S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y . Aminabadi, A. A. Awan, J. Rasley, and Y . He, “Deepspeed-moe: Advancing mixture-of- experts inference and training to power next-generation AI scale,” in International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, ser. Proceedings of Machine Learning Research, K. Chaud...

  53. [54]

    Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping

    G. Wang, C. Zhang, Z. Shen, A. Li, and O. Ruwase, “Domino: Eliminating communication in LLM training via generic tensor slicing and overlapping,”CoRR, vol. abs/2409.15241, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2409.15241

  54. [55]

    A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,

    S. Singh, O. Ruwase, A. A. Awan, S. Rajbhandari, Y . He, and A. Bhatele, “A hybrid tensor-expert-data parallelism approach to optimize mixture-of-experts training,” inProceedings of the 37th International Conference on Supercomputing, ICS 2023, Orlando, FL, USA, June 21-23, 2023, K. A. Gallivan, E. Gallopoulos, D. S. Nikolopoulos, and R. Beivide, Eds. ACM...

  55. [56]

    Pytorch symmetricmemory: Harnessing nvlink programmability with ease,

    Y . Wang, H. He, and L. Wehrstedt, “Pytorch symmetricmemory: Harnessing nvlink programmability with ease,” Feb 2025, pyTorch Developer Forum. [Online]. Available: https://dev-discuss.pytorch.org/t/ pytorch-symmetricmemory-harnessing-nvlink-programmability-with-ease/ 2798

  56. [57]

    Petuum: A new platform for distributed machine learning on big data,

    E. P. Xing, Q. Ho, W. Dai, J. K. Kim, J. Wei, S. Lee, X. Zheng, P. Xie, A. Kumar, and Y . Yu, “Petuum: A new platform for distributed machine learning on big data,”IEEE Trans. Big Data, vol. 1, no. 2, pp. 49–67, 2015. [Online]. Available: https://doi.org/10.1109/TBDATA.2015.2472014

  57. [58]

    Context parallelism for scalable million-token inference,

    A. Yang, J. Yang, A. Ibrahim, X. Xie, B. Tang, G. Sizov, J. Reizenstein, J. Park, and J. Huang, “Context parallelism for scalable million-token inference,”CoRR, vol. abs/2411.01783, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2411.01783 12

  58. [59]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V . Kerkez, M. Khabsa, I. Kloumann, A. Koren...

  59. [60]

    Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,

    P. Vaithilingam, T. Zhang, and E. L. Glassman, “Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models,” inCHI ’22: CHI Conference on Human Factors in Computing Systems, New Orleans, LA, USA, 29 April 2022 - 5 May 2022, Extended Abstracts, S. D. J. Barbosa, C. Lampe, C. Appert, and D. A. Shamma, Eds....

  60. [61]

    Triton-distributed: Programming overlapping kernels on distributed ai systems with the triton compiler,

    S. Zheng, W. Bao, Q. Hou, X. Zheng, J. Fang, C. Huang, T. Li, H. Duanmu, R. Chen, R. Xu, Y . Guo, N. Zheng, Z. Jiang, X. Di, D. Wang, J. Ye, H. Lin, L.-W. Chang, L. Lu, Y . Liang, J. Zhai, and X. Liu, “Triton-distributed: Programming overlapping kernels on distributed ai systems with the triton compiler,” 2025. [Online]. Available: https://arxiv.org/abs/2...

  61. [62]

    Overlap communication with dependent computation via decomposition in large deep learning models,

    S. Wang, J. Wei, A. Sabne, A. Davis, B. Ilbeyi, B. Hechtman, D. Chen, K. S. Murthy, M. Maggioni, Q. Zhang, S. Kumar, T. Guo, Y . Xu, and Z. Zhou, “Overlap communication with dependent computation via decomposition in large deep learning models,” inProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and O...

  62. [66]

    Comet: Fine-grained computation-communication overlapping for mixture- of-experts,

    S. Zhang, N. Zheng, H. Lin, Z. Jiang, W. Bao, C. Jiang, Q. Hou, W. Cui, S. Zheng, L. Chang, Q. Chen, and X. Liu, “Comet: Fine-grained computation-communication overlapping for mixture- of-experts,”CoRR, vol. abs/2502.19811, 2025. [Online]. Available: https://doi.org/10.48550/arXiv.2502.19811

  63. [67]

    Pytorch FSDP: experiences on scaling fully sharded data parallel,

    Y . Zhao, A. Gu, R. Varma, L. Luo, C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y . Hao, A. Mathews, and S. Li, “Pytorch FSDP: experiences on scaling fully sharded data parallel,” Proc. VLDB Endow., vol. 16, no. 12, pp. 3848–3860, 2023. [Online]. Available: https://www.vldb....

  64. [69]

    Tilelink: Generating efficient compute-communication overlapping kernels using tile-centric primitives,

    S. Zheng, J. Fang, X. Zheng, Q. Hou, W. Bao, N. Zheng, Z. Jiang, D. Wang, J. Ye, H. Lin, L.-W. Chang, and X. Liu, “Tilelink: Generating efficient compute-communication overlapping kernels using tile-centric primitives,” inEighth Conference on Machine Learning and Systems,

  65. [70]

    Available: https://openreview.net/forum?id=ccjvBkTRRe 13

    [Online]. Available: https://openreview.net/forum?id=ccjvBkTRRe 13

  66. [2022]

    Available: http://papers.nips.cc/paper files/paper/2022/ hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html

    [Online]. Available: http://papers.nips.cc/paper files/paper/2022/ hash/67d57c32e20fd0a7a302cb81d36e40d5-Abstract-Conference.html

  67. [2023]

    Available: https://doi.org/10.48550/arXiv.2310.01889

    [Online]. Available: https://doi.org/10.48550/arXiv.2310.01889

  68. [2024]

    Available: https://arxiv.org/abs/2407.21783

    [Online]. Available: https://arxiv.org/abs/2407.21783

  69. [2025]

    Available: https://arxiv.org/abs/2412.19437

    [Online]. Available: https://arxiv.org/abs/2412.19437