REVIEW 4 major objections 4 minor 45 references
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read HelixPipe claims that splitting each transformer layer into parameterized and attention parts and scheduling attention across pipeline stages in parallel removes the dominant quadratic cost from pipeline idle time, yielding up to 26%…
desk verdict A genuinely new scheduling idea—attention parallel partition—with real speedups on fast interconnects, but the headline result rests on an overlap condition the paper itself shows fails on A800 at 32k, and the quantitative analysis is asserted rather than derived. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the attention parallel partition: each transformer layer is divided into pre-attention, attention, and post-attention, the two parameterized parts are assigned to a home pipeline stage, and the non-parameterized attention for different micro-batches rotates across stages so it executes in parallel. The two-fold FILO micro-batch schedule runs two micro-batches per step rather than one, doubling the bubble relative to the naive FILO schedule but enabling the roughly $2bsh$ per-layer p2p communication to be hidden behind attention computation. Recomputation without attention stashes only the attention inputs and residuals, recomputing pre-attention and post-attention activations before the backward pass, which cuts activation memory by a factor of four at the cost of one third more pipeline bubble time. Chunked MLP splits the MLP forward and backward into smaller pieces to prevent memory fragmentation caused by the combination of long sequences, recomputation, and the two-fold schedule.
What would settle it
Run the same 7B/128k configuration on a cluster where per-layer attention time is shorter than per-layer p2p communication time, as the paper itself observes on A800 GPUs at 32k sequence length, and measure effective throughput against 1F1B: if HelixPipe no longer beats the baseline, the overlap assumption is violated. A direct microbenchmark would compare the critical-path time with and without the two-fold overlap to see whether communication adds to the pipeline.
Extended reading notes
Core claim
HelixPipe's central claim is that a transformer layer can be partitioned into a parameterized pre-attention part, a non-parameterized attention part, and a parameterized post-attention part, and the attention part can be scheduled independently of layer ownership. The attention of micro-batch $i$ of layer $l$ runs on stage $((l+i+1) \bmod p)$, so multiple attention computations proceed in parallel across pipeline stages while the parameterized pieces stay on their assigned stages. The two-fold FILO schedule then executes two micro-batches at a time so that the communication of one overlaps with the computation of the other. The paper derives that this removes attention from the pipeline bubble: with recomputation without attention, the bubble becomes $8(p-1)(t_{\text{pre}}+t_{\text{post}})$ instead of $3(p-1)(t_{\text{pre}}+t_{\text{attn}}+t_{\text{post}})$ for 1F1B, while activation memory drops to $4bshmL/p$ when combining the FILO schedule with recomputation without attention.
Load-bearing premise
The load-bearing premise is that the extra inter-stage communication, about $2bsh$ per layer per micro-batch, can be completely hidden behind attention computation by the two-fold FILO schedule; on hardware where attention is faster than the interconnect, the overlap fails and the speedup disappears.
Editorial extensions
If this is right
- As sequence lengths grow, attention occupies a larger fraction of layer execution time, so removing attention from the pipeline bubble gives HelixPipe an increasing throughput advantage over layer-wise pipelines.
- At a fixed global token budget per iteration, HelixPipe saturates the pipeline with fewer micro-batches than 1F1B or ZB1P, which matters for long-sequence training where batch size is constrained.
- HelixPipe balances activation memory across pipeline stages instead of concentrating it in early stages, enabling longer sequences under a fixed per-GPU memory budget.
- The method is orthogonal to intra-layer sequence parallelisms such as Megatron sequence parallelism, Ulysses, ring attention, and context parallelism, so it can be combined with them to extend sequence length at both the intra-layer and inter-layer levels.
- On clusters with fast interconnects, the two-fold schedule is claimed to hide communication behind attention entirely, allowing HelixPipe to scale to any pipeline size; on slower interconnects or with faster attention kernels the overlap degrades and the speedup shrinks.
Reading between the lines
- The helix mapping effectively turns attention into a parameter-free, data-parallel workload; a natural extension is to apply the same scheduling idea to other non-parameterized operators in a transformer, such as activation functions or normalization-free residual paths, though the paper does not propose this.
- The requirement that the number of micro-batches be divisible by the pipeline size, doubled for the two-fold schedule, may constrain small-batch fine-tuning regimes; a testable extension is measuring HelixPipe's throughput when this divisibility does not hold.
- The paper's overlap analysis assumes NCCL communication consumes only a few SMs and leaves computation largely undisturbed; on GPUs or runtime configurations where communication kernels occupy more SMs, the overlap benefit could be smaller than measured.
- The recomputation-without-attention strategy costs up to 20% throughput at short sequence lengths, so a hybrid policy that disables recomputation when pre- and post-attention time is nontrivial would likely extend HelixPipe's advantage to shorter contexts.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes HelixPipe, a pipeline parallelism method for training long-sequence transformers. HelixPipe partitions each transformer layer into pre-attention, attention, and post-attention components, assigns attention computations of different micro-batches to different pipeline stages in a rotating pattern, and uses a two-fold first-in-last-out (FILO) schedule to balance memory and overlap communication with computation. It also contributes recomputation without attention and chunked MLP to reduce memory. The paper gives quantitative formulas for pipeline bubble time and activation memory, and reports experiments on H20 and A800 clusters with 1.3B, 3B, and 7B models at sequence lengths from 32k to 128k, claiming throughput improvements over 1F1B, ZB1P, and AdaPipe, including a 26% speedup for a 7B model at 128k sequence length on 64 H20 GPUs.
Significance. If the results hold, HelixPipe is a meaningful advance in pipeline parallelism for long-context transformer training. The idea of removing the quadratic attention computation from pipeline bubbles by scheduling attention across stages is clever and well motivated. The paper ships code, reports consistent speedups on long sequences and fast interconnects, and includes memory measurements that support the qualitative claims. The main significance is moderated by three conditions: the quantitative analysis in Section 4.5 is asserted rather than derived, the communication-overlap assumption is shown in Section 5.3 to fail on A800 at 32k, and the closest inter-layer long-context baseline, WeiPipe, is not evaluated. These issues affect the strength of the central claims but are addressable in revision.
major comments (4)
- [Section 4.5, Table 2] The HelixPipe pipeline bubble formula in Table 2 is written as 8(p-1)(t_pre+t_post) with no L/p factor, whereas the 1F1B and ZB1P rows include L/p. If t_pre and t_post are per-layer times as used in Eqs. (1) and (3), the per-stage bubble in HelixPipe should also scale with the number of layer components assigned to a stage; otherwise the notation should be redefined explicitly. The text in Section 4.5 asserts the formula from Figure 7 rather than deriving it, and a step-by-step derivation is needed to make the comparison quantitative and check the factor of L/p.
- [Section 4.5, Table 2] The activation memory formula 4bsh m L/p is not derived. The text says each stage has L/p layers and each layer stashes 16bsh (4bsh after recomputation), then jumps to the final expression. The factor m needs justification: in the FILO schedule the peak likely occurs at the end of the forward phase when all m micro-batches are live, which is structurally different from the p-i factor in 1F1B. This should be stated explicitly and the peak memory point should be identified in the schedule.
- [Section 5.3] The load-bearing assumption of the two-fold FILO schedule is that the 2bsh per-layer p2p communication is fully hidden behind attention computation. Figure 9 and the text in Section 5.3 show that this fails on A800 at sequence length 32k, where attention is faster than inter-node communication and the advantage of HelixPipe disappears. The paper should state this as an explicit applicability condition (e.g., t_attn >= t_comm) and should not claim, as it does in Section 5.3, that HelixPipe 'can scale to any number of nodes' on H20 based only on experiments up to 8 nodes.
- [Sections 5.1 and 6.1] WeiPipe is identified in Section 6.1 as the closest prior work on inter-layer pipeline parallelism for long-context training, but it is not included in the experimental comparison. Given the paper's claim to outperform existing methods, the absence of a WeiPipe baseline weakens that claim; either add an empirical comparison or explicitly restrict the claim to the evaluated baselines and explain why WeiPipe cannot be included.
minor comments (4)
- [Section 4.2] There is a typo 'pre-attetion' in the description of the attention parallel partition.
- [Section 4.3.2] The word 'Figre' appears in the sentence referring to the purple block; it should read 'Figure 6b'.
- [Sections 5.2 and 5.3] The phrase 'descent speedup' should be 'decent speedup' in the discussion of A800 results.
- [Section 5.1] The word 'orthorgonal' should be 'orthogonal' in the sentence about sequence parallelism and pipeline parallelism.
Circularity Check
No significant circularity: HelixPipe's analytical and empirical claims are self-contained and do not reduce to their own inputs.
full rationale
The paper's central claims are the attention parallel partition and the two-fold FILO schedule. The analytical results (Table 2 and Section 4.5) are derived from the proposed schedule's idle units and the activation sizes per layer, not fitted to the experimental throughput numbers. The bubble-time expressions such as 8(p−1)(t_pre+t_post) are read off the schedule diagrams in Figure 7 under the stated execution-time ratio convention, and the memory expression 4bshmL/p follows from the per-layer stashed activation size (4bsh) multiplied by the number of layers and micro-batches per stage. These are construction-level derivations from the schedule definition rather than predictions defined in terms of fitted constants. The empirical speedups are measured against 1F1B, ZB1P, and AdaPipe, with no fitted parameter being renamed as a prediction. The paper's own admission in Section 5.3 that the overlap fails on A800 at 32k is an honest, falsifiable limitation, not a circular step. The citations to prior work by overlapping authors (WeiPipe, Hanayo, WallFacer) appear only in related-work discussion and do not carry the load-bearing argument; no uniqueness theorem or prior result by the same authors is invoked to forbid alternatives or justify the schedule. The formulas in Table 2 may be asserted more compactly than fully derived, and the baseline comparison omits WeiPipe, but those are completeness or correctness-risk concerns, not circularity. Overall, the derivation chain is self-contained and the experimental evidence is independent of the analytical construction.
Assumptions & free parameters
free parameters (1)
- chunk size c for chunked MLP
assumptions (4)
- domain assumption The attention part (scaled dot-product and softmax) has no model parameters and can be executed on any pipeline stage without weight exchange.
- domain assumption Inter-stage p2p communication of about 2bsh per layer per micro-batch can be fully overlapped with attention computation via the two-fold FILO schedule.
- domain assumption The number of micro-batches m must be divisible by the pipeline size p.
- domain assumption Backward pass time is approximately twice the forward pass time for each transformer layer.
Cite this review
Pith. "Pith review of HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism." pith.science (2026). https://pith.science/paper/QAUZV7WS
@misc{pith2026250700394,
author = {Pith},
title = {Pith review of: HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism},
year = {2026},
howpublished = {\url{https://pith.science/paper/QAUZV7WS}},
note = {Machine review of arXiv:2507.00394}
}
read the original abstract
As transformer sequence lengths grow, existing pipeline parallelisms incur suboptimal performance due to the quadratic attention computation and the substantial memory overhead. To relieve these challenges, we propose HelixPipe, a novel pipeline parallelism for long sequence transformer training. First, HelixPipe introduces attention parallel partition, which schedules attention computations of different micro batches across different pipeline stages in parallel, reducing pipeline bubbles. Second, it employs a two-fold first-in-last-out micro batch schedule to balance memory usage and overlap communication with computation. Additionally, HelixPipe utilizes recomputation without attention and chunked MLP to mitigate fragmentation and enable longer sequences. Experiments demonstrate that HelixPipe gains increasing advantages with longer sequence lengths, and outperforms existing methods in throughput and scalability across varying pipeline sizes, model sizes, and cluster configurations. Notably, it achieves a 26\% speedup over baseline methods when training a 7B model with 128k sequence length on 64 H20 GPUs. Code is available at https://github.com/code-tunnel/Megatron-LM/tree/dev.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al. 2024. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature (2024), 1–3. HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Par...
work page 2024
-
[2]
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al. 2024. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954 (2024)
arXiv 2024
-
[3]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877–1901
2020
-
[4]
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashat- tention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Information Processing Systems 35 (2022), 16344–16359
work page 2022
-
[5]
DeepSeek-AI. 2024. DeepSeek-V2: A Strong, Economical, and Efficient Mixture- of-Experts Language Model. arXiv:2405.04434 [cs.CL]
arXiv 2024
-
[6]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
arXiv 2024
-
[7]
Shiqing Fan, Yi Rong, Chen Meng, Zongyan Cao, Siyu Wang, Zhen Zheng, Chuan Wu, Guoping Long, Jun Yang, Lixue Xia, et al. 2021. DAPPLE: A pipelined data parallel approach for training large models. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. 431–445
work page 2021
-
[8]
Jiarui Fang and Shangchun Zhao. 2024. A Unified Sequence Parallelism Approach for Long Context Generative AI. arXiv preprint arXiv:2405.07719 (2024)
arXiv 2024
Show all 45 references
-
[9]
Cong Guo, Rui Zhang, Jiale Xu, Jingwen Leng, Zihan Liu, Ziyu Huang, Minyi Guo, Hao Wu, Shouren Zhao, Junping Zhao, et al . 2024. GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching. In Proceedings of the 29th ...
2024
-
[10]
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al . 2024. DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence. arXiv preprint arXiv:2401.14196 (2024)
2024 arXiv
-
[11]
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al . 2019. Gpipe: Efficient training of giant neural networks using pipeline parallelism. Advances in neural information processing systems 32 (2019)
2019
-
[12]
Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Reza Yaz- dani Aminadabi, Shuaiwen Leon Song, Samyam Rajbhandari, and Yuxiong He
-
[13]
Taebum Kim, Hyoungjoo Kim, Gyeong-In Yu, and Byung-Gon Chun. 2023. Bpipe: Memory-balanced pipeline parallelism for training large language models. In International Conference on Machine Learning . PMLR, 16639–16653
2023
-
[14]
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. 2024. HunyuanVideo: A Systematic Framework For Large Video Generative Models. arXiv preprint arXiv:2412.03603 (2024)
2024 arXiv
-
[15]
Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems 5 (2023)
2023
-
[16]
Joel Lamy-Poirier. 2023. Breadth-first pipeline parallelism.Proceedings of Machine Learning and Systems 5 (2023)
2023
-
[17]
Dacheng Li, Rulin Shao, Anze Xie, Eric P Xing, Xuezhe Ma, Ion Stoica, Joseph E Gonzalez, and Hao Zhang. 2024. Distflashattn: Distributed memory-efficient at- tention for long-context llms training. In First Conference on Language Modeling
2024
-
[18]
Shigang Li and Torsten Hoefler. 2021. Chimera: efficiently training large-scale neural networks with bidirectional pipelines. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . 1–14
2021
-
[19]
Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li, and Yang You. 2023. Sequence Parallelism: Long Sequence Training from System Perspective. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 2391–2404
2023
-
[20]
Zhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo, Hao Zhang, Dawn Song, and Ion Stoica. 2021. Terapipe: Token-level pipeline parallelism for training large-scale language models. In International Conference on Machine Learning . PMLR, 6543–6552
2021
-
[21]
Junfeng Lin, Ziming Liu, Yang You, Jun Wang, Weihao Zhang, and Rong Zhao
-
[22]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Cheng- gang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024)
2024 arXiv
-
[23]
Hao Liu, Matei Zaharia, and Pieter Abbeel. 2023. Ring attention with blockwise transformers for near-infinite context. arXiv preprint arXiv:2310.01889 (2023)
2023 arXiv
-
[24]
Ziming Liu, Shenggan Cheng, Haotian Zhou, and Yang You. 2023. Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training Efficiency. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . 1–13
2023
-
[25]
Ziming Liu, Shaoyu Wang, Shenggan Cheng, Zhongkai Zhao, Kai Wang, Xuan- lei Zhao, James Demmel, and Yang You. 2024. WallFacer: Harnessing Multi- dimensional Ring Parallelism for Efficient Long Sequence Model Training. arXiv:2407.00611 [cs.DC] https://arxiv.org/abs/2407.00611
2024
-
[26]
Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen, and Matei Zaharia
-
[27]
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al. 2021. Efficient large-scale language model training on gpu clusters using megatron-lm. In Pro...
2021
-
[28]
NVIDIA. 2025. Context Parallelism. Retrieved April 9, 2025 from https://docs.nvidia.com/megatron-core/developer-guide/latest/api- guide/context_parallel.html
2025
-
[29]
NVIDIA. 2025. NCCL. Retrieved April 13, 2025 from https://github.com/NVIDIA/ nccl
2025
-
[30]
Hyungjun Oh, Junyeol Lee, Hyeongju Kim, and Jiwon Seo. 2022. Out-of-order backprop: An effective scheduling technique for deep learning. In Proceedings of the Seventeenth European Conference on Computer Systems . 435–452
2022
-
[31]
Penghui Qi, Xinyi Wan, Guangxing Huang, and Min Lin. 2024. Zero Bubble (Almost) Pipeline Parallelism. InThe Twelfth International Conference on Learning Representations
2024
-
[32]
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. Zero: Memory optimizations toward training trillion parameter models. In SC20: Inter- national Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–16
2020
-
[33]
John K Salmon, Mark A Moraes, Ron O Dror, and David E Shaw. 2011. Parallel random numbers: as easy as 1, 2, 3. InProceedings of 2011 international conference for high performance computing, networking, storage and analysis . 1–12
2011
-
[34]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053 (2019)
2019 arXiv
-
[35]
Zhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng, Shengqi Chen, Haojie Wang, and Wenguang Chen. 2024. AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and Partitioning. In Proceedings of the 29th ACM International Conference on Architectural Support for Pr...
2024
-
[36]
Hugo Touvron and Louis Martin et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv:2307.09288
2023 arXiv
-
[37]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. 2023. LLaMA: Open and Efficient Foundation ...
2023 arXiv
-
[38]
Eric P Xing, Qirong Ho, Wei Dai, Jin-Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu. 2015. Petuum: A new platform for distributed machine learning on big data. In Proceedings of the 21th ACM SIGKDD International Conference on Knowl...
2015
-
[39]
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bho- janapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2020. Large Batch Optimization for Deep Learning: Training BERT in 76 minutes. In International Conference on Learning Representa...
2020
-
[40]
Tailing Yuan, Yuliang Liu, Xucheng Ye, Shenglong Zhang, Jianchao Tan, Bin Chen, Chengru Song, and Di Zhang. 2024. Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hy- brid Parallelism. In 2024 USENIX Annual Technical C...
2024
-
[41]
Zheng Zhang, Yaqi Xia, Hulin Wang, Donglin Yang, Chuang Hu, Xiaobo Zhou, and Dazhao Cheng. 2024. MPMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism. IEEE Transactions on Parallel and Distributed Systems (2024)
2024
-
[42]
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yan- ping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P Xing, et al. 2022. Alpa: Automating inter-and{Intra-Operator} parallelism for distributed deep learning. In 16th USENIX Symposium on Operating Sys...
2022
-
[2021]
InInternational Conference on Machine Learning
Memory-efficient pipeline-parallel dnn training. InInternational Conference on Machine Learning. PMLR, 7937–7947
-
[2024]
In Proceedings of the 43rd ACM Symposium on Principles of Distributed Computing
System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models. In Proceedings of the 43rd ACM Symposium on Principles of Distributed Computing. 121–130
-
[2025]
In Proceedings of the 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming
WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long- Context Large Model Training. In Proceedings of the 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming . 225–238
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.