REVIEW 2 major objections 5 minor 45 references
BiDiRL reclaims idle GPUs on both sides of asynchronous LLM RL, raising training throughput by up to 1.94× without changing how the model learns.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 04:41 UTC pith:IE22MEPK
load-bearing objection Solid systems paper: same-budget bidirectional borrow under a hot-switch envelope is new relative to StreamRL/AReaL/ROLL, and the 1.05–1.94× throughput story is well measured; the “no effect on convergence” half is thinner than the abstract implies. the 2 major comments →
Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper shows that residual idle time in staleness-bounded disaggregated LLM RL is two-sided and can be harvested by bidirectional, model-guided borrowing inside a fixed GPU budget. With a hot-switch runtime, a scheduling-aware static partition, and an admission rule that compares predicted stage speedup against measured switch overhead, BiDiRL raises training throughput by up to 1.94× across workloads, models, and hardware without altering GRPO convergence behavior.
What carries the argument
The hybrid time-space multiplexing stack: a hot-switch runtime that swaps rollout and training roles with measured overhead, a static planner that returns a hot-switch-compatible resource envelope, and a bidirectional scheduler that admits temporary borrowing only when stage-time models predict net benefit and then splits work between primary and auxiliary replicas.
Load-bearing premise
That short early training curves plus the claim that logical sample groups are preserved under preemption are enough to guarantee that bidirectional placement never changes long-run learning behavior.
What would settle it
Run the same GRPO workload for several hundred steps with and without bidirectional borrowing; if final reward or sample statistics diverge once residual bubbles become large, the claim that placement is learning-neutral fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. BiDiRL is a hybrid time–space multiplexing system for disaggregated, asynchronous LLM RL post-training. It keeps separate committed rollout and training pools but allows either pool to temporarily host the other stage via a hot-switch runtime, a scheduling-aware static planner that selects a hot-switch-compatible resource envelope from stage-time models, and a bidirectional runtime scheduler (Rollouter-on-TrainPoll and Trainer-on-RollPoll) that admits borrowing only when predicted benefit exceeds measured switch cost and splits work with online-calibrated models. On two 32-GPU testbeds (A6000 and H100), across response lengths, staleness bounds, resource budgets, model sizes, and text/multimodal datasets, the paper reports up to 1.94× training throughput over veRL, AReaL, and ROLL, with ablations attributing gains to model-guided bidirectional borrowing, and claims no effect on GRPO convergence.
Significance. If the results hold, BiDiRL addresses a concrete and recurring inefficiency in modern disaggregated RL stacks: residual idle windows that remain after asynchronous overlap and static partitioning. The combination of a hot-switch-compatible planner, measured switch costs, benefit-over-overhead admission, and two-sided borrowing under a fixed GPU budget is a clear systems contribution relative to one-sided elastic rollout or fixed-pool async designs (Table 1). Strengths include multi-baseline, multi-hardware end-to-end evaluation (Figure 5), ablations isolating both borrow directions and model-guided admission (Figure 6), validated stage-time models with median errors of ~3% (Figure 7), and explicit hot-switch cost measurements (Table 5). These make the throughput claim falsifiable and useful for the RL systems community.
major comments (2)
- The joint central claim pairs large throughput gains with “without affecting convergence behavior,” but the empirical support for the second half is thin. Figure 8 reports only the first 60 Geo3K reward steps under s=1 and s=2 versus veRL, with small last-point gaps (+0.017 / +0.000). Sections 5.2 and 6.3 and Table 2 argue that partial-rollout resume and ordered chunk merge preserve logical GRPO groups, yet borrowing still changes weight-sync timing, can interrupt/resume partial groups, and reorders chunk futures before merge. Under the same staleness bounds the system is designed to exploit, short-horizon reward agreement does not rule out long-run divergence of the effective sample stream or gradient timing. Please either (i) extend convergence runs to a substantially longer horizon (and, ideally, a second dataset/model) under the same s settings used in the throughput sweeps, or (ii)
- End-to-end speedups in §7.2 are measured under “the same node-aligned rollout/trainer partitioning as the compared systems,” so the static planner’s selected envelope (Algorithm 1, §4) is not the primary driver of the headline 1.05×–1.94× numbers; its role is mainly to supply hot-switch-compatible layouts and stage models. Figure 7 shows that partition choice matters (up to 2.09× variation for 4B) and that the planner tracks the measured-best partition in the displayed sweeps, but the paper should more clearly separate (a) gains from bidirectional runtime scheduling under a fixed common partition from (b) gains from planner-chosen partitions. Without that separation, readers may over-attribute end-to-end speedups to static planning. A short table or paragraph that reports BiDiRL under the planner-selected partition versus the baseline-aligned partition would make the two contributions lo
minor comments (5)
- Abstract and §3.1 describe hot-switch overhead as “negligible,” while Table 5 reports C_in/C_out of several seconds (e.g., 3.58–7.70 s). §7.3 correctly treats these costs as non-negligible for short windows and gates admission on them. Align the abstract/intro wording with the measured costs and admission rule.
- Figure 5 caption and §7.1 mark unsupported/OOM settings with ×/OOM, but a single consolidated table of which baseline configurations were excluded (and why) would improve reproducibility of the speedup ranges.
- Notation for the resource envelope E in Eq. (2) introduces ρ_r, ρ_t and M_r, M_t; Algorithm 1 then returns d_r, d_t in the best tuple. A one-line clarification that d is induced from (g, ρ) would reduce minor ambiguity between layout and replica count.
- In §5.2.1, prompt groups are split by replica capacity (Eq. 6) without length prediction; §8 notes this limitation. A brief quantitative note on how often interrupted auxiliary groups return partial prefixes would help readers judge resume overhead in practice.
- Typos/polish: “1 .94×” spacing appears repeatedly in the abstract and §1; “Trainpoll/Rollpoll” capitalization is inconsistent with “TrainPoll/RollPoll” in Figure 3.
Circularity Check
No significant circularity: measured throughput and ablations stand independent of fitted stage models.
full rationale
BiDiRL is an empirical systems paper. The central throughput claim (up to 1.94× vs veRL/AReaL/ROLL) is established by end-to-end wall-clock measurements on two 32-GPU testbeds across workloads, not by algebraic rearrangement of fitted parameters. Stage-time models M_r and M_t (Appendix A.1) are calibrated from profiling and used only as ranking/admission heuristics for static partition search and borrow decisions; they are separately validated against held-out measured stage times (median errors ~3%) and further checked by ablations (no-borrow, one-direction, opportunistic). Hot-switch costs are measured, not defined into the speedup. Convergence is argued from logical sample preservation plus short reward curves, which is an evidence-strength issue rather than a definitional loop. Related-work citations (StreamRL, AReaL, veRL, etc.) supply baselines and context, not a self-citation uniqueness chain that forces the result. No equation equates the claimed speedup to a quantity that is the fit by construction.
Axiom & Free-Parameter Ledger
free parameters (2)
- Stage-model coefficients (τ_s, α_s, r_s, α_comm, β_comm, τ_pre/dec, β_tok, β_hist, …)
- Measured hot-switch costs C_in, C_out, C_grad
axioms (3)
- domain assumption Disaggregated rollout and training pools plus a staleness bound s correctly model modern asynchronous LLM RL (partial rollout, GRPO groups, micro-batch chunks).
- domain assumption Primary and auxiliary workers of the same stage can share an identical model layout so that hot switching needs no process restart or resharding.
- ad hoc to paper Replica-max semantics and p-norm compute/communication overlap are adequate predictors of stage time for ranking partitions and admitting borrows.
invented entities (3)
-
Hot-switch runtime
no independent evidence
-
Scheduling-aware static planner / resource envelope E
no independent evidence
-
Bidirectional scheduler (Rollouter-on-TrainPoll + Trainer-on-RollPoll)
no independent evidence
read the original abstract
It is well established that the reasoning capabilities of large language models (LLMs) can be improved by applying reinforcement learning (RL) in a post-training stage. In a standard RL iteration, the current model (the policy) generates experience through rollouts, and the resulting data is then used to update the policy during training. High-performance RL frameworks such as StreamRL and AReaL employ a disaggregated architecture and asynchronous rollouts to better exploit both rollout and training resources, thereby increasing overall system throughput. Nonetheless, across varying RL setups (e.g., hardware configurations, model scales, staleness levels, and hyperparameters) and under changing workloads, it remains common for both rollout and training resources to experience idle periods. In this paper, we present BiDiRL, a hybrid time-space multiplexing architecture for asynchronous, disaggregated RL designed to reduce resource idleness. First, we develop a hot-switch runtime that enables rapid switching between rollout and training resources with negligible overhead. Second, we propose a static, scheduling-aware planner based on time-performance modeling that chooses a hot-switch-friendly resource partition, so that rollout and training durations are roughly balanced at a coarse level. Third, at execution time, we introduce a bidirectional scheduler that further exploits runtime bubbles through fine-grained resource switching, allowing the bottleneck stage to temporarily borrow idle resources from the other pool. Across a wide range of workloads, datasets, and models on two 32-GPU testbeds, BiDiRL increases RL training throughput by up to 1.94x compared with RL systems including veRL, AReaL, and ROLL, without affecting convergence behavior.
Figures
Reference graph
Works this paper leans on
-
[1]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman
-
[2]
Training Verifiers to Solve Math Word Problems.arXiv preprint arXiv:2110.14168(2021). arXiv:2110.14168
Pith/arXiv arXiv 2021
-
[3]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai D...
-
[4]
Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xi- angyu Yue. 2025. Video-R1: Reinforcing Video Reasoning in MLLMs. arXiv:2503.21776 [cs.CV] doi:10.48550/arXiv.2503.21776
-
[5]
Wei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu, Zhiyu Mei, Chuyi He, Shusheng Xu, Guo Wei, Jun Mei, Jiashu Wang, Tongkai Yang, Binhang Yuan, and Yi Wu. 2025. AReaL: A Large-Scale Asynchronous Rein- forcement Learning System for Language Reasoning. doi:10.48550/ ARXIV.2505.24298
Pith/arXiv arXiv 2025
-
[6]
Wei Gao, Yuheng Zhao, Dakai An, Tianyuan Wu, Lunxi Cao, Shaopan Xiong, Ju Huang, Weixun Wang, Siran Yang, Wenbo Su, et al. 2026. RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail Batching. In23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26). 849–866
2026
-
[7]
Wei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong, Weixun Wang, Dakai An, Lunxi Cao, Dilxat Muhtar, Zichen Liu, Haizhou Zhao, et al
-
[8]
RollArt: Scaling Agentic RL Training via Disaggregated Infras- tructure.arXiv preprint arXiv:2512.22560(2025). arXiv:2512.22560
arXiv 2025
-
[9]
Jingkai He, Tianjian Li, Erhu Feng, Dong Du, Qian Liu, Tao Liu, Yubin Xia, and Haibo Chen. 2025. History Rhymes: Accelerating LLM Rein- forcement Learning with RhymeRL. doi:10.48550/ARXIV.2508.18588 13
-
[10]
Jian Hu, Xibin Wu, Wei Shen, Jason Klein Liu, Weixun Wang, Songlin Jiang, Haoran Wang, Hao Chen, Bin Chen, Wenkai Fang, Xianyu, Yu Cao, Haotian Xu, and Yiming Liu. 2025. OpenRLHF: A Ray-Based Easy- to-Use, Scalable and High-Performance RLHF Framework. InProceed- ings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demons...
-
[11]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Sto- ica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InProceedings of the 29th Symposium on Operating Systems Principles. ACM, Koblenz Germany, 611–626. doi:10.1145/3600006.3613165
-
[12]
Kinman Lei, Yuyang Jin, Mingshu Zhai, Kezhao Huang, Haoxing Ye, and Jidong Zhai. [n. d.]. Puzzle: Efficiently Aligning Large Language Models through Light-Weight Context Switch. ([n. d.])
-
[13]
Jiacai Liu, Chaojie Wang, Chris Liu, Liang Zeng, Rui Yan, Yiwen Sun, and Yang Liu. 2025. DAPO: Improving Multi-Step Reasoning Abili- ties of Large Language Models with Direct Advantage-Based Policy Optimization. InAdvances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (Eds.), Vol....
2025
-
[14]
Xiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu, Wentao Ma, Aobo Kong, Fei Huang, Jianbin Jiao, and Junge Zhang. 2025. EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforce- ment Learning. InProceedings of the 63rd Annual Meeting of the Associ- ation for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende...
-
[15]
Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu. 2021. Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning. In Proceedings of the 59th Annual Meeting of the Association for Computa- tional Linguistics and the 11th International Joint Conference on Natural Language Process...
-
[16]
Zhiyu Mei, Wei Fu, Kaiwei Li, Guangju Wang, Huanchen Zhang, and Yi Wu. 2025. ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation. InProceedings of Machine Learning and Systems, Vol. 7
2025
-
[17]
Jordan, and Ion Stoica
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica. 2018. Ray: A Distributed Framework for Emerging AI Applications. In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, 561–577
2018
-
[18]
NVIDIA, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, Yunhao Ge, Jinwei Gu, Siddharth Gururani, Ethan He, Jiahui Huang, Jacob Huff- man, Pooya Jannaty, Jin...
-
[19]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wain- wright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Chris- tiano, Jan Leike, and Ryan Lowe. 2022. Training Language Models to Follow Instructions with Hum...
2022
-
[20]
Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R Ganger, and Eric P Xing. [n. d.]. Pollux: Co-adaptive Cluster Scheduling for Goodput- Optimized Deep Learning. ([n. d.])
-
[21]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo
-
[22]
arXiv:2402.03300 [cs.CL] doi:10.48550/ arXiv.2402.03300
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300 [cs.CL] doi:10.48550/ arXiv.2402.03300
-
[23]
Gerald Shen, Zhilin Wang, Olivier Delalleau, Jiaqi Zeng, Yi Dong, Daniel Egert, Shengyang Sun, Jimmy Zhang, Sahil Jain, Ali Taghibakhshi, Markel Sanz Ausin, Ashwath Aithal, and Oleksii Kuchaiev. 2024. NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment. arXiv:2405.01481 [cs.CL] doi:10.48550/arXiv.2405.01481
-
[24]
Guangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang, Chaobo Jia, Xibin Wu, Yuqi Wu, Xiang Li, Chi Zhang, Yanghua Peng, Haibin Lin, Xin Liu, and Chuan Wu. 2026. Laminar: A Scalable Asynchronous RL Post-Training Framework. InProceedings of the 21st European Confer- ence on Computer Systems. ACM, McEwan Hall/The University of Edin- burgh Edinburgh Scotland U...
-
[25]
Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2025. Hybrid- Flow: A Flexible and Efficient RLHF Framework. InProceedings of the Twentieth European Conference on Computer Systems. ACM, Rotterdam Netherlands, 1279–1297. doi:10.1145/3689031.3696075
-
[26]
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 [cs.CL] doi:10.48550/arXiv.1909.08053
-
[27]
Xin Tan, Yicheng Feng, Yu Zhou, Yimin Jiang, Yibo Zhu, and Hong Xu. 2026. OrchestrRL: Dynamic Compute and Network Orchestra- tion for Disaggregated RL.arXiv preprint arXiv:2601.01209(2026). arXiv:2601.01209
arXiv 2026
-
[28]
2026.EArl:Efficient Agentic RL Post- Training for LLMs under Dynamic Context Lengths
Zheyue Tan, Tuo Shi, Huining Yuan, Zelai Xu, Chao Yu, Boxun Li, Yu Wang, and Bo Zhao. 2026.EArl:Efficient Agentic RL Post- Training for LLMs under Dynamic Context Lengths. InProceedings of the Sixth European Workshop on Machine Learning and Systems. ACM, Edinburgh Scotland Uk, 41–48. doi:10.1145/3805621.3807632
-
[29]
Weixun Wang, Shaopan Xiong, Gengru Chen, Wei Gao, Sheng Guo, Yancheng He, Ju Huang, Jiaheng Liu, Zhendong Li, Xiaoyang Li, Zichen Liu, Haizhou Zhao, Dakai An, Lunxi Cao, Qiyang Cao, Wanxi Deng, Feilei Du, Yiliang Gu, Jiahe Li, Xiang Li, Mingjie Liu, Yijia Luo, Zihe Liu, Yadao Wang, Pei Wang, Tianyuan Wu, Yanan Wu, Yuheng Zhao, Shuaibing Zhao, Jin Yang, Si...
-
[30]
Yan Wang, Wenjie Luo, Junjie Bai, Yulong Cao, Tong Che, Ke Chen, Yuxiao Chen, Jenna Diamond, Yifan Ding, Wenhao Ding, et al. 2025. 14 Alpamayo-R1: Bridging Reasoning and Action Prediction for Gen- eralizable Autonomous Driving in the Long Tail.arXiv preprint arXiv:2511.00088(2025). arXiv:2511.00088
Pith/arXiv arXiv 2025
-
[31]
Zhixin Wang, Tianyi Zhou, Liming Liu, Ao Li, Jiarui Hu, Dian Yang, Yinhui Lu, Jinlong Hou, Siyuan Feng, Yuan Cheng, et al. 2026. DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post- Training. InForty-Third International Conference on Machine Learning (ICML)
2026
-
[32]
Chuan Wu. 2024. A Framework for Training Large Language Models for Code Generation via Proximal Policy Optimization. InNL2Code Workshop of ACM KDD (25/08/2024-29/08/2024, Barcelona)
2024
-
[33]
Junde Wu, Jiayuan Zhu, Yuyuan Liu, Min Xu, and Yueming Jin. 2025. Agentic Reasoning: A Streamlined Framework for Enhancing LLM Rea- soning with Agentic Tools. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Moham- mad Taher Pilehvar (Ed...
-
[34]
Tianyuan Wu, Lunxi Cao, Yining Wei, Wei Gao, Yuheng Zhao, Dakai An, Shaopan Xiong, Zhiqiang Lv, Ju Huang, Siran Yang, Yinghao Yu, Jiamang Wang, Lin Qu, and Wei Wang. 2026. Weave: Efficient Co- Scheduling for Disaggregated RL Post-Training. In20th USENIX Sym- posium on Operating Systems Design and Implementation (OSDI 26). USENIX Association, Seattle, WA, USA
2026
-
[35]
Morley Mao, Arvind Krishnamurthy, and Ion Stoica
Yongji Wu, Xueshen Liu, Haizhong Zheng, Juncheng Gu, Beidi Chen, Z. Morley Mao, Arvind Krishnamurthy, and Ion Stoica. 2025. RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs. doi:10.48550/ARXIV.2510.19225
-
[36]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...
-
[37]
arXiv:2505.09388 [cs.CL] doi:10.48550/ arXiv.2505.09388
Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] doi:10.48550/ arXiv.2505.09388
-
[38]
Zhewei Yao, Reza Yazdani Aminabadi, Olatunji Ruwase, Samyam Rajb- handari, Xiaoxia Wu, Ammar Ahmad Awan, Jeff Rasley, Minjia Zhang, Conglong Li, Connor Holmes, et al. 2023. Deepspeed-Chat: Easy, Fast and Affordable Rlhf Training of Chatgpt-like Models at All Scales. arXiv preprint arXiv:2308.01320(2023). arXiv:2308.01320
Pith/arXiv arXiv 2023
-
[39]
Yiqi Zhang, Huiqiang Jiang, Xufang Luo, Zhihe Yang, Chengruidong Zhang, Yifei Shen, Dongsheng Li, Yuqing Yang, Lili Qiu, and Yang You
-
[40]
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling. doi:10.48550/ARXIV.2603.23414
-
[41]
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. 2023. PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.Proceedings of the VLDB Endowment16, 12 ...
arXiv 2023
-
[42]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Effi- cient Execution of Structured Language Model Programs. InAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paqu...
-
[43]
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). USENIX Association, 193–210
2024
-
[44]
Yinmin Zhong, Zili Zhang, Xiaoniu Song, Hanpeng Hu, Chao Jin, Bingyang Wu, Nuo Chen, Yukun Chen, Yu Zhou, Changyi Wan, Hongyu Zhou, Yimin Jiang, Yibo Zhu, and Daxin Jiang. 2025. StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation. arXiv:2504.15930 doi:10.48550/arXiv.2504.15930
-
[45]
Yinmin Zhong, Zili Zhang, Bingyang Wu, Shengyu Liu, Yukun Chen, Changyi Wan, Hanpeng Hu, Lei Xia, Ranchen Ming, Yibo Zhu, and Xin Jin. 2025. Optimizing RLHF Training for Large Language Models with Stage Fusion. In22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, 489–503. A Appendix A.1 Stage-Time Models Th...
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.