Pith. sign in

REVIEW 2 major objections 5 minor 45 references

BiDiRL reclaims idle GPUs on both sides of asynchronous LLM RL, raising training throughput by up to 1.94× without changing how the model learns.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 04:41 UTC pith:IE22MEPK

load-bearing objection Solid systems paper: same-budget bidirectional borrow under a hot-switch envelope is new relative to StreamRL/AReaL/ROLL, and the 1.05–1.94× throughput story is well measured; the “no effect on convergence” half is thinner than the abstract implies. the 2 major comments →

arxiv 2607.09207 v1 pith:IE22MEPK submitted 2026-07-10 cs.DC

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training

classification cs.DC
keywords LLM RL post-trainingdisaggregated architectureasynchronous rolloutsbidirectional resource schedulinghot-switch runtimestaleness-bounded trainingresource bubblesGPU scheduling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Disaggregated, asynchronous reinforcement-learning post-training for large language models still leaves GPUs idle: a fixed split between generation (rollout) and weight-update (training) pools cannot keep pace with shifting response lengths, staleness limits, and parallelism constraints. BiDiRL treats those idle windows as a two-timescale scheduling problem. Before a job starts, a planner picks a resource envelope that roughly balances the two stages and keeps both pools hot-switchable. At runtime a lightweight switch mechanism lets the bottleneck stage borrow idle devices from the other pool, but only when a profiled time model predicts that the gain exceeds the measured switch cost. On two 32-GPU platforms the design lifts end-to-end token throughput by as much as 1.94 times relative to existing systems while preserving the logical samples that the learning algorithm consumes.

Core claim

The paper shows that residual idle time in staleness-bounded disaggregated LLM RL is two-sided and can be harvested by bidirectional, model-guided borrowing inside a fixed GPU budget. With a hot-switch runtime, a scheduling-aware static partition, and an admission rule that compares predicted stage speedup against measured switch overhead, BiDiRL raises training throughput by up to 1.94× across workloads, models, and hardware without altering GRPO convergence behavior.

What carries the argument

The hybrid time-space multiplexing stack: a hot-switch runtime that swaps rollout and training roles with measured overhead, a static planner that returns a hot-switch-compatible resource envelope, and a bidirectional scheduler that admits temporary borrowing only when stage-time models predict net benefit and then splits work between primary and auxiliary replicas.

Load-bearing premise

That short early training curves plus the claim that logical sample groups are preserved under preemption are enough to guarantee that bidirectional placement never changes long-run learning behavior.

What would settle it

Run the same GRPO workload for several hundred steps with and without bidirectional borrowing; if final reward or sample statistics diverge once residual bubbles become large, the claim that placement is learning-neutral fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. BiDiRL is a hybrid time–space multiplexing system for disaggregated, asynchronous LLM RL post-training. It keeps separate committed rollout and training pools but allows either pool to temporarily host the other stage via a hot-switch runtime, a scheduling-aware static planner that selects a hot-switch-compatible resource envelope from stage-time models, and a bidirectional runtime scheduler (Rollouter-on-TrainPoll and Trainer-on-RollPoll) that admits borrowing only when predicted benefit exceeds measured switch cost and splits work with online-calibrated models. On two 32-GPU testbeds (A6000 and H100), across response lengths, staleness bounds, resource budgets, model sizes, and text/multimodal datasets, the paper reports up to 1.94× training throughput over veRL, AReaL, and ROLL, with ablations attributing gains to model-guided bidirectional borrowing, and claims no effect on GRPO convergence.

Significance. If the results hold, BiDiRL addresses a concrete and recurring inefficiency in modern disaggregated RL stacks: residual idle windows that remain after asynchronous overlap and static partitioning. The combination of a hot-switch-compatible planner, measured switch costs, benefit-over-overhead admission, and two-sided borrowing under a fixed GPU budget is a clear systems contribution relative to one-sided elastic rollout or fixed-pool async designs (Table 1). Strengths include multi-baseline, multi-hardware end-to-end evaluation (Figure 5), ablations isolating both borrow directions and model-guided admission (Figure 6), validated stage-time models with median errors of ~3% (Figure 7), and explicit hot-switch cost measurements (Table 5). These make the throughput claim falsifiable and useful for the RL systems community.

major comments (2)
  1. The joint central claim pairs large throughput gains with “without affecting convergence behavior,” but the empirical support for the second half is thin. Figure 8 reports only the first 60 Geo3K reward steps under s=1 and s=2 versus veRL, with small last-point gaps (+0.017 / +0.000). Sections 5.2 and 6.3 and Table 2 argue that partial-rollout resume and ordered chunk merge preserve logical GRPO groups, yet borrowing still changes weight-sync timing, can interrupt/resume partial groups, and reorders chunk futures before merge. Under the same staleness bounds the system is designed to exploit, short-horizon reward agreement does not rule out long-run divergence of the effective sample stream or gradient timing. Please either (i) extend convergence runs to a substantially longer horizon (and, ideally, a second dataset/model) under the same s settings used in the throughput sweeps, or (ii)
  2. End-to-end speedups in §7.2 are measured under “the same node-aligned rollout/trainer partitioning as the compared systems,” so the static planner’s selected envelope (Algorithm 1, §4) is not the primary driver of the headline 1.05×–1.94× numbers; its role is mainly to supply hot-switch-compatible layouts and stage models. Figure 7 shows that partition choice matters (up to 2.09× variation for 4B) and that the planner tracks the measured-best partition in the displayed sweeps, but the paper should more clearly separate (a) gains from bidirectional runtime scheduling under a fixed common partition from (b) gains from planner-chosen partitions. Without that separation, readers may over-attribute end-to-end speedups to static planning. A short table or paragraph that reports BiDiRL under the planner-selected partition versus the baseline-aligned partition would make the two contributions lo
minor comments (5)
  1. Abstract and §3.1 describe hot-switch overhead as “negligible,” while Table 5 reports C_in/C_out of several seconds (e.g., 3.58–7.70 s). §7.3 correctly treats these costs as non-negligible for short windows and gates admission on them. Align the abstract/intro wording with the measured costs and admission rule.
  2. Figure 5 caption and §7.1 mark unsupported/OOM settings with ×/OOM, but a single consolidated table of which baseline configurations were excluded (and why) would improve reproducibility of the speedup ranges.
  3. Notation for the resource envelope E in Eq. (2) introduces ρ_r, ρ_t and M_r, M_t; Algorithm 1 then returns d_r, d_t in the best tuple. A one-line clarification that d is induced from (g, ρ) would reduce minor ambiguity between layout and replica count.
  4. In §5.2.1, prompt groups are split by replica capacity (Eq. 6) without length prediction; §8 notes this limitation. A brief quantitative note on how often interrupted auxiliary groups return partial prefixes would help readers judge resume overhead in practice.
  5. Typos/polish: “1 .94×” spacing appears repeatedly in the abstract and §1; “Trainpoll/Rollpoll” capitalization is inconsistent with “TrainPoll/RollPoll” in Figure 3.

Circularity Check

0 steps flagged

No significant circularity: measured throughput and ablations stand independent of fitted stage models.

full rationale

BiDiRL is an empirical systems paper. The central throughput claim (up to 1.94× vs veRL/AReaL/ROLL) is established by end-to-end wall-clock measurements on two 32-GPU testbeds across workloads, not by algebraic rearrangement of fitted parameters. Stage-time models M_r and M_t (Appendix A.1) are calibrated from profiling and used only as ranking/admission heuristics for static partition search and borrow decisions; they are separately validated against held-out measured stage times (median errors ~3%) and further checked by ablations (no-borrow, one-direction, opportunistic). Hot-switch costs are measured, not defined into the speedup. Convergence is argued from logical sample preservation plus short reward curves, which is an evidence-strength issue rather than a definitional loop. Related-work citations (StreamRL, AReaL, veRL, etc.) supply baselines and context, not a self-citation uniqueness chain that forces the result. No equation equates the claimed speedup to a quantity that is the fit by construction.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 3 invented entities

As a systems paper the load-bearing content is the design and the measurements, not free mathematical constants. The free parameters are the fitted coefficients inside the stage-time models and the measured switch costs; the axioms are standard RL-system assumptions; the invented entities are the three BiDiRL mechanisms themselves.

free parameters (2)
  • Stage-model coefficients (τ_s, α_s, r_s, α_comm, β_comm, τ_pre/dec, β_tok, β_hist, …)
    Calibrated from profiling runs and used both for static partition selection and for runtime admission; online profiler continues to update them.
  • Measured hot-switch costs C_in, C_out, C_grad
    Empirically timed on the target hardware (Table 5) and inserted into the benefit-over-overhead test; treated as constants for each model size.
axioms (3)
  • domain assumption Disaggregated rollout and training pools plus a staleness bound s correctly model modern asynchronous LLM RL (partial rollout, GRPO groups, micro-batch chunks).
    Stated in §2 and used throughout the time models and scheduler.
  • domain assumption Primary and auxiliary workers of the same stage can share an identical model layout so that hot switching needs no process restart or resharding.
    Enforced by the static planner (§4.1) and required for the measured switch costs to remain small.
  • ad hoc to paper Replica-max semantics and p-norm compute/communication overlap are adequate predictors of stage time for ranking partitions and admitting borrows.
    Appendix A.1; validated only by the reported median/p90 errors on the profiled points.
invented entities (3)
  • Hot-switch runtime no independent evidence
    purpose: Make rollouter and trainer roles exchangeable on a committed pool with preemption and recoverable pending work.
    Core mechanism that turns the planner envelope into an executable bidirectional system (§3.2).
  • Scheduling-aware static planner / resource envelope E no independent evidence
    purpose: Select a hot-switch-compatible, rate-balanced GPU partition before training starts.
    Algorithm 1; supplies the layouts and stage models used by the runtime.
  • Bidirectional scheduler (Rollouter-on-TrainPoll + Trainer-on-RollPoll) no independent evidence
    purpose: Admit and size temporary borrows of idle resources when predicted benefit exceeds measured switch cost.
    §5; the component that harvests residual bubbles.

pith-pipeline@v1.1.0-grok45 · 28110 in / 2732 out tokens · 38308 ms · 2026-07-13T04:41:13.841615+00:00 · methodology

0 comments
read the original abstract

It is well established that the reasoning capabilities of large language models (LLMs) can be improved by applying reinforcement learning (RL) in a post-training stage. In a standard RL iteration, the current model (the policy) generates experience through rollouts, and the resulting data is then used to update the policy during training. High-performance RL frameworks such as StreamRL and AReaL employ a disaggregated architecture and asynchronous rollouts to better exploit both rollout and training resources, thereby increasing overall system throughput. Nonetheless, across varying RL setups (e.g., hardware configurations, model scales, staleness levels, and hyperparameters) and under changing workloads, it remains common for both rollout and training resources to experience idle periods. In this paper, we present BiDiRL, a hybrid time-space multiplexing architecture for asynchronous, disaggregated RL designed to reduce resource idleness. First, we develop a hot-switch runtime that enables rapid switching between rollout and training resources with negligible overhead. Second, we propose a static, scheduling-aware planner based on time-performance modeling that chooses a hot-switch-friendly resource partition, so that rollout and training durations are roughly balanced at a coarse level. Third, at execution time, we introduce a bidirectional scheduler that further exploits runtime bubbles through fine-grained resource switching, allowing the bottleneck stage to temporarily borrow idle resources from the other pool. Across a wide range of workloads, datasets, and models on two 32-GPU testbeds, BiDiRL increases RL training throughput by up to 1.94x compared with RL systems including veRL, AReaL, and ROLL, without affecting convergence behavior.

Figures

Figures reproduced from arXiv: 2607.09207 by Chu Xiaowen, Shi Shaohuai, Tan Zhiqiang, Wang Maoxin, Wang Qiang, Wang Sijie, Yin Yiming.

Figure 1
Figure 1. Figure 1: Profiling evidence for resource bubbles in disaggre￾gated LLM RL. Rollout and training exhibit different scaling trends [39] (top row). Static partitions still leave resource bubbles (bottom row). duration of one training window as 𝑇step ≈ 𝑇wait +𝑇consume, (1) where 𝑇wait is the time until the rollout buffer contains enough valid rollout groups under the staleness bound, and 𝑇consume is the time for traine… view at source ↗
Figure 2
Figure 2. Figure 2: Motivation for bidirectional scheduling in disag￾gregated LLM RL. With bounded off-policy execution, rollout and training can overlap but bubbles remain when (a) rollout is slower, (b) training is slower, or (c) the staleness bound forces trainers to wait for fresh samples after rollouters run ahead. budget into committed rollout and training pools. This par￾titioning problem appears in existing disaggrega… view at source ↗
Figure 4
Figure 4. Figure 4: Bidirectional scheduling. The two directions share an overhead-aware admission interface but use different work units and recovery rules. etc.). Based on the profiling information and workload mod￾els, BiDiRL formulates each borrowing opportunity as an admission-control and workload-splitting problem. A bor￾rowed window is admitted only when the predicted benefit outweighs the measured hot-switch overhead,… view at source ↗
Figure 5
Figure 5. Figure 5: End-to-end throughput comparison across A6000/H100 workloads and resource settings. Bars show raw throughput; speedup ranges compare BiDiRL with the strongest valid external baseline. OOM, failed, and unsupported runs are excluded; × marks unsupported settings and ‡ marks H100 results [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: A6000 ablation throughput summary. The figure isolates the contribution of model-guided bidirectional scheduling. Bars show raw throughput; speedup ranges compare BiDiRL with the strongest valid ablated variant, excluding failed runs. Bidirectional scheduling. Both borrowing directions are necessary because different runtime states expose differ￾ent idle pools. Across the ablation sweeps, BiDiRL is 1.02×– … view at source ↗
Figure 7
Figure 7. Figure 7: Static planning and model validation. Partition sweeps compare measured throughput with planner pre￾dictions, where (𝑅,𝑇 ) denotes rollout/trainer devices. The rollout and trainer panels compare measured and predicted stage time. 0 10 20 30 40 50 60 step 0.40 0.45 0.50 0.55 Geo3K reward BiDiRL, s=1 BiDiRL, s=2 veRL, s=1 veRL, s=2 [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Convergence behavior over the first 60 training steps. BiDiRL changes computation placement and timing while preserving the logical rollout groups and training sam￾ples consumed by GRPO. Stage-time model validation. The stage-time models provide the ranking signal used by both static planning and runtime admission; Appendix A.1 gives their definitions. Fig￾ure 7c shows the rollout model and Figure 7d shows… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

45 extracted references · 1 canonical work pages

  1. [1]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman

  2. [2]

    arXiv:2110.14168

    Training Verifiers to Solve Math Word Problems.arXiv preprint arXiv:2110.14168(2021). arXiv:2110.14168

  3. [3]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai D...

  4. [4]

    Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, and Xi- angyu Yue. 2025. Video-R1: Reinforcing Video Reasoning in MLLMs. arXiv:2503.21776 [cs.CV] doi:10.48550/arXiv.2503.21776

  5. [5]

    Wei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu, Zhiyu Mei, Chuyi He, Shusheng Xu, Guo Wei, Jun Mei, Jiashu Wang, Tongkai Yang, Binhang Yuan, and Yi Wu. 2025. AReaL: A Large-Scale Asynchronous Rein- forcement Learning System for Language Reasoning. doi:10.48550/ ARXIV.2505.24298

  6. [6]

    Wei Gao, Yuheng Zhao, Dakai An, Tianyuan Wu, Lunxi Cao, Shaopan Xiong, Ju Huang, Weixun Wang, Siran Yang, Wenbo Su, et al. 2026. RollPacker: Taming Long-Tail Rollouts for RL Post-Training with Tail Batching. In23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26). 849–866

  7. [7]

    Wei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong, Weixun Wang, Dakai An, Lunxi Cao, Dilxat Muhtar, Zichen Liu, Haizhou Zhao, et al

  8. [8]

    arXiv:2512.22560

    RollArt: Scaling Agentic RL Training via Disaggregated Infras- tructure.arXiv preprint arXiv:2512.22560(2025). arXiv:2512.22560

  9. [9]

    Jingkai He, Tianjian Li, Erhu Feng, Dong Du, Qian Liu, Tao Liu, Yubin Xia, and Haibo Chen. 2025. History Rhymes: Accelerating LLM Rein- forcement Learning with RhymeRL. doi:10.48550/ARXIV.2508.18588 13

  10. [10]

    Jian Hu, Xibin Wu, Wei Shen, Jason Klein Liu, Weixun Wang, Songlin Jiang, Haoran Wang, Hao Chen, Bin Chen, Wenkai Fang, Xianyu, Yu Cao, Haotian Xu, and Yiming Liu. 2025. OpenRLHF: A Ray-Based Easy- to-Use, Scalable and High-Performance RLHF Framework. InProceed- ings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demons...

  11. [11]

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Sto- ica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InProceedings of the 29th Symposium on Operating Systems Principles. ACM, Koblenz Germany, 611–626. doi:10.1145/3600006.3613165

  12. [12]

    Kinman Lei, Yuyang Jin, Mingshu Zhai, Kezhao Huang, Haoxing Ye, and Jidong Zhai. [n. d.]. Puzzle: Efficiently Aligning Large Language Models through Light-Weight Context Switch. ([n. d.])

  13. [13]

    Jiacai Liu, Chaojie Wang, Chris Liu, Liang Zeng, Rui Yan, Yiwen Sun, and Yang Liu. 2025. DAPO: Improving Multi-Step Reasoning Abili- ties of Large Language Models with Direct Advantage-Based Policy Optimization. InAdvances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen (Eds.), Vol....

  14. [14]

    Xiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu, Wentao Ma, Aobo Kong, Fei Huang, Jianbin Jiao, and Junge Zhang. 2025. EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforce- ment Learning. InProceedings of the 63rd Annual Meeting of the Associ- ation for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende...

  15. [15]

    Pan Lu, Ran Gong, Shibiao Jiang, Liang Qiu, Siyuan Huang, Xiaodan Liang, and Song-Chun Zhu. 2021. Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning. In Proceedings of the 59th Annual Meeting of the Association for Computa- tional Linguistics and the 11th International Joint Conference on Natural Language Process...

  16. [16]

    Zhiyu Mei, Wei Fu, Kaiwei Li, Guangju Wang, Huanchen Zhang, and Yi Wu. 2025. ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation. InProceedings of Machine Learning and Systems, Vol. 7

  17. [17]

    Jordan, and Ion Stoica

    Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica. 2018. Ray: A Distributed Framework for Emerging AI Applications. In13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, 561–577

  18. [18]

    NVIDIA, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, Yunhao Ge, Jinwei Gu, Siddharth Gururani, Ethan He, Jiahui Huang, Jacob Huff- man, Pooya Jannaty, Jin...

  19. [19]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wain- wright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Chris- tiano, Jan Leike, and Ryan Lowe. 2022. Training Language Models to Follow Instructions with Hum...

  20. [20]

    Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R Ganger, and Eric P Xing. [n. d.]. Pollux: Co-adaptive Cluster Scheduling for Goodput- Optimized Deep Learning. ([n. d.])

  21. [21]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo

  22. [22]

    arXiv:2402.03300 [cs.CL] doi:10.48550/ arXiv.2402.03300

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300 [cs.CL] doi:10.48550/ arXiv.2402.03300

  23. [23]

    Gerald Shen, Zhilin Wang, Olivier Delalleau, Jiaqi Zeng, Yi Dong, Daniel Egert, Shengyang Sun, Jimmy Zhang, Sahil Jain, Ali Taghibakhshi, Markel Sanz Ausin, Ashwath Aithal, and Oleksii Kuchaiev. 2024. NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment. arXiv:2405.01481 [cs.CL] doi:10.48550/arXiv.2405.01481

  24. [24]

    Guangming Sheng, Yuxuan Tong, Borui Wan, Wang Zhang, Chaobo Jia, Xibin Wu, Yuqi Wu, Xiang Li, Chi Zhang, Yanghua Peng, Haibin Lin, Xin Liu, and Chuan Wu. 2026. Laminar: A Scalable Asynchronous RL Post-Training Framework. InProceedings of the 21st European Confer- ence on Computer Systems. ACM, McEwan Hall/The University of Edin- burgh Edinburgh Scotland U...

  25. [25]

    Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2025. Hybrid- Flow: A Flexible and Efficient RLHF Framework. InProceedings of the Twentieth European Conference on Computer Systems. ACM, Rotterdam Netherlands, 1279–1297. doi:10.1145/3689031.3696075

  26. [26]

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 [cs.CL] doi:10.48550/arXiv.1909.08053

  27. [27]

    Xin Tan, Yicheng Feng, Yu Zhou, Yimin Jiang, Yibo Zhu, and Hong Xu. 2026. OrchestrRL: Dynamic Compute and Network Orchestra- tion for Disaggregated RL.arXiv preprint arXiv:2601.01209(2026). arXiv:2601.01209

  28. [28]

    2026.EArl:Efficient Agentic RL Post- Training for LLMs under Dynamic Context Lengths

    Zheyue Tan, Tuo Shi, Huining Yuan, Zelai Xu, Chao Yu, Boxun Li, Yu Wang, and Bo Zhao. 2026.EArl:Efficient Agentic RL Post- Training for LLMs under Dynamic Context Lengths. InProceedings of the Sixth European Workshop on Machine Learning and Systems. ACM, Edinburgh Scotland Uk, 41–48. doi:10.1145/3805621.3807632

  29. [29]

    Weixun Wang, Shaopan Xiong, Gengru Chen, Wei Gao, Sheng Guo, Yancheng He, Ju Huang, Jiaheng Liu, Zhendong Li, Xiaoyang Li, Zichen Liu, Haizhou Zhao, Dakai An, Lunxi Cao, Qiyang Cao, Wanxi Deng, Feilei Du, Yiliang Gu, Jiahe Li, Xiang Li, Mingjie Liu, Yijia Luo, Zihe Liu, Yadao Wang, Pei Wang, Tianyuan Wu, Yanan Wu, Yuheng Zhao, Shuaibing Zhao, Jin Yang, Si...

  30. [30]

    Yan Wang, Wenjie Luo, Junjie Bai, Yulong Cao, Tong Che, Ke Chen, Yuxiao Chen, Jenna Diamond, Yifan Ding, Wenhao Ding, et al. 2025. 14 Alpamayo-R1: Bridging Reasoning and Action Prediction for Gen- eralizable Autonomous Driving in the Long Tail.arXiv preprint arXiv:2511.00088(2025). arXiv:2511.00088

  31. [31]

    Zhixin Wang, Tianyi Zhou, Liming Liu, Ao Li, Jiarui Hu, Dian Yang, Yinhui Lu, Jinlong Hou, Siyuan Feng, Yuan Cheng, et al. 2026. DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post- Training. InForty-Third International Conference on Machine Learning (ICML)

  32. [32]

    Chuan Wu. 2024. A Framework for Training Large Language Models for Code Generation via Proximal Policy Optimization. InNL2Code Workshop of ACM KDD (25/08/2024-29/08/2024, Barcelona)

  33. [33]

    Junde Wu, Jiayuan Zhu, Yuyuan Liu, Min Xu, and Yueming Jin. 2025. Agentic Reasoning: A Streamlined Framework for Enhancing LLM Rea- soning with Agentic Tools. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Moham- mad Taher Pilehvar (Ed...

  34. [34]

    Tianyuan Wu, Lunxi Cao, Yining Wei, Wei Gao, Yuheng Zhao, Dakai An, Shaopan Xiong, Zhiqiang Lv, Ju Huang, Siran Yang, Yinghao Yu, Jiamang Wang, Lin Qu, and Wei Wang. 2026. Weave: Efficient Co- Scheduling for Disaggregated RL Post-Training. In20th USENIX Sym- posium on Operating Systems Design and Implementation (OSDI 26). USENIX Association, Seattle, WA, USA

  35. [35]

    Morley Mao, Arvind Krishnamurthy, and Ion Stoica

    Yongji Wu, Xueshen Liu, Haizhong Zheng, Juncheng Gu, Beidi Chen, Z. Morley Mao, Arvind Krishnamurthy, and Ion Stoica. 2025. RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs. doi:10.48550/ARXIV.2510.19225

  36. [36]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...

  37. [37]

    arXiv:2505.09388 [cs.CL] doi:10.48550/ arXiv.2505.09388

    Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] doi:10.48550/ arXiv.2505.09388

  38. [38]

    Zhewei Yao, Reza Yazdani Aminabadi, Olatunji Ruwase, Samyam Rajb- handari, Xiaoxia Wu, Ammar Ahmad Awan, Jeff Rasley, Minjia Zhang, Conglong Li, Connor Holmes, et al. 2023. Deepspeed-Chat: Easy, Fast and Affordable Rlhf Training of Chatgpt-like Models at All Scales. arXiv preprint arXiv:2308.01320(2023). arXiv:2308.01320

  39. [39]

    Yiqi Zhang, Huiqiang Jiang, Xufang Luo, Zhihe Yang, Chengruidong Zhang, Yifei Shen, Dongsheng Li, Yuqing Yang, Lili Qiu, and Yang You

  40. [40]

    doi:10.48550/ARXIV.2603.23414

    SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling. doi:10.48550/ARXIV.2603.23414

  41. [41]

    Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. 2023. PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.Proceedings of the VLDB Endowment16, 12 ...

  42. [42]

    Gonzalez, Clark Barrett, and Ying Sheng

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Effi- cient Execution of Structured Language Model Programs. InAdvances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paqu...

  43. [43]

    Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xu- anzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving. In18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). USENIX Association, 193–210

  44. [44]

    Yinmin Zhong, Zili Zhang, Xiaoniu Song, Hanpeng Hu, Chao Jin, Bingyang Wu, Nuo Chen, Yukun Chen, Yu Zhou, Changyi Wan, Hongyu Zhou, Yimin Jiang, Yibo Zhu, and Daxin Jiang. 2025. StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation. arXiv:2504.15930 doi:10.48550/arXiv.2504.15930

  45. [45]

    Yinmin Zhong, Zili Zhang, Bingyang Wu, Shengyu Liu, Yukun Chen, Changyi Wan, Hanpeng Hu, Lei Xia, Ranchen Ming, Yibo Zhu, and Xin Jin. 2025. Optimizing RLHF Training for Large Language Models with Stage Fusion. In22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). USENIX Association, 489–503. A Appendix A.1 Stage-Time Models Th...