Pith. sign in

← back to paper

Review history

arxiv: 2606.26997 · 2 revisions

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning

  1. 2026-07-12 CONDITIONAL HIGH v1.1.0-grok45 novelty 6.0
    37120 ms 17676 in 3179 out 2026-07-12T11:51:43.222556+00:00
  2. 2026-06-26 UNVERDICTED LOW v0.9.1-grok novelty 6.0
    66072 ms 5861 in 1384 out 2026-06-26T02:53:43.201588+00:00