REVIEW 3 major objections 5 minor 41 references
πR² makes a pretrained VLA replan closed-loop at ~25 Hz by splitting fresh proprioception from stale vision-language features and emitting actions in one denoising step per call.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 00:44 UTC pith:LKOKXEE6
load-bearing objection Core idea is plausible and the sim study is solid, but the headline 4x speedup is confounded by a 2-GPU vs 1-GPU comparison and the real-world evidence is too thin to take the 30% gains at face value. the 3 major comments →
πR²: Reactive Real-time Flow Policies
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's claim is that the two bottlenecks of chunked flow policies — open-loop chunks and slow perception-to-action latency — are removable by construction. πR² assigns each position p of the action chunk its own noise level via a three-region staircase τ⋆,d: the front d slots are clamped clean (in-flight actions already sent to the robot), the interior ramps linearly from clean to noise, and the tail holds d pure-noise slots. One Euler denoising step per policy call advances the schedule by d slots, so d clean actions are released at the front while the freshly emitted actions become the new inpainting conditioning; the per-call cost is a single function evaluation of the action head. T
What carries the argument
The load-bearing object is the delay-parameterized staircase noise schedule τ⋆,d (Eq. 3), a three-region per-position noise profile: a clamped-clean front that treats in-flight actions as inpainting conditioning, a linear ramp over the interior, and a pure-noise tail. Combined with the asynchronous fast/slow conditioning split (fresh proprioception, cached VLM features with learned delay embedding), it lets one denoising step per call emit d clean actions and makes the model adaptive to measured hardware latency, because training samples d uniformly and the schedule reproduces itself after each slide.
Load-bearing premise
The whole reactivity gain rests on the claim that fresh joint angles, torques, and fingertip forces carry enough information about the immediate situation that the action head can ignore vision and language for up to five control ticks (200 ms); if that claim fails, the asynchronous slow channel adds nothing and πR² is just a latency-adaptive schedule running on stale observations.
What would settle it
On the real Catch Book task, keep the staircase schedule but replace the cached visual feature with one from d_vis = 8 ticks (320 ms) earlier while proprioception stays fresh; if the gripper consistently fails to close on the falling book, the proprioceptive channel alone was not carrying the reaction.
If this is right
- Any pretrained VLA with a flow-matching action head can be fine-tuned into this reactive mode with a one-line DiT change (per-position AdaLN) and a frozen backbone, preserving the semantic grounding of the base model.
- Actions are no longer committed for a full chunk: each 40 ms tick conditions the emitted action on fresh joint angles, torques, and fingertip forces, so contact events can alter the motion mid-chunk.
- A single trained model adapts to varying hardware and network latency at deployment, because the schedule is parameterized by the measured per-call delay d.
- Vision-language staleness up to about 200 ms (d_vis=5 ticks) is tolerated without losing task success, indicating that the slow channel provides coarse guidance while the fast channel drives local correction.
- In the simulation deployment study, πR²'s advantage over naive asynchronous execution and training-time action conditioning widens as the underlying latency budget grows.
Where Pith is reading between the lines
- (Editorial extension) The same slow/fast split should transfer to any expensive conditioning stream, such as video tokens or world-model state predictions, whenever the cheap stream (proprioception or joint encoders) plausibly carries the high-frequency information.
- (Editorial extension) A sharp test of the paper's central premise: on a contact-rich task, feed the final denoising step a deliberately wrong visual feature (e.g., a frame from a different episode) and keep proprioception fresh; if success degrades sharply, proprioception alone is not sufficient for local corrections and the async split is doing less work than claimed.
- (Editorial extension) The method's gains should scale with how contact-rich and time-critical the task is; on quasi-static pick-and-place tasks the 4× latency reduction may not translate into measurable success changes, a boundary case the paper does not explore.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes πR², a method to make flow-matching VLA policies reactive by (1) splitting conditioning into a fast proprioceptive channel and a slow vision-language channel updated asynchronously, and (2) using a delay-parameterized staircase noise schedule (Eq. 3) that emits clean actions from a diffusion-forcing buffer with one denoising step per call. The authors fine-tune GR00T-N1.7 on a real xArm6+XHand and report ~25 Hz closed-loop replanning (4× speedup over the base pipeline), with success-rate gains up to 23% in simulation and up to 30% in the real world. They also present simulation ablations (execution-horizon sweep, effective-delay study, vision-staleness sweep) to support the mechanism.
Significance. The paper tackles an important and timely problem: large VLA action-chunking policies are too slow for reactive closed-loop control. The two ideas—asynchronous fast/slow conditioning and a latency-adaptive diffusion-forcing schedule—are simple, architecture-agnostic, and compatible with pretrained VLAs. The simulation study is well-designed (controlled delays, d_vis ablation, h-sweep), and the real-world tasks are genuinely reactive. If the performance claims survive an equal-hardware comparison, this would be a useful contribution. However, the headline real-world speedup is currently confounded by the use of 2 GPUs for πR² and 1 GPU for baselines, and the zero-delay schedule is under-specified.
major comments (3)
- [§4.2, Table 1, Appendix A.2] The headline real-world speedup (abstract: ~4× faster, ~25 Hz; §4.2: d=1–2 vs d=4–5) is measured with πR² on 2×RTX A5000 and all baselines on 1×RTX A5000. Appendix A.2 states this explicitly and argues that sharing one GPU 'would conflate the measured d,' but no experiment controls for the extra hardware. The 25 Hz claim and the Table 1 success gains could therefore be due to the additional GPU, not to the proposed schedule/async split. Please provide an equal-hardware comparison (e.g., πR² with both workers on one A5000, or baselines with the VLM on a second A5000) and report the measured control-tick delay for each configuration.
- [§3.3 vs §4.1.1] The inference cycle is described as shifting the schedule right by d slots and emitting positions [d,2d), so d clean actions are released per call and the buffer slides by d. For d=0, Eq. (3) gives an empty front and tail and a degenerate slide, yet §4.1.1 evaluates 'πR² ... emits one action per call (h=1)' in a zero-delay setting. The paper does not specify the step-size rule that produces exactly one emitted action in this case. This matters because the zero-delay h-sweep is used to claim that one-step amortized denoising matches h=1 flow. Please clarify the d=0 schedule/protocol, or treat the experiment as d=1 and adjust the interpretation.
- [§4.2 and Table 1] The real-world evaluation does not include an ablation that removes the asynchronous fast/slow split. The success-rate gains over Train-Time RTC (up to 30%) are attributed to proprioceptive reactivity, but the only direct evidence for this mechanism is the simulation d_vis sweep with a low-dimensional vision proxy. The real-world policy uses rich VLM features and allows up to d_vis=5 ticks (200 ms) of visual staleness during training, a regime not covered by the sim. Please add a real-world run of 'πR² w/o async' (or a d_vis ablation) to confirm that the async split—rather than the schedule or the additional GPU—is responsible for the real-world improvement.
minor comments (5)
- [§4.2] Typo: 'Tmaihe same pattern holds' should be 'The same pattern holds.'
- [§3.2 vs Appendix A.2] Notation is inconsistent: §3.2 uses d_vlm for the vision-language staleness, while §4.1.2 and Appendix A.2 use d_vis. Please unify.
- [Abstract and §4.2] The abstract says '~25Hz on an A5000 GPU,' but deployment uses 2×A5000. Please specify the hardware count in the abstract or change the wording to avoid implying a single A5000.
- [Table 1] The Catch Book columns list only SR and no Prog, even though the task is described as having one subgoal. Please add the Prog column or explain the omission.
- [Fig. 3(right) caption] The caption says 'for each datapoint, d indicates the effective delay of proprioception, while d_v is visual delay,' but the text does not specify which d_vis values correspond to the plotted w/ async curve. Please clarify whether the curve is a specific d_vis value or an average over the d_vis sweep.
Circularity Check
No significant circularity: πR2's success claims are empirically measured, its schedule is a stated design choice, and prior-work citations are external and non-load-bearing.
full rationale
The paper's central quantitative claims (25 Hz replanning, ~4x speedup, success improvements) are direct measurements of wall-clock latency and task outcomes, not derived from assumptions that already contain the results. Eq. 3 introduces the staircase schedule as a design choice and explicitly frames it as a diffusion-forcing generalization of external prior work (Train-Time RTC, streaming diffusion), rather than as a prediction fitted to the evaluation. Eq. 4's one-step emission property follows from the chosen per-position Δτ advances, and the paper clearly labels this as construction. The simulation deployment study assigns effective delays using a stated computation-cost model and then measures success; the real-world experiments measure actual delay ticks (d=4–5 for baselines, d=1–2 for πR2) and report success rates. No parameter is fitted to the evaluation and then renamed a prediction. The only author-overlapping citation (Open X-Embodiment, a large collaboration) supports a background statement about VLA scaling and is not load-bearing; no uniqueness theorem or ansatz is imported from the authors' own prior work. The deployment hardware imbalance (πR2 on two GPUs vs baselines on one) is a benchmarking fairness issue that the paper discloses, but it is not a circularity of the kind where an output is equivalent to an input by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- d_max (training max delay) =
5 (sim and real); 10 for RTC baseline
- α (standard-flow warm-up probability) =
0.2
- j (symmetric jitter magnitude) =
not reported
- d_max^vis (max vision delay) =
5 ticks (200 ms at 25 Hz)
axioms (5)
- standard math Standard flow matching objective and conditional interpolation (Eq. 1) are valid for action-chunking policies.
- standard math Per-position noise scheduling as in diffusion forcing (Eq. 2) preserves the training objective when noise levels differ across chunk positions.
- domain assumption For dynamic contact-rich tasks, fresh proprioception carries sufficient information for local reactive corrections, while stale vision-language features up to ~200 ms provide adequate global context.
- domain assumption A single Euler step of per-position size up to d/(H−2d) produces accurate enough clean actions for control; the velocity field is smooth enough for this large step.
- domain assumption The measured deployment delay d and vision staleness d_vis are within the training ranges and the learned delay embedding transfers from zero-init to real measured delays.
read the original abstract
Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies ill-suited for dynamic, closed-loop control. We present $\pi\mathbf{R}^2$, which makes these policies reactive and real-time while retaining large backbones, expressive multi-modal policies, and multi-action prediction. Built on the per-position noise schedule of diffusion forcing, $\pi\mathbf{R}^2$ contributes two ideas. First, it splits conditioning into a fast channel (proprioception, fresh every tick) and an asynchronously updated slow channel (vision-language features), so the policy reacts to proprioception within a chunk while tolerating stale vision. Second, a latency-adaptive flow schedule treats in-flight actions as inpainting conditioning and emits actions in one denoising step per call, letting one trained model adapt to varying hardware latency. Requiring minimal modification to existing architectures, $\pi\mathbf{R}^2$ can be finetuned from a pretrained policy: applied to GR00T-N1.7 on a real xArm6+XHand platform, it replans closed-loop roughly $4\times$ faster than the base policy (~$25$Hz on an A5000 GPU), acting on a fresh observation every $40$ms. Across simulation and real-world manipulation tasks, $\pi\mathbf{R}^2$ improves the success rate by up to $23\%$ in simulation and $30\%$ in the real world over the strongest baseline. Project page: https://pi-r2-flow.github.io/
Figures
Reference graph
Works this paper leans on
-
[1]
Black, N
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al.π0: A vision-language-action flow model for general robot control. InRSS, 2025
2025
-
[2]
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. V...
Pith/arXiv arXiv 2025
-
[4]
Brohan, N
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. In RSS, 2023
2023
-
[5]
Zitkovich, T
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In CoRL, 2023
2023
-
[6]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model. InCoRL, 2024
2024
-
[7]
H. Liu, C. Li, Q. Wu, and Y . J. Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
2023
-
[8]
Alayrac, J
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Mil- lican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022
2022
-
[9]
S. Bai, Y . Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025
Pith/arXiv arXiv 2025
-
[10]
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805, 2023
Pith/arXiv arXiv 2023
-
[11]
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Al- abdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Boˇsnjak, X. Chen, M. Minderer, P. V oigtlaender, I. Bica, I. Balazevic, J. Puigcer...
Pith/arXiv arXiv 2024
-
[12]
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y . Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, J. Lu, T. Anderson, E. Bransom, K. Ehsani, H. Ngo, Y . Chen, A. Patel, M. Yatskar, C. Callison-Burch, A. Head, R. Hendrix, F. Bastani, E. VanderBilt, N. Lambert, Y . Chou, A. Chheda, J. Sparks, S. Skjonsberg, M. Schmitz, A. Sarnat, B. Bischoff, P...
Pith/arXiv arXiv 2024
-
[13]
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. InIJRR, 2023
2023
-
[14]
C. Yuan, C. Wen, T. Zhang, and Y . Gao. General flow as foundation affordance for scalable robot learning.arXiv preprint arXiv:2401.11439, 2024
Pith/arXiv arXiv 2024
-
[15]
C. Pan, G. Anantharaman, N.-C. Huang, C. Jin, D. Pfrommer, C. Yuan, F. Permenter, G. Qu, N. Boffi, G. Shi, et al. Much ado about noising: Dispelling the myths of generative robotic control.arXiv preprint arXiv:2512.01809, 2025
arXiv 2025
-
[16]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware.arXiv preprint arXiv:2304.13705, 2023
Pith/arXiv arXiv 2023
-
[17]
T. T. Zhang, D. Pfrommer, C. Pan, N. Matni, and M. Simchowitz. Action chunking and ex- ploratory data collection yield exponential improvements in behavior cloning for continuous control.arXiv preprint arXiv:2507.09061, 2025
arXiv 2025
-
[18]
B. Chen, D. Mart ´ı Mons ´o, Y . Du, M. Simchowitz, R. Tedrake, and V . Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advances in Neural Information Processing Systems, 37:24081–24125, 2024
2024
-
[19]
S. H. Høeg, Y . Du, and O. Egeland. Streaming diffusion policy: Fast policy synthesis with variable noise diffusion models.arXiv preprint arXiv:2406.04806, 2024
Pith/arXiv arXiv 2024
-
[20]
Z. Fu, T. Z. Zhao, and C. Finn. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation, 2024. URLhttps://arxiv.org/abs/2401.02117
Pith/arXiv arXiv 2024
-
[21]
T. Z. Zhao, J. Tompson, D. Driess, P. Florence, K. Ghasemipour, C. Finn, and A. Wahid. Aloha unleashed: A simple recipe for robot dexterity.arXiv preprint arXiv:2410.13126, 2024
Pith/arXiv arXiv 2024
-
[22]
Q. Zhang, Z. Liu, H. Fan, G. Liu, B. Zeng, and S. Liu. Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow matching for robot manipulation, 2024. URLhttps: //arxiv.org/abs/2412.04987
Pith/arXiv arXiv 2024
-
[23]
G. Yan, J. Zhu, Y . Deng, S. Yang, R.-Z. Qiu, X. Cheng, M. Memmel, R. Krishna, A. Goyal, X. Wang, and D. Fox. Maniflow: A general robot manipulation policy via consistency flow training, 2025. URLhttps://arxiv.org/abs/2509.01819
Pith/arXiv arXiv 2025
-
[24]
O. X.-E. Collaboration, A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Her- zog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakr- ishna, A. W...
Pith/arXiv arXiv 2023
-
[25]
Khazatsky, K
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. InRSS, 2024
2024
-
[26]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Pith/arXiv arXiv 2024
-
[27]
NVIDIA, :, J. Bjorck, F. Casta ˜neda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y . Xie, Y . Xu, Z. Xu, S. Ye, Z...
Pith/arXiv arXiv 2025
-
[28]
H. Xie, B. Wen, J. Zheng, Z. Chen, F. Hong, H. Diao, and Z. Liu. Dynamicvla: A vision- language-action model for dynamic object manipulation.arXiv preprint arXiv:2601.22153, 2026
arXiv 2026
-
[29]
Black, M
K. Black, M. Galliker, and S. Levine. Real-time execution of action chunking flow policies. Advances in Neural Information Processing Systems, 38:33383–33407, 2026
2026
- [30]
- [31]
-
[32]
Y . Lu, Z. Liu, X. Fan, Z. Yang, J. Hou, J. Li, K. Ding, and H. Zhao. Faster: Rethinking real-time flow vlas, 2026. URLhttps://arxiv.org/abs/2603.19199. 12
Pith/arXiv arXiv 2026
-
[33]
S. Ye, Y . Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y . L. Tan, C. Zhu, J. Xiang, et al. World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026
Pith/arXiv arXiv 2026
-
[34]
M. J. Kim, Y . Gao, T.-Y . Lin, Y .-C. Lin, Y . Ge, G. Lam, P. Liang, S. Song, M.-Y . Liu, C. Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning. InICLR, 2026
2026
-
[35]
S. Li, Y . Gao, D. Sadigh, and S. Song. Unified video action model. InRSS, 2025
2025
-
[36]
C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta. Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets. InRSS, 2025
2025
-
[37]
Y . Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling.arXiv preprint arXiv:2210.02747, 2022
Pith/arXiv arXiv 2022
-
[38]
F. Zhang and M. Gienger. Affordance-based robot manipulation with flow matching.arXiv preprint arXiv:2409.01083, 2024
arXiv 2024
-
[39]
K. Song, B. Chen, M. Simchowitz, Y . Du, R. Tedrake, and V . Sitzmann. History-guided video diffusion.arXiv preprint arXiv:2502.06764, 2025
Pith/arXiv arXiv 2025
-
[40]
J. Tang, Y . Sun, Y . Zhao, S. Yang, Y . Lin, Z. Zhang, J. Hou, Y . Lu, Z. Liu, and S. Han. Vlash: Real-time vlas via future-state-aware asynchronous inference.arXiv preprint arXiv:2512.01031, 2025
Pith/arXiv arXiv 2025
-
[41]
K. Zakka, B. Tabanpour, Q. Liao, M. Haiderbhai, S. Holt, J. Y . Luo, A. Allshire, E. Frey, K. Sreenath, L. A. Kahrs, et al. Mujoco playground.arXiv preprint arXiv:2502.08844, 2025
Pith/arXiv arXiv 2025
-
[42]
Perez, F
E. Perez, F. Strub, H. De Vries, V . Dumoulin, and A. Courville. Film: Visual reasoning with a general conditioning layer. InProceedings of the AAAI conference on artificial intelligence, volume 32, 2018. 13 A Implementation Details A.1 Simulation Task.The Leap Cube Reorientation task in MuJoCo Playground [41] requires a16-DoF Leap Hand to orient a cube w...
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.