Pith. sign in

REVIEW 3 major objections 5 minor 41 references

πR² makes a pretrained VLA replan closed-loop at ~25 Hz by splitting fresh proprioception from stale vision-language features and emitting actions in one denoising step per call.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 00:44 UTC pith:LKOKXEE6

load-bearing objection Core idea is plausible and the sim study is solid, but the headline 4x speedup is confounded by a 2-GPU vs 1-GPU comparison and the real-world evidence is too thin to take the 30% gains at face value. the 3 major comments →

arxiv 2607.26055 v1 pith:LKOKXEE6 submitted 2026-07-28 cs.RO cs.AIcs.LG

πR²: Reactive Real-time Flow Policies

classification cs.RO cs.AIcs.LG
keywords reactive controlflow matchingaction chunkingdiffusion forcingvision-language-action modelsproprioceptionlatency adaptationdexterous manipulation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

πR² sets out to restore closed-loop reactivity to action-chunking flow policies, the dominant architecture for robotic foundation models, without giving up large pretrained backbones, multi-modal action distributions, or chunk-based training. Its central proposal is to decouple the policy's conditioning into a fast channel (proprioception, refreshed every control tick) and a slow channel (vision and language features, refreshed asynchronously in the background), and to pair this with a latency-adaptive noise schedule in which in-flight actions are treated as inpainting conditioning and d clean actions are emitted per single denoising step. If correct, the paper shows that a large VLA can be fine-tuned (with a frozen backbone) to replan at about 25 Hz on a single GPU — roughly 4× faster than the base pipeline — and that this recovers reactivity in contact-rich tasks, raising success by up to 23% in simulation and 30% in the real world over the strongest baseline. The sympathetic reading is that expensive semantic perception can be relegated to a slow background thread because dynamic manipulation reacts through the body's own sensors.

Core claim

The paper's claim is that the two bottlenecks of chunked flow policies — open-loop chunks and slow perception-to-action latency — are removable by construction. πR² assigns each position p of the action chunk its own noise level via a three-region staircase τ⋆,d: the front d slots are clamped clean (in-flight actions already sent to the robot), the interior ramps linearly from clean to noise, and the tail holds d pure-noise slots. One Euler denoising step per policy call advances the schedule by d slots, so d clean actions are released at the front while the freshly emitted actions become the new inpainting conditioning; the per-call cost is a single function evaluation of the action head. T

What carries the argument

The load-bearing object is the delay-parameterized staircase noise schedule τ⋆,d (Eq. 3), a three-region per-position noise profile: a clamped-clean front that treats in-flight actions as inpainting conditioning, a linear ramp over the interior, and a pure-noise tail. Combined with the asynchronous fast/slow conditioning split (fresh proprioception, cached VLM features with learned delay embedding), it lets one denoising step per call emit d clean actions and makes the model adaptive to measured hardware latency, because training samples d uniformly and the schedule reproduces itself after each slide.

Load-bearing premise

The whole reactivity gain rests on the claim that fresh joint angles, torques, and fingertip forces carry enough information about the immediate situation that the action head can ignore vision and language for up to five control ticks (200 ms); if that claim fails, the asynchronous slow channel adds nothing and πR² is just a latency-adaptive schedule running on stale observations.

What would settle it

On the real Catch Book task, keep the staircase schedule but replace the cached visual feature with one from d_vis = 8 ticks (320 ms) earlier while proprioception stays fresh; if the gripper consistently fails to close on the falling book, the proprioceptive channel alone was not carrying the reaction.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Any pretrained VLA with a flow-matching action head can be fine-tuned into this reactive mode with a one-line DiT change (per-position AdaLN) and a frozen backbone, preserving the semantic grounding of the base model.
  • Actions are no longer committed for a full chunk: each 40 ms tick conditions the emitted action on fresh joint angles, torques, and fingertip forces, so contact events can alter the motion mid-chunk.
  • A single trained model adapts to varying hardware and network latency at deployment, because the schedule is parameterized by the measured per-call delay d.
  • Vision-language staleness up to about 200 ms (d_vis=5 ticks) is tolerated without losing task success, indicating that the slow channel provides coarse guidance while the fast channel drives local correction.
  • In the simulation deployment study, πR²'s advantage over naive asynchronous execution and training-time action conditioning widens as the underlying latency budget grows.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • (Editorial extension) The same slow/fast split should transfer to any expensive conditioning stream, such as video tokens or world-model state predictions, whenever the cheap stream (proprioception or joint encoders) plausibly carries the high-frequency information.
  • (Editorial extension) A sharp test of the paper's central premise: on a contact-rich task, feed the final denoising step a deliberately wrong visual feature (e.g., a frame from a different episode) and keep proprioception fresh; if success degrades sharply, proprioception alone is not sufficient for local corrections and the async split is doing less work than claimed.
  • (Editorial extension) The method's gains should scale with how contact-rich and time-critical the task is; on quasi-static pick-and-place tasks the 4× latency reduction may not translate into measurable success changes, a boundary case the paper does not explore.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes πR², a method to make flow-matching VLA policies reactive by (1) splitting conditioning into a fast proprioceptive channel and a slow vision-language channel updated asynchronously, and (2) using a delay-parameterized staircase noise schedule (Eq. 3) that emits clean actions from a diffusion-forcing buffer with one denoising step per call. The authors fine-tune GR00T-N1.7 on a real xArm6+XHand and report ~25 Hz closed-loop replanning (4× speedup over the base pipeline), with success-rate gains up to 23% in simulation and up to 30% in the real world. They also present simulation ablations (execution-horizon sweep, effective-delay study, vision-staleness sweep) to support the mechanism.

Significance. The paper tackles an important and timely problem: large VLA action-chunking policies are too slow for reactive closed-loop control. The two ideas—asynchronous fast/slow conditioning and a latency-adaptive diffusion-forcing schedule—are simple, architecture-agnostic, and compatible with pretrained VLAs. The simulation study is well-designed (controlled delays, d_vis ablation, h-sweep), and the real-world tasks are genuinely reactive. If the performance claims survive an equal-hardware comparison, this would be a useful contribution. However, the headline real-world speedup is currently confounded by the use of 2 GPUs for πR² and 1 GPU for baselines, and the zero-delay schedule is under-specified.

major comments (3)
  1. [§4.2, Table 1, Appendix A.2] The headline real-world speedup (abstract: ~4× faster, ~25 Hz; §4.2: d=1–2 vs d=4–5) is measured with πR² on 2×RTX A5000 and all baselines on 1×RTX A5000. Appendix A.2 states this explicitly and argues that sharing one GPU 'would conflate the measured d,' but no experiment controls for the extra hardware. The 25 Hz claim and the Table 1 success gains could therefore be due to the additional GPU, not to the proposed schedule/async split. Please provide an equal-hardware comparison (e.g., πR² with both workers on one A5000, or baselines with the VLM on a second A5000) and report the measured control-tick delay for each configuration.
  2. [§3.3 vs §4.1.1] The inference cycle is described as shifting the schedule right by d slots and emitting positions [d,2d), so d clean actions are released per call and the buffer slides by d. For d=0, Eq. (3) gives an empty front and tail and a degenerate slide, yet §4.1.1 evaluates 'πR² ... emits one action per call (h=1)' in a zero-delay setting. The paper does not specify the step-size rule that produces exactly one emitted action in this case. This matters because the zero-delay h-sweep is used to claim that one-step amortized denoising matches h=1 flow. Please clarify the d=0 schedule/protocol, or treat the experiment as d=1 and adjust the interpretation.
  3. [§4.2 and Table 1] The real-world evaluation does not include an ablation that removes the asynchronous fast/slow split. The success-rate gains over Train-Time RTC (up to 30%) are attributed to proprioceptive reactivity, but the only direct evidence for this mechanism is the simulation d_vis sweep with a low-dimensional vision proxy. The real-world policy uses rich VLM features and allows up to d_vis=5 ticks (200 ms) of visual staleness during training, a regime not covered by the sim. Please add a real-world run of 'πR² w/o async' (or a d_vis ablation) to confirm that the async split—rather than the schedule or the additional GPU—is responsible for the real-world improvement.
minor comments (5)
  1. [§4.2] Typo: 'Tmaihe same pattern holds' should be 'The same pattern holds.'
  2. [§3.2 vs Appendix A.2] Notation is inconsistent: §3.2 uses d_vlm for the vision-language staleness, while §4.1.2 and Appendix A.2 use d_vis. Please unify.
  3. [Abstract and §4.2] The abstract says '~25Hz on an A5000 GPU,' but deployment uses 2×A5000. Please specify the hardware count in the abstract or change the wording to avoid implying a single A5000.
  4. [Table 1] The Catch Book columns list only SR and no Prog, even though the task is described as having one subgoal. Please add the Prog column or explain the omission.
  5. [Fig. 3(right) caption] The caption says 'for each datapoint, d indicates the effective delay of proprioception, while d_v is visual delay,' but the text does not specify which d_vis values correspond to the plotted w/ async curve. Please clarify whether the curve is a specific d_vis value or an average over the d_vis sweep.

Circularity Check

0 steps flagged

No significant circularity: πR2's success claims are empirically measured, its schedule is a stated design choice, and prior-work citations are external and non-load-bearing.

full rationale

The paper's central quantitative claims (25 Hz replanning, ~4x speedup, success improvements) are direct measurements of wall-clock latency and task outcomes, not derived from assumptions that already contain the results. Eq. 3 introduces the staircase schedule as a design choice and explicitly frames it as a diffusion-forcing generalization of external prior work (Train-Time RTC, streaming diffusion), rather than as a prediction fitted to the evaluation. Eq. 4's one-step emission property follows from the chosen per-position Δτ advances, and the paper clearly labels this as construction. The simulation deployment study assigns effective delays using a stated computation-cost model and then measures success; the real-world experiments measure actual delay ticks (d=4–5 for baselines, d=1–2 for πR2) and report success rates. No parameter is fitted to the evaluation and then renamed a prediction. The only author-overlapping citation (Open X-Embodiment, a large collaboration) supports a background statement about VLA scaling and is not load-bearing; no uniqueness theorem or ansatz is imported from the authors' own prior work. The deployment hardware imbalance (πR2 on two GPUs vs baselines on one) is a benchmarking fairness issue that the paper discloses, but it is not a circularity of the kind where an output is equivalent to an input by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The ledger shows the central claim depends primarily on the domain assumption that proprioception suffices for local control while vision can be stale, and on the empirical validity of a large single Euler step. No new entities are postulated. The unchanged base architecture means the method inherits the pretrained backbone's assumptions without adding independent evidence for them.

free parameters (4)
  • d_max (training max delay) = 5 (sim and real); 10 for RTC baseline
    Chosen by hand; defines the range of delays sampled at training (Sec. 3.3, Tables 2/5). The schedule is only validated up to this value; deployment delays beyond d_max are out of distribution.
  • α (standard-flow warm-up probability) = 0.2
    Chosen by hand (Tables 2/5); controls how often the network trains on standard flow to enable warm-starting the buffer. No ablation showing sensitivity.
  • j (symmetric jitter magnitude) = not reported
    Used in Alg. 1 line 9 (τ_p ← clip(τ_p + δ_p, 0, 1), δ_p ~ Uniform[−j, j]) to absorb per-call delay variation, but its numerical value never appears in the paper, impeding replication and leaving the schedule's robustness claim partially underspecified.
  • d_max^vis (max vision delay) = 5 ticks (200 ms at 25 Hz)
    Chosen by hand (Sec. A.2); sets the training range for the slow-channel delay embedding. The real-world deployment d_vis presumably falls in this range, but the paper does not report the measured distribution.
axioms (5)
  • standard math Standard flow matching objective and conditional interpolation (Eq. 1) are valid for action-chunking policies.
    Taken from Lipman et al. [37]; used throughout.
  • standard math Per-position noise scheduling as in diffusion forcing (Eq. 2) preserves the training objective when noise levels differ across chunk positions.
    From Chen et al. [18]; the paper extends but does not re-derive this.
  • domain assumption For dynamic contact-rich tasks, fresh proprioception carries sufficient information for local reactive corrections, while stale vision-language features up to ~200 ms provide adequate global context.
    Stated in Sec. 3.2 ('proprioception carries sufficient information for local reactive corrections') and tested only via the sim d_vis sweep; it is the load-bearing assumption for the fast/slow split.
  • domain assumption A single Euler step of per-position size up to d/(H−2d) produces accurate enough clean actions for control; the velocity field is smooth enough for this large step.
    The schedule (Eq. 3) is designed so one NFE per call emits d actions; the paper validates empirically (Sec. 4.1.1) but provides no bound or analysis.
  • domain assumption The measured deployment delay d and vision staleness d_vis are within the training ranges and the learned delay embedding transfers from zero-init to real measured delays.
    Training samples d ~ Uniform{1,...,d_max} and d_vis ~ Uniform{0,...,d_max^vis} (Alg. 1); at deployment the same embeddings are used (Alg. 2). No evidence that out-of-range delays work.

pith-pipeline@v1.3.0-alltime-deepseek · 18450 in / 18613 out tokens · 161372 ms · 2026-08-01T00:44:02.037915+00:00 · methodology

0 comments
read the original abstract

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies ill-suited for dynamic, closed-loop control. We present $\pi\mathbf{R}^2$, which makes these policies reactive and real-time while retaining large backbones, expressive multi-modal policies, and multi-action prediction. Built on the per-position noise schedule of diffusion forcing, $\pi\mathbf{R}^2$ contributes two ideas. First, it splits conditioning into a fast channel (proprioception, fresh every tick) and an asynchronously updated slow channel (vision-language features), so the policy reacts to proprioception within a chunk while tolerating stale vision. Second, a latency-adaptive flow schedule treats in-flight actions as inpainting conditioning and emits actions in one denoising step per call, letting one trained model adapt to varying hardware latency. Requiring minimal modification to existing architectures, $\pi\mathbf{R}^2$ can be finetuned from a pretrained policy: applied to GR00T-N1.7 on a real xArm6+XHand platform, it replans closed-loop roughly $4\times$ faster than the base policy (~$25$Hz on an A5000 GPU), acting on a fresh observation every $40$ms. Across simulation and real-world manipulation tasks, $\pi\mathbf{R}^2$ improves the success rate by up to $23\%$ in simulation and $30\%$ in the real world over the strongest baseline. Project page: https://pi-r2-flow.github.io/

Figures

Figures reproduced from arXiv: 2607.26055 by Shubham Tulsiani, Sungjae Park.

Figure 1
Figure 1. Figure 1: Overview of πR2 . Top. While standard diffusion/flow matching relies on stale observa￾tions to predict actions via iterative denoising, πR2 disentangles the observation into a fast channel (proprioception) and a slow channel (image encoding, VLM embedding, etc.), and uses up-to-date observations for each denoising step. Bottom. To incorporate latency and smooth execution, we adopt an adaptive noise schedul… view at source ↗
Figure 2
Figure 2. Figure 2: Inference cycle, one call. (a) Buffer at τ ⋆,d at the start of the call. (b) One Euler substep (one NFE) shifts the schedule right by d slots: positions [d, 2d) reach τ=1 and are released, and the ramp extends through the back (c) The buffer slides d positions: just-emitted actions become the new front conditioning, d fresh-noise slots (τ=0) are appended – the schedule reproduces exactly. the front of the … view at source ↗
Figure 3
Figure 3. Figure 3: Simulation results. Left. Without any inference delay, executing a smaller number of actions within the chunk benefits the performance for a flow matching policy. πR2 also achieves the same performance via replanning every timestep, while the only last denoising step is conditioned on up-to-date state. Right. When there is an inference delay for each component, πR2 reduces the effective delay d by reducing… view at source ↗
Figure 4
Figure 4. Figure 4: Real-world manipulation tasks. Arrows in each image show the desired task outcome [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: πR2 reacts to proprioception, while baselines run a stale plan (Tidy Up Book). Fin￾gertip force (solid, left axis) and the emitted action (dashed, right axis) over time, for πR2 and Train-Time RTC; numbered markers link plot times to the overhead frames above. πR2 grips just enough (∼50 N), while RTC reacts late and over-grips to ∼120 N, dropping the book. across all metrics, with the largest gains on reac… view at source ↗
Figure 6
Figure 6. Figure 6: Leap Cube Reorientation goals. The Leap Hand grasps a cube and must rotate it until the blue face points up at one of four target yaw poses (0 ◦ , 90◦ , 180◦ , 270◦ ). The bottom caption shows the corresponding goal quaternions (w, x, y, z). Dataset. Four PPO experts (different seeds) generate 50 trajectories each, 200 total at 50 Hz. The simulator injects per-step sensing noise into the observation, sampl… view at source ↗
Figure 7
Figure 7. Figure 7: Real-world workspace. xArm6 + XHand with a single overhead RGB camera. Hardware and Workspace. A 6-DoF xArm6 manipulator car￾ries a 12-DoF XHand at its wrist; a single overhead 640 × 480 RGB camera looks down on the workspace from above ( [PITH_FULL_IMAGE:figures/full_fig_p016_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Reactivity on the remaining tasks. Fingertip force (solid, left axis) and the emitted action (dashed, right axis) over time for πR2 and Train-Time RTC, with overhead frames at the marked times. Top: Catch Book. Middle: Insert Box. Bottom: Don’t Spill. 20 [PITH_FULL_IMAGE:figures/full_fig_p020_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

41 extracted references · 21 linked inside Pith

  1. [1]

    Black, N

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al.π0: A vision-language-action flow model for general robot control. InRSS, 2025

  2. [2]

    Intelligence, K

    P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. V...

  3. [4]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. In RSS, 2023

  4. [5]

    Zitkovich, T

    B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In CoRL, 2023

  5. [6]

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model. InCoRL, 2024

  6. [7]

    H. Liu, C. Li, Q. Wu, and Y . J. Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

  7. [8]

    Alayrac, J

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Mil- lican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

  8. [9]

    S. Bai, Y . Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025

  9. [10]

    G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805, 2023

  10. [11]

    Beyer, A

    L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Al- abdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Boˇsnjak, X. Chen, M. Minderer, P. V oigtlaender, I. Bica, I. Balazevic, J. Puigcer...

  11. [12]

    Deitke, C

    M. Deitke, C. Clark, S. Lee, R. Tripathi, Y . Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, J. Lu, T. Anderson, E. Bransom, K. Ehsani, H. Ngo, Y . Chen, A. Patel, M. Yatskar, C. Callison-Burch, A. Head, R. Hendrix, F. Bastani, E. VanderBilt, N. Lambert, Y . Chou, A. Chheda, J. Sparks, S. Skjonsberg, M. Schmitz, A. Sarnat, B. Bischoff, P...

  12. [13]

    C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. InIJRR, 2023

  13. [14]

    C. Yuan, C. Wen, T. Zhang, and Y . Gao. General flow as foundation affordance for scalable robot learning.arXiv preprint arXiv:2401.11439, 2024

  14. [15]

    C. Pan, G. Anantharaman, N.-C. Huang, C. Jin, D. Pfrommer, C. Yuan, F. Permenter, G. Qu, N. Boffi, G. Shi, et al. Much ado about noising: Dispelling the myths of generative robotic control.arXiv preprint arXiv:2512.01809, 2025

  15. [16]

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware.arXiv preprint arXiv:2304.13705, 2023

  16. [17]

    T. T. Zhang, D. Pfrommer, C. Pan, N. Matni, and M. Simchowitz. Action chunking and ex- ploratory data collection yield exponential improvements in behavior cloning for continuous control.arXiv preprint arXiv:2507.09061, 2025

  17. [18]

    B. Chen, D. Mart ´ı Mons ´o, Y . Du, M. Simchowitz, R. Tedrake, and V . Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion.Advances in Neural Information Processing Systems, 37:24081–24125, 2024

  18. [19]

    S. H. Høeg, Y . Du, and O. Egeland. Streaming diffusion policy: Fast policy synthesis with variable noise diffusion models.arXiv preprint arXiv:2406.04806, 2024

  19. [20]

    Z. Fu, T. Z. Zhao, and C. Finn. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation, 2024. URLhttps://arxiv.org/abs/2401.02117

  20. [21]

    T. Z. Zhao, J. Tompson, D. Driess, P. Florence, K. Ghasemipour, C. Finn, and A. Wahid. Aloha unleashed: A simple recipe for robot dexterity.arXiv preprint arXiv:2410.13126, 2024

  21. [22]

    Zhang, Z

    Q. Zhang, Z. Liu, H. Fan, G. Liu, B. Zeng, and S. Liu. Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow matching for robot manipulation, 2024. URLhttps: //arxiv.org/abs/2412.04987

  22. [23]

    G. Yan, J. Zhu, Y . Deng, S. Yang, R.-Z. Qiu, X. Cheng, M. Memmel, R. Krishna, A. Goyal, X. Wang, and D. Fox. Maniflow: A general robot manipulation policy via consistency flow training, 2025. URLhttps://arxiv.org/abs/2509.01819

  23. [24]

    O. X.-E. Collaboration, A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Her- zog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakr- ishna, A. W...

  24. [25]

    Khazatsky, K

    A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. InRSS, 2024

  25. [26]

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024

  26. [27]

    Bjorck, F

    NVIDIA, :, J. Bjorck, F. Casta ˜neda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, G. Wang, Z. Wang, J. Wang, Q. Wang, J. Xiang, Y . Xie, Y . Xu, Z. Xu, S. Ye, Z...

  27. [28]

    H. Xie, B. Wen, J. Zheng, Z. Chen, F. Hong, H. Diao, and Z. Liu. Dynamicvla: A vision- language-action model for dynamic object manipulation.arXiv preprint arXiv:2601.22153, 2026

  28. [29]

    Black, M

    K. Black, M. Galliker, and S. Levine. Real-time execution of action chunking flow policies. Advances in Neural Information Processing Systems, 38:33383–33407, 2026

  29. [30]

    Black, A

    K. Black, A. Z. Ren, M. Equi, and S. Levine. Training-time action conditioning for efficient real-time chunking.arXiv preprint arXiv:2512.05964, 2025

  30. [31]

    Jiang, X

    S. Jiang, X. Fang, N. Roy, T. Lozano-P ´erez, L. P. Kaelbling, and S. Ancha. Streaming flow policy: Simplifying diffusion/flow-matching policies by treating action trajectories as flow trajectories, 2025. URLhttps://arxiv.org/abs/2505.21851

  31. [32]

    Y . Lu, Z. Liu, X. Fan, Z. Yang, J. Hou, J. Li, K. Ding, and H. Zhao. Faster: Rethinking real-time flow vlas, 2026. URLhttps://arxiv.org/abs/2603.19199. 12

  32. [33]

    S. Ye, Y . Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y . L. Tan, C. Zhu, J. Xiang, et al. World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026

  33. [34]

    M. J. Kim, Y . Gao, T.-Y . Lin, Y .-C. Lin, Y . Ge, G. Lam, P. Liang, S. Song, M.-Y . Liu, C. Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning. InICLR, 2026

  34. [35]

    S. Li, Y . Gao, D. Sadigh, and S. Song. Unified video action model. InRSS, 2025

  35. [36]

    C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta. Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets. InRSS, 2025

  36. [37]

    Lipman, R

    Y . Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling.arXiv preprint arXiv:2210.02747, 2022

  37. [38]

    Zhang and M

    F. Zhang and M. Gienger. Affordance-based robot manipulation with flow matching.arXiv preprint arXiv:2409.01083, 2024

  38. [39]

    K. Song, B. Chen, M. Simchowitz, Y . Du, R. Tedrake, and V . Sitzmann. History-guided video diffusion.arXiv preprint arXiv:2502.06764, 2025

  39. [40]

    J. Tang, Y . Sun, Y . Zhao, S. Yang, Y . Lin, Z. Zhang, J. Hou, Y . Lu, Z. Liu, and S. Han. Vlash: Real-time vlas via future-state-aware asynchronous inference.arXiv preprint arXiv:2512.01031, 2025

  40. [41]

    Zakka, B

    K. Zakka, B. Tabanpour, Q. Liao, M. Haiderbhai, S. Holt, J. Y . Luo, A. Allshire, E. Frey, K. Sreenath, L. A. Kahrs, et al. Mujoco playground.arXiv preprint arXiv:2502.08844, 2025

  41. [42]

    Perez, F

    E. Perez, F. Strub, H. De Vries, V . Dumoulin, and A. Courville. Film: Visual reasoning with a general conditioning layer. InProceedings of the AAAI conference on artificial intelligence, volume 32, 2018. 13 A Implementation Details A.1 Simulation Task.The Leap Cube Reorientation task in MuJoCo Playground [41] requires a16-DoF Leap Hand to orient a cube w...