{"id":"226c8f99-6c81-42cb-a34b-57549412720d","arxiv_id":"2603.15620","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"DOMINO supplies 110K+ dynamic dual-arm trajectories across 35 tasks, and PUMA’s optical-flow history plus object-centric future queries raise dynamic success rate by 6.3 points over strong VLA baselines.","lead":"Robots that grab moving objects fail when trained only on still scenes. This paper releases a large dynamic-manipulation benchmark (DOMINO) and a VLA model (PUMA) that uses past motion and short-horizon object forecasts, lifting success by about 6 points in simulation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Main SOTA claim rests on Level-1 kinematic sim only; absolute SR remains low and higher-order dynamics are largely untested.","rationale":"The reader already flags the same soft spot (kinematic Level-1 sim, low absolute SR, no real-robot data) as the reason for CONDITIONAL rather than ACCEPT. My stress-test simply sharpens the load-bearing point: the quantitative SOTA number that anchors the strongest claim is measured only in the regime where motion is perfectly predictable and contact-free. Because the paper itself demonstrates rapid degradation on higher-order dynamics for a simpler policy, the absence of the same measurement for PUMA is the single most consequential missing check. The architectural ablations (Table 5) and transfer results remain internally consistent and useful; they simply do not yet underwrite a general claim of “dynamics-aware” superiority. Keeping the verdict CONDITIONAL with high confidence is therefore appropriate; no stronger rejection is warranted given the clear empirical gains inside the stated setting and the promised open-source release.","tokens_in":28651,"tokens_out":598,"duration_ms":6700,"concrete_test":"Retrain and evaluate the full PUMA configuration (N=4, Hist. Flow + Aux. Pred.) on the same five tasks used for Fig. 4 under Level-2 and Level-3 dynamics at α=0.1; if the absolute SR gain over the Qwen3-VL OpenVLA-OFT baseline shrinks below 3 pp or absolute SR falls below 8%, the headline SOTA claim does not generalize beyond Level-1 kinematics.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (PUMA yields a 6.3 pp absolute SR gain, 17.20% vs ~10.9% Qwen3-VL OpenVLA-OFT) is measured almost exclusively under the easiest regime: Level-1 constant-velocity kinematic bodies at α=0.1 on Aloha-AgileX (explicit experimental setup §4 and Finding 1). Objects are instantiated as kinematic bodies immune to unintended physical disturbances (§2.2), so contact, slip, and reaction forces never appear. Fig. 4 already shows ACT SR collapsing from ~48% MS / 34% SR on Level 1 to ~13% MS / 4% SR on Level 3; no corresponding multi-level or multi-embodiment numbers are reported for PUMA itself. Consequently the 6.3 pp gain may reflect improved linear extrapolation under perfect motion rather than the claimed “genuine short-horizon object dynamics understanding.” The absolute success rate of 17% further indicates that even the best model still fails on the large majority of episodes, so the practical significance of the architectural advance remains open.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that current Vision-Language-Action (VLA) models fail on dynamic manipulation primarily because of scarce dynamic data and single-frame architectures that lack spatiotemporal reasoning. It introduces DOMINO, a SAPIEN/RoboTwin-based benchmark of 35 dual-arm tasks organized into three hierarchical dynamic levels (constant-velocity, high-order polynomial, and stochastic/abrupt), parameterized by a speed coefficient α, with >110K expert trajectories and a multi-metric suite (Success Rate plus Manipulation Score that combines route completion with safety penalties). On this benchmark the authors systematically evaluate existing VLAs, show large static-to-dynamic degradation, and propose PUMA: a Qwen3-VL-based architecture that injects scene-centric historical Farneback optical flow and trains specialized world queries with a cosine-similarity loss against frozen DINOv2 features of GroundingDINO+SAM2-masked future frames (train-only). Reported results claim a 6.3 pp absolute SR gain (17.20 % vs. ~10.9 % for a Qwen3-VL OpenVLA-OFT re-implementation) under Level-1 α=0.1 conditions, plus evidence that dynamic training transfers to static tasks and that co-training further improves dynamic performance.","tokens_in":29082,"tokens_out":1183,"duration_ms":10424,"significance":"If the claims hold, the work supplies both a missing evaluation infrastructure and a concrete architectural recipe for endowing VLAs with short-horizon object dynamics. DOMINO’s hierarchical taxonomy, multi-embodiment coverage, and multi-dimensional metrics fill a genuine gap left by static-world simulators (RLBench, CALVIN, RoboTwin, etc.). The open release of code and data, the controlled ablations (optical flow vs. raw frames, prediction horizon N, co-training), and the oracle GT-trajectory probe are strengths that make the empirical findings reproducible and falsifiable. Even if absolute success rates remain modest, a well-characterized dynamic benchmark and a simple, train-only prediction head constitute a useful stepping stone for the community.","major_comments":[{"comment":"The headline 6.3 pp SR claim (Tab. 3, abstract) and all primary ablations (Tab. 5) are obtained exclusively under Level-1 constant-velocity kinematic bodies at α=0.1 on a single embodiment (Aloha-AgileX), as stated in §4 experimental setup. Objects are instantiated as kinematic bodies “immune to unintended physical disturbances” (§2.2), so contact, slip and reaction forces never appear. Fig. 4 already shows ACT collapsing from ~34 % SR on Level 1 to ~4 % on Level 3, yet no corresponding multi-level or multi-embodiment numbers are reported for PUMA itself. Consequently the central claim that the architecture induces “genuine short-horizon object dynamics understanding” rests on the easiest regime and may largely reflect improved linear extrapolation under perfect motion. At minimum the authors should report PUMA SR/MS on Levels 2–3 and on at least one additional embodiment, or explicitly","section":"§4, Tab. 3, Fig. 4"},{"comment":"Absolute success remains low (17.2 %). The oracle experiment (Tab. 2) shows that even perfect future trajectories raise SR only marginally while improving MS, and the authors themselves note residual control jitter. This raises the question whether the reported architectural gain is practically meaningful for real dynamic manipulation. A clearer discussion of remaining failure modes (timing, grasp stability under residual motion, dual-arm coordination) and of the gap between kinematic simulation and contact-rich dynamics would strengthen the paper’s claims about generalizability.","section":"Tab. 2, §4.2–4.3"}],"minor_comments":[{"comment":"The abstract and introduction repeatedly state “6.3 % absolute improvement” without naming the precise baseline (Qwen3-VL OpenVLA-OFT re-implementation). Adding the baseline number in the abstract would avoid ambiguity.","section":"Abstract"},{"comment":"Notation for the dynamics coefficient α and the three hierarchical levels is introduced in §2.3 but used inconsistently later (sometimes “DOMINO@0.1”, sometimes “Level 1”). A short glossary or consistent subscript would help.","section":"§2.3"},{"comment":"Optical-flow resolution is fixed at 64×64 and Farneback parameters are not listed; a one-sentence statement of the exact OpenCV settings (or a pointer to the released code) would improve reproducibility.","section":"§3.1, Supp. C.1"},{"comment":"Several tables in the supplement (Tabs. 7–11) contain minor formatting inconsistencies (e.g., missing percent signs, truncated task names). These do not affect the main claims but should be cleaned.","section":"Supplementary Tables"}],"recommendation":"major_revision","confidential_remarks":"The work is a solid empirical contribution that belongs in a strong CV/robotics venue once the evaluation scope is either broadened or clearly caveated. The kinematic-body limitation is the single load-bearing issue; if the authors can supply Level-2/3 and multi-embodiment numbers for PUMA (even on a subset of tasks), the paper becomes much stronger. I see no circularity or novelty-disclosure problems."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful part of this paper is DOMINO. They actually built the missing piece: 35 dual-arm tasks, three explicit dynamics tiers, 110K+ trajectories, multi-embodiment, and a continuous MS/RC metric that is more informative than binary SR. That alone is worth having. The systematic static-to-dynamic collapse numbers (π0.5 44.8% → 7.5%, etc.) and the transfer/co-training results (dynamic data helps static, mixing helps more) are clean and useful.\n\nPUMA is a sensible engineering response: Farneback flow on compressed history + train-only world queries supervised by cosine to DINO features of GroundingDINO+SAM2-masked future frames. Ablations line up (flow beats raw frames; N=4 best; +4.9 pp from co-training). Against the Qwen3-VL OpenVLA-OFT reimplementation they get 17.2% vs 10.9% SR. Math is ordinary L1 + cosine, no circularity, code/data promised.\n\nThe stress-test concern is fair and already visible in the paper. Main tables are Level-1 α=0.1 kinematic bodies on Aloha-AgileX; objects ignore contact forces. Fig. 4 already shows ACT collapsing on Level 2/3 and they never report the same multi-level numbers for PUMA. Absolute SR of 17% means most episodes still fail. So the architectural claim is “better linear extrapolation under perfect motion,” not yet “genuine short-horizon dynamics understanding under contact.” That is a limitation of scope, not a hidden flaw in the reported numbers.\n\nWho it is for: anyone building or evaluating VLAs that claim real-world reactivity. The benchmark will get used; the architecture is a clear, reproducible baseline. Soft spots are the usual sim-to-real and easy-regime caveats, stated proportionately. I would send it to referees; they will ask for Level-2/3 and multi-embodiment PUMA numbers, which is the right ask. Worth reading and citing for the dataset and the transfer findings.","headline":"Solid new dynamic-manipulation benchmark plus a clean history+prediction VLA recipe; the 6.3 pp SOTA claim is real but sits almost entirely on easy Level-1 kinematic sim.","tokens_in":29649,"tokens_out":542,"would_cite":true,"duration_ms":6702,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"History-aware optical flow plus short-horizon object prediction raises dynamic manipulation success by 6.3 points and transfers to static tasks.","keywords":["dynamic manipulation","vision-language-action","optical flow","object-centric prediction","embodied AI","robotics benchmark","spatiotemporal reasoning"],"falsifier":"Remove the world-query loss and replace historical optical flow with raw frames (or deploy the same model on real robots whose objects move freely under contact); if success on Level-1 DOMINO@0.1 falls back to the single-frame baseline, the claim that the architecture induces useful dynamics awareness is false.","tokens_in":29557,"feed_emoji":"🦾","tokens_out":949,"duration_ms":16995,"temperature":0.7,"pith_summary":"Vision-language-action models succeed when objects sit still but collapse once targets move, mainly because they see only single frames and almost no dynamic training data exist. This paper builds DOMINO—a 35-task, 110K-trajectory benchmark with three motion-complexity levels and multi-metric scoring—and shows that strong VLAs lose most of their static success rate under those conditions. It then introduces PUMA, which feeds the model compressed historical optical flow and trains specialized world queries to anticipate object-centric features of future frames. The architecture reaches 17.2 percent success, a 6.3-point absolute gain over re-implemented baselines, and policies trained only on dynamic data transfer zero-shot to static scenes. A reader who wants robots that can work on conveyors, assembly lines, or next to people should care: static-only methods are not enough, and dynamic data plus explicit history-plus-prediction is a concrete path forward.","feed_headline":"History plus prediction lifts moving-target robot success 6.3 points","feed_subtitle":"DOMINO benchmark and PUMA show dynamic training also transfers to static tasks","key_machinery":"PUMA: a VLA that encodes history as Farneback optical-flow maps on compressed multi-view frames and uses learnable world queries that, training-only, are forced to match future object-centric DINO features extracted via GroundingDINO+SAM2; the shared backbone thereby learns short-horizon object dynamics that regularize action-chunk prediction.","core_discovery":"Single-frame VLAs suffer large success-rate drops when objects move; fine-tuning them on dynamic data recovers only a few points. Coupling scene-centric historical optical flow with an auxiliary object-centric future-feature objective (world queries supervised by cosine similarity to DINO features of masked future frames) yields a dynamics-aware policy that reaches 17.2 percent success on DOMINO@0.1—6.3 points above the strongest baseline—and produces representations that transfer to static tasks.","pith_inferences":["If the same optical-flow-plus-world-query recipe holds under real contact-rich dynamics, the kinematic-body simulation gap for short-horizon interception may be smaller than often assumed.","Raw historical frames hurt while explicit flow helps, suggesting future VLAs should treat motion fields as first-class inputs rather than hoping transformers discover them.","Extending the prediction horizon beyond N=4 or closing the loop with the same world queries is a direct next test for the remaining Level-2/3 failures.","The two-stage dry-run then kinematic back-calculation pipeline could convert many existing static datasets into dynamic counterparts at low cost."],"forward_implications":["Single-frame VLAs will remain unreliable on any task whose objects move independently of the robot until they incorporate explicit history and short-horizon prediction.","Dynamic training data produces spatiotemporal features that transfer zero-shot to static scenes, so dynamic data is not a narrow specialization.","Co-training static and dynamic data further raises dynamic success, indicating complementary structural and reactive priors.","Hierarchical motion levels (constant velocity, high-order curves, abrupt stochastic segments) and multi-metric scoring become necessary for evaluating reactive policies."],"fun_headline_variants":["PUMA's history flow plus future queries lift dynamic robot success 6.3 points","DOMINO shows single-frame VLAs fail on movers; PUMA recovers with spatiotemporal forecasti","Dynamic optical-flow history and world queries yield 17.2% success on moving-target tasks","Training VLAs on DOMINO dynamic data also boosts static manipulation transfer","Scene history plus short-horizon prediction closes dynamic VLA gap by 6.3 points"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Matching frozen visual features of future object masks, together with low-resolution optical flow of scripted kinematic trajectories, is enough to teach genuine short-horizon physical anticipation rather than shallow motion tracking.","fun_headline_variants_meta":{"raw":{"variants":["PUMA's history flow plus future queries lift dynamic robot success 6.3 points","DOMINO shows single-frame VLAs fail on movers; PUMA recovers with spatiotemporal forecasting","Dynamic optical-flow history and world queries yield 17.2% success on moving-target tasks","Training VLAs on DOMINO dynamic data also boosts static manipulation transfer","Scene history plus short-horizon prediction closes dynamic VLA gap by 6.3 points"]},"model":"grok-4.5","effort":"low","cost_usd":0.002878,"raw_usage":{"total_tokens":1073,"prompt_tokens":794,"num_sources_used":0,"completion_tokens":98,"cost_in_usd_ticks":28780000,"prompt_tokens_details":{"text_tokens":794,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":181,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":794,"tokens_out":98,"duration_ms":3809,"temperature":1.0,"reasoning_tokens":181,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T00:12:56.453428+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Remove the world-query loss and replace historical optical flow with raw frames (or deploy the same model on real robots whose objects move freely under contact); if success on Level-1 DOMINO@0.1 falls back to the single-frame baseline, the claim that the architecture induces useful dynamics awareness is false.","supporting_citations":[],"review_version":2}