{"id":"30c6605d-9956-43d7-ad84-fb0362d01f71","arxiv_id":"2505.09723","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"EnerVerse-AC generates realistic multi-view robot videos conditioned on action sequences and shows early evidence it can augment training data and rank policy performance like a real robot.","lead":"This paper builds a robot world model, EnerVerse-AC, that generates future camera views from an action sequence, and tests it as both a data generator and a policy evaluator. If it works at scale, robot policies could be tested and trained with less physical hardware.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluator validity is not established for policies whose action sequences leave the training/failure distribution; §4.5 shows hallucination when failure coverage is missing, and no OOD measurement is reported.","rationale":"I read the paper in good faith: the central claim is empirical, not derived, so the appropriate test is whether the reported evidence rules out the main failure mode. The reader's weakest assumption is correct and is the most load-bearing: EVAC can only be a reliable evaluator if it remains accurate on the action sequences emitted by the policy under test, not just on training or failure trajectories. The paper's own Section 4.5 demonstrates the exact hallucination risk when failure data is absent, so the evaluation claim depends on unquantified failure coverage. I also note that Figure 7's success-rate numbers are ambiguous as extracted; if the bars are interleaved by task as Real: 28/85/25/88 vs. EVAC: 100/55/90/50, the absolute discrepancies are 20-72 percentage points and the task-level ordering is not consistent, but if the values are grouped by legend the two sets are close. This ambiguity, plus the absence of per-task counts, correlation coefficients, confidence intervals, and inter-rater agreement, means the four-task evidence is too thin to support the generalization in the Abstract. However, the paper does provide qualitative multi-view results and acknowledges limitations, and no clear internal contradiction is apparent beyond the need to clarify the figure. Keeping the conditional verdict, with a required out-of-distribution stress test and exact per-task numbers, is the honest midpoint.","tokens_in":11289,"tokens_out":8250,"duration_ms":84546,"concrete_test":"Take the trained GO-1 policy and perturb its action chunks with increasing noise (e.g., 0%, 5%, 10%, 20% of the action range) across the four tasks; run 40 real-robot rollouts and 40 EVAC rollouts per noise level, with human raters blinded to source, and compute per-level success rates plus a Spearman correlation with 95% confidence intervals. Also report the mean per-step end-effector delta between perturbed and training actions to establish distributional coverage. If EVAC scores track real-world scores only at 0% noise, or diverge as noise increases, the evaluator is not valid for out-of-distribution policies; resolving Figure 7's ambiguous bar ordering and raw counts would also settle whether the current four-task evidence supports the claimed correlation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Abstract; §4.3) is that EVAC-generated video evaluations correlate with real-robot success. This requires EVAC to render physically faithful outcomes for the exact action chunks produced by the policy under test. The paper only evaluates GO-1 on four tasks (§4.3; Appendix A.4.1), all close to the AgiBot-World data used to train both EVAC and the policy. It never quantifies the distance between a tested policy's action distribution and EVAC's training distribution. Worse, §4.5 is an admission of the failure mode: without failure trajectories, EVAC 'hallucinates' a successful grasp when the gripper misses the bottle. The evaluator therefore has no demonstrated robustness for policies that produce atypical actions (e.g., noisy or partially trained policies), and the reported Figure 7 has no error bars, no correlation coefficient, and no inter-rater agreement, so the 'high correlation' claim rests on four unquantified task-level points. A plausible-looking but physically wrong generated video can mislead human raters exactly in the regime where an evaluator is most needed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces EVAC (EnerVerse-AC), an action-conditional video-diffusion world model for robot manipulation built on the EnerVerse architecture. The method injects robotic actions through spatial action maps and a delta-action attention module, extends generation to multiple camera views with ray map encodings, and trains on both successful and failure trajectories. EVAC is proposed for two applications: as a data engine that augments small human-collected demonstration sets with generated trajectories, and as a policy evaluator that scores policies from generated video instead of physical robot rollouts. The experiments include qualitative generation results, four task-level comparisons between EVAC-based and real-robot evaluation, an ablation of failure data, and a data-augmentation experiment reporting an increase in success rate from 0.28 to 0.36.","tokens_in":11484,"tokens_out":4138,"duration_ms":43111,"significance":"If the evaluator claim were quantitatively established, EVAC would be a practically valuable tool for reducing hardware costs in robot policy development. The paper has notable strengths: the authors promise code, checkpoints, and datasets; the qualitative figures show genuinely action-controllable and multi-view consistent generation; and the failure-data ablation in Section 4.5 is a concrete, honest demonstration of a known failure mode. However, the central quantitative claims are not backed by sufficient statistical evidence: there are no error bars or seed counts for the data-engine result, no objective video-quality metrics, no inter-rater agreement for the evaluator study, and no quantitative comparison against the EnerVerse baseline on which the method builds. The contribution is therefore plausible but not yet established at the standard claimed in the abstract.","major_comments":[{"comment":"The headline claim of a high correlation between EVAC-based evaluation and real-robot evaluation is not statistically supported. Figure 7 shows success rates for only four tasks and three training-step points, with no error bars, no confidence intervals, no correlation coefficient, and no inter-rater agreement metric even though the text states that three evaluators were used. Since each task was evaluated 40 times according to Appendix A.4.1, the authors should report bootstrap confidence intervals, a rank correlation across the task-level points, and an agreement statistic such as Cohen's kappa; without these, the claim is indistinguishable from chance agreement over four points.","section":"§4.3, Figure 7"},{"comment":"Evaluator validity is established only for the GO-1 policy on four tasks that are close to the AgiBot-World training distribution, and robustness to out-of-distribution action sequences is never measured. The evaluator feeds action chunks sampled from the policy under test into EVAC, but the paper does not quantify how far those actions are from EVAC's training action distribution. Section 4.5 demonstrates the concrete failure mode: when failure trajectories are absent, EVAC hallucinates a successful grasp even though no contact occurred. This is precisely the regime in which a hardware-free evaluator is most needed, such as for partially trained or noisy policies. An experiment that injects action noise or tests policies at varied training stages with reported action-distance diagnostics is necessary to support the generalization claim.","section":"§4.3, §4.5, Appendix A.4.1"},{"comment":"The data-engine result rests on a single pair of success rates (0.28 baseline versus 0.36 augmented) with no number of seeds, no standard deviation, no per-condition episode counts, and no significance test. With a difference of 0.08, random seed variation in policy training could easily explain the observed gap. The experiment should be repeated over multiple seeds with mean and variance reported, and the meaning of '30% additional trajectories' should be made precise, including the absolute number of generated trajectories used.","section":"§4.4, Table 1"},{"comment":"No objective video-quality metrics (such as FVD, LPIPS, action-follow error, or cross-view consistency) are reported for the generation claims; statements that videos remain sharp and reliable for up to 30 chunks are supported only by qualitative stills. The delta-action ablation in Figure 9 is also qualitative. Because the paper positions EVAC as a world simulator, quantitative fidelity metrics and a direct comparison against the EnerVerse baseline are needed to establish that the proposed architecture, rather than the curated training data, improves over the prior system.","section":"§4.2, Appendix A.2"}],"minor_comments":[{"comment":"The heading 'Mutli-Level Action Condition Injection' contains a typo and should read 'Multi-Level Action Condition Injection'.","section":"§3.1"},{"comment":"The notation 'O∈R V×(H+K)×3×h×w' appears to be missing a multiplication sign between V and (H+K); it should be 'R^{V × (H+K) × 3 × h × w}'.","section":"§3, first equation"},{"comment":"The left panel reports success rates as 28%, 100%, 85%, 55%, 25%, 90%, 88%, and 50% without an explicit legend mapping each bar to the four tasks and two evaluation modes; adding direct labels above each bar would greatly improve readability.","section":"§4.3, Figure 7"},{"comment":"Several references are incomplete or informal, including [20] (missing title and venue details) and [21] (a blog post with no author or title); these should be completed or reformatted.","section":"References"},{"comment":"The limitations section does not mention the distribution-shift limitation of the evaluator application, which is the most serious limitation from the paper's own Section 4.5; a sentence acknowledging this would make the stated limitations more consistent with the experimental evidence.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an incremental extension of EnerVerse by a heavily overlapping author group, yet it contains no direct quantitative comparison with EnerVerse or with other action-conditioned world models. The absence of such a baseline is a novelty and positioning concern that should be raised with the authors during revision, in addition to the statistical gaps described in the main report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick read: EVAC is EnerVerse plus explicit action conditioning, ray maps for moving wrist cameras, and failure-trajectory training. That combination is genuinely new relative to the cited EnerVerse and UniPi work, and the qualitative figures show controllable generation, including distingishing toss from shake via the delta-action attention. The data-engine experiment, though small, is a real result: 20 demos plus 30% synthetic data lifts success from 0.28 to 0.36 on a bottle-from-box task. Credit where due: the authors built and trained the thing, and the failure-data ablation in §4.5 is an honest admission that without failure coverage the model hallucinates.\n\nWhere it gets soft is the evaluator claim. The paper says 'high correlation' between EVAC and real-robot success, but Figure 7 is four task-level points with no error bars, no correlation coefficient, and no inter-rater agreement. The stress-test note lands: §4.5 demonstrates exactly the failure mode—a missed grasp is rendered as success when failure data is absent—and the paper never measures how far the tested policy's actions are from EVAC's training distribution. So evaluator validity is only shown for four in-distribution tasks. That is the load-bearing claim, and it needs more than a qualitative trend.\n\nAnother gap: no comparison with EnerVerse or other action-conditioned world models. Since EVAC is explicitly built on EnerVerse, the reader cannot tell what the new modules actually buy beyond the two qualitative ablations. Also, no error bars on the data-engine gain and no artifacts in the preprint, despite the project page promise.\n\nThe citation pattern is fine; building on EnerVerse and AgiBot-World is legitimate, and the failure-trajectory sourcing is described. No circularity: the evaluator is checked against real-robot outcomes.\n\nBottom line: credible systems paper with a real but narrow experimental base. The right move is peer review with heavy revision expected: add statistics, an OOD analysis of policy actions, a comparison against EnerVerse, and release at least the model or key scripts. A researcher working on video-based world models for manipulation will get value; a researcher looking for a validated evaluation alternative should wait.\n\nRecommendation: send to serious referees, but flag the evaluator evidence as the central issue.","headline":"EVAC is a plausible engineering extension of EnerVerse with an under-supported evaluator claim; it deserves serious refereeing but not acceptance at current evidence.","tokens_in":12043,"tokens_out":1856,"would_cite":false,"duration_ms":19226,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An action-conditional video world model can stand in for real-robot testing, ranking policies by generated rollouts.","keywords":["action-conditioned world model","video generation","robotic manipulation","policy evaluation","data augmentation","multi-view video generation","failure trajectories","latent diffusion model"],"falsifier":"Run a policy on a task that shifts its action distribution, for example by adding increasing position or velocity offsets to its action outputs, and compare EVAC's success-rate rankings against real-robot rollouts from the same initial frames. If the rankings diverge as the offset grows, or if EVAC shows successful grasps in videos where the generated gripper never closes, the evaluator claim is falsified for out-of-distribution actions.","tokens_in":11094,"feed_emoji":"🎥","tokens_out":8424,"duration_ms":78821,"temperature":0.7,"pith_summary":"EVAC is an action-conditional world model: given an initial observation and a sequence of end-effector actions, it generates the multi-view video frames that would follow, including frames from moving wrist cameras. The paper's central claim is empirical: when human raters score EVAC-generated rollouts, the resulting success rates track real-robot evaluations across four manipulation tasks and across training checkpoints of the same policy. A second claim is that the model can serve as a data engine, synthesizing new trajectories from a few human demonstrations; adding 30 percent synthetic data raised a policy's success rate from 0.28 to 0.36. The authors also argue that training on failure trajectories is necessary to avoid hallucinating successful grasps that never physically occurred. If the correlation with real-world outcomes holds beyond the tested tasks, robot developers could iterate on policies using generated video instead of physical hardware.","feed_headline":"Generated video can rank robot policies without hardware","feed_subtitle":"A world model that turns predicted actions into future views tracks real-robot success rates across tasks and training steps.","key_machinery":"The load-bearing mechanism is the multi-level action-conditioning stream. At each timestep, the end-effector pose and gripper openness are rendered into an action map, with pixel coordinates from calibrated cameras, orientation unit vectors, and a shaded circle encoding grip state; a frozen vision encoder turns this map into features concatenated with the image latents. A separate delta action attention module computes differences between consecutive actions and injects those motion cues through cross-attention, giving the model access to speed and acceleration. For moving wrist cameras, ray maps defined by camera origins and directions are concatenated into the input features so the model knows where each view is looking at each instant. The generation itself is chunk-wise autoregressive diffusion with a sparse memory of four history frames, which keeps the video visually consistent over multiple 16-frame chunks. The training mix also includes failure trajectories, and the paper shows this is what prevents the model from fabricating a grasp when the gripper never contacts the object.","core_discovery":"On its own terms, the paper establishes that an embodied world model can be made action-controllable and can double as a low-cost evaluator. EVAC builds on a diffusion-based video generator and injects the robot's action at two levels: spatial-aware pose maps that draw the end-effector's position, orientation, and gripper state into image space, and a delta action attention module that encodes velocities and accelerations between consecutive actions. Multi-view consistency, including dynamic wrist cameras, is handled by ray map encoding of camera origins and directions. Trained on a large-scale manipulation dataset supplemented with deliberately collected failure trajectories, EVAC generates video that human raters score consistently with real rollouts: reported success rates for four tasks and three training checkpoints follow the same trends (for example, bottle retrieval 28% real versus 25% EVAC; training checkpoints 40/40, 61/63, 79/76 real versus EVAC). The same generator, run in reverse, synthesizes augmented training trajectories that improved policy success from 0.28 to 0.36 on a bottle-extraction task.","pith_inferences":["If the correlation generalizes, the bottleneck in robot manipulation shifts from physical testing to world-model fidelity and video review, making large-scale policy search and training-time evaluation far cheaper.","The same action-conditioned mechanism could be used for test-time planning, not just evaluation: select action sequences by generating and scoring their outcomes before execution.","The delta action attention finding suggests that explicit conditioning on velocity and acceleration is a broadly useful recipe for physically plausible video generation, with applications outside robotics.","A prudent extension would measure how close a queried policy's actions are to EVAC's training distribution and use that distance as a confidence bound on evaluator results."],"forward_implications":["Policies can be ranked and compared during development by generating rollouts rather than deploying on robots, cutting hardware and setup costs.","Human evaluators or video-language models can score generated rollouts; EVAC's results show it picks out the same performance trends as real tests, including progress across training steps.","Small expert datasets can be expanded: augmenting 20 demonstrations with 30 percent EVAC-generated trajectories raised a bottle-extraction policy's success rate from 0.28 to 0.36.","Including failure data in world-model training suppresses hallucinations, so generated video can be trusted to show failed grasps instead of inventing successes.","Multi-view generation with ray maps supports dynamic wrist-camera observations, which are important for dexterous manipulation policies."],"supporting_citations":[{"why":"Supplies the chunk-wise autoregressive diffusion architecture and sparse memory mechanism that EVAC extends with action conditioning.","marker":"[1]"},{"why":"Provides the large-scale manipulation trajectory dataset, including mined failure cases, and the policy model used in the evaluator and data-engine experiments.","marker":"[2]"},{"why":"Motivates the paper's premise that video generation models can act as world simulators for embodied agents.","marker":"[3]"},{"why":"Defines the latent diffusion formulation used to generate frames in the video generator.","marker":"[7]"},{"why":"Supplies the UNet-based video diffusion model adopted as the generation baseline.","marker":"[23]"},{"why":"Provides the frozen vision encoder that converts rendered action maps into features for the diffusion model.","marker":"[29]"},{"why":"Used for qualitative cross-checks in a standard simulator benchmark after fine-tuning EVAC on simulated trajectories.","marker":"[33]"}],"fun_headline_variants":["Action-fed video world model scores robot policies as well as real runs","Generated future frames gauge robot success without a physical robot","World model turns predicted actions into realistic robot tests","Action-conditional video forecaster grades robot policies cost-effectively","Synthetic video from predicted actions estimates robot skill accurately"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluator's usefulness rests on the assumption that the policy being tested produces actions similar to those EVAC saw during training; the paper does not measure how far a policy can deviate before the generated video becomes plausible-looking but physically wrong.","fun_headline_variants_meta":{"raw":{"variants":["Action-fed video world model scores robot policies as well as real runs","Generated future frames gauge robot success without a physical robot","World model turns predicted actions into realistic robot tests","Action-conditional video forecaster grades robot policies cost-effectively","Synthetic video from predicted actions estimates robot skill accurately"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000914,"raw_usage":{"total_tokens":3924,"prompt_tokens":941,"completion_tokens":2983,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":2903}},"tokens_in":557,"tokens_out":2983,"duration_ms":20519,"temperature":1.0,"reasoning_tokens":2903,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:25:46.550605+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a policy on a task that shifts its action distribution, for example by adding increasing position or velocity offsets to its action outputs, and compare EVAC's success-rate rankings against real-robot rollouts from the same initial frames. If the rankings diverge as the offset grows, or if EVAC shows successful grasps in videos where the generated gripper never closes, the evaluator claim is falsified for out-of-distribution actions.","supporting_citations":[],"review_version":1}