{"id":"f105c246-d472-47df-99b4-0f63c90e4963","arxiv_id":"2607.02642","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Long-horizon action-faithful consistency, not short-term visual realism, dominates world-model reliability for robot policy evaluation; GigaWorld-1 implements that roadmap and gains 14.9% on evaluator-alignment metrics.","lead":"This paper shows that robot world models work as policy evaluators mainly when they stay action-faithful over long rollouts, not when they just look photorealistic. It ships a paired real/sim benchmark (WMBench), large ablations, and GigaWorld-1, a model built from that design map.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Headline 14.9% rests on a paper-chosen six-metric average, not on closed-loop ranking correlation ρ with real success rates.","rationale":"The paper is a solid methods/benchmark contribution: large controlled ablations, open release, and clear evidence that long-horizon action fidelity and memory matter more than short-horizon photorealism. The reader’s CONDITIONAL verdict and domain-coverage caveat are right. The sharper internal soft spot is metric substitution: Eq. 4 defines evaluator quality as policy-outcome ranking agreement, yet the 14.9% and Table 9 AVG are a curated diagnostic mean whose link to closed-loop π-ranking is only correlational (via WMES on challenge videos) and not demonstrated at the same scale as the ablations. That does not overturn the design roadmap or the qualitative closed-loop plots, but it means the strongest numerical claim overstates what was measured. A single ρ comparison on the existing closed-loop set would settle whether the gap is cosmetic or material; until then CONDITIONAL remains appropriate, with slightly tighter language on what “evaluator-alignment” numbers actually are. Agreement with the reader is partial: same overall caution, different primary weak link (metric identity vs OOD coverage).","tokens_in":40465,"tokens_out":757,"duration_ms":7168,"concrete_test":"On the same held-out WMBench closed-loop episodes used for Figs. 16–17, compute Spearman/Pearson ρ between real-robot success rates and world-model success rates (VLM or human WMES) across all available policy checkpoints for GigaWorld-1-Plus vs Wan 2.2 5B (and Cosmos-Predict2.5) under identical post-training and rollout protocol. If Δρ is < ~0.1 or non-significant while the six-metric average still shows +14.9%, the headline claim should be restated as diagnostic-average improvement, not evaluator-alignment improvement.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim packages three insights plus a 14.9% gain of GigaWorld-1-Plus over Wan 2.2 5B (Table 9: 0.6834 vs 0.5948). That number is the mean of six retained diagnostics (Aesthetic, Image, JEPA, Semantic, Subject, Trajectory) after shared post-training, selected because they correlate with human WMES on challenge rollouts (Findings 1–2, Figs. 4–5). The paper’s own primary evaluator target, however, is ranking/success agreement ρ = Corr(S_real(π), S_wm(π)) across policies (Eq. 4, Sec. 3). Closed-loop evidence is only qualitative/bias plots on four tasks and a few checkpoints (Figs. 16–17, Sec. 6.5.5), with residual optimistic bias on contact-sensitive failures that the authors report. Thus the headline percentage can improve while the quantity that would actually justify “reliable policy evaluator” (policy ranking fidelity under closed-loop interaction) remains unquantified at the same scale. The reader correctly flags domain-bound generalization; the more immediate load-bearing gap is that the reported 14.9% is not the same object as the claimed evaluator-alignment quantity.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper studies world models as surrogate evaluators of robot policies, arguing that real-robot evaluation is the bottleneck for embodied foundation models. It introduces WMBench (paired teleoperation and policy-rollout data over eight manipulation task families), analyzes seven video world models, four action encodings, and 324k+ annotated rollouts (including CVPR 2026 challenge submissions), and reports three design insights: long-horizon action-faithful consistency dominates short-term visual realism; pretraining gains require balancing general physical priors with robot controllability; and action interface, memory, and evaluator-oriented post-training strongly affect real-world alignment. These are instantiated in GigaWorld-1 (Wan-based, ~13k hours multi-source data, pixel-aligned EE/ray control, hierarchical memory, progressive training), which improves a six-metric average by 14.9% over Wan 2.2 5B under matched post-training (Table 9). Code, models, and data are released.","tokens_in":40923,"tokens_out":1660,"duration_ms":20631,"significance":"If the design claims hold, the work is a substantial contribution to scalable robot policy evaluation: it reframes world models as external policy evaluators rather than only data engines or planners, provides a large paired real/sim benchmark and metric analysis (Figs. 4–5, WMES), and ships a concrete open roadmap (data mixture, spatially aligned control, memory, distillation) with reproducible artifacts. The controlled ablations (Tables 2–4, 9) and community-scale annotation are genuine strengths relative to prior proof-of-concept evaluator papers. The practical value depends on whether closed-loop ranking fidelity—not only diagnostic video metrics—is demonstrated at the same scale as the headline gains.","major_comments":[{"comment":"Sec. 3, Eq. (4) defines the primary evaluator target as ranking/success agreement ρ = Corr(S_real(π), S_wm(π)) across policies. The abstract, intro, and Sec. 6.5.1 instead headline a 14.9% gain of GigaWorld-1-Plus over Wan 2.2 5B on a paper-chosen six-metric average (Aesthetic, Image, JEPA, Semantic, Subject, Trajectory; Table 9: 0.6834 vs 0.5948). Those diagnostics are justified by correlation with human WMES on challenge rollouts (Findings 1–2, Figs. 4–5), but they are not ρ. Closed-loop evidence in Sec. 6.5.5 is limited to task-level success-rate scatter and Gen−Real bias bars on four tasks/subtasks (Figs. 16–17), without a quantified ρ (or Kendall/Spearman ranking) over a multi-checkpoint policy set for GigaWorld-1 vs baselines. Please report ρ (with CIs) under the same closed-loop protocol for the main models, or reframe the 14.9% claim so it is not presented as the primary evaluato","section":"Sec. 3 Eq. (4); Sec. 6.5.1 Table 9; Sec. 6.5.5 Figs. 16–17"},{"comment":"Sec. 6.5.5 and Fig. 17 acknowledge residual optimistic bias on contact-sensitive failures (e.g., pour/press subtasks). Because the paper’s own Finding 3 and Table 4 argue that long-horizon action-faithful consistency—not short-horizon visual quality—dominates evaluator reliability, the closed-loop calibration gap is load-bearing. Either quantify how often GigaWorld-1 flips real success/failure relative to baselines (confusion matrices or per-subtask agreement rates on the full WMBench closed-loop set), or temper claims that the model is “specially optimized for policy evaluation” until failure-mode calibration is measured at the same scale as the diagnostic average.","section":"Sec. 6.5.5; Fig. 17; Finding 3 / Table 4"},{"comment":"WMES and several diagnostic metrics (Perspectivity, Instruction Following, Interaction Quality, Semantic Alignment) rely on VLM judges (Sec. 4.3; Finding 4–5; Table 1). The LoRA VLM is trained to predict human WMES with score-token weight 8.0, then used for scalable outcome assessment. Metric selection for Table 9 is guided by correlation with that same WMES family. This is not fatal circularity—human WMES and real success labels remain external—but it creates mild self-reinforcement risk if models are tuned toward VLM-preferred appearance/geometry. Please report (i) human-only vs VLM-only ranking of the main models on a held-out subset, and (ii) sensitivity of the six-metric average and any ρ to excluding VLM-judged diagnostics.","section":"Sec. 4.3; Findings 4–5; Table 1; Table 9"}],"minor_comments":[{"comment":"Fig. 1 and abstract claim “324,000+ analyzed rollouts” and “12K+ hours training data”; Table 6 totals ~12,980 hours. Align the rounded figures and clarify whether 324k counts segments or full closed-loop episodes (Sec. 4.1 says segments chained into episodes of 20–30 segments).","section":"Fig. 1; Abstract; Sec. 4.1; Table 6"},{"comment":"Table 3 ranks control interfaces on Trajectory Accuracy and motion metrics for Wan 2.1 1.3B only. A short note on whether channel-concat remains best for the 5B backbone (GigaWorld-1-Plus) would strengthen Finding 9.","section":"Table 3; Finding 9"},{"comment":"Sec. 4.1 train/test split is episode-disjoint within the same eight task families and platforms. Sec. 7 correctly flags limited coverage of mobile/dexterous/safety-critical settings; a brief explicit statement in Sec. 6.5 that OOD claims (Fig. 15) are appearance/content shifts, not embodiment or policy-family shifts, would prevent over-reading.","section":"Sec. 4.1; Sec. 6.5.4; Sec. 7"},{"comment":"Typo in Fig. 3 caption/step labels: “Train Wodel Model” should be “Train World Model.” Several figure panels (e.g., Fig. 2 Chinese annotation fragment in the source) should be cleaned for the camera-ready version.","section":"Fig. 2; Fig. 3"},{"comment":"Cosmos-3 is mentioned as planned but unavailable (footnote in Finding 6). Either update the comparison if multiview access is obtained, or move the note to a single limitations paragraph to avoid dangling promises.","section":"Finding 6 footnote; Sec. 7"}],"recommendation":"major_revision","confidential_remarks":"The empirical scale and open release make this a strong candidate if the authors close the ρ vs six-metric gap; without that, the headline 14.9% risks being cited as “evaluator reliability” when it is primarily diagnostic video quality under matched post-training. Scope fits cs.RO / embodied evaluation venues well. No integrity concerns beyond the usual need to keep challenge-submission and internal pipeline claims clearly separated from the public benchmark protocol."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a real systems paper on a real bottleneck. The new pieces are WMBench (paired teleop + policy rollouts, episode-disjoint, eight manipulation families), the 324k annotated challenge rollouts with human WMES, the metric–WMES correlation study, and a concrete design map (balanced PhysData/robot data, channel-concat action maps, hierarchical memory + first-frame anchor, progressive AR + distillation) realized in open GigaWorld-1 Nano/Plus.\n\nWhat they do well is the controlled comparison. Seven backbones, four action interfaces, data-mixture ablations, long-horizon PSNR/FID/FVD tables, and closed-loop success-rate plots against real robots. Findings that long-horizon action fidelity beats short-term photorealism, that static-consistency metrics can rank degenerate models highly, and that spatially aligned control plus memory matter are backed by the tables and figures, not slogans. Full release of code, models, data, and toolkits is genuine infrastructure value. Citations cover WorldEval, WorldGym, Ctrl-World, Cosmos, Wan, etc.; the framing is not oversold as inventing the evaluator idea.\n\nSoft spots, in proportion. The stress-test is right on the packaging: the 14.9% (0.6834 vs 0.5948) is the mean of six retained diagnostics after shared post-training, selected because they correlate with WMES. The paper’s own primary target is ρ = Corr(S_real, S_wm) across policies (Eq. 4). Closed-loop evidence is bias plots and task-level alignment on four tasks/subtasks (Figs. 16–17), not a large-scale ranking correlation at the same scale as the 324k study. Residual optimistic bias on contact failures is reported honestly. Domain is still eight families, same platforms/cameras; OOD is qualitative. VLM judges and the LoRA WMES predictor introduce mild circularity risk if people optimize to them. None of that sinks the work; it means the strongest claim should be read as “better on the evaluator-relevant diagnostics and closer closed-loop calibration,” not “proven general ranking fidelity.”\n\nMath is standard (flow matching, AR diffusion, SLERP, NDTW). Data scale and open release are the main assets. For anyone building or using video world models for policy iteration, this is worth reading and citing. I would send it to referees; it is methods/benchmark work that advances the conversation even if the headline number needs tighter language.","headline":"Solid large-scale empirical roadmap for world models as robot policy evaluators; the 14.9% headline is on a chosen diagnostic average, not quantified closed-loop ranking ρ.","tokens_in":41670,"tokens_out":611,"would_cite":true,"duration_ms":7904,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Robot policy evaluators succeed by staying action-faithful over long horizons, not by looking more photorealistic.","keywords":["world models","robot policy evaluation","WMBench","action-conditioned video generation","long-horizon rollout","embodied AI","GigaWorld-1","closed-loop evaluation"],"falsifier":"Take a set of policies whose real-robot success ranking is known on tasks outside WMBench’s eight families or on a different embodiment; if the world model’s closed-loop ranking correlation collapses while short-horizon visual metrics remain high, the central claim that long-horizon action faithfulness on this benchmark is sufficient fails.","tokens_in":41311,"feed_emoji":"🤖","tokens_out":1020,"duration_ms":9561,"temperature":0.7,"pith_summary":"Robot foundation models cannot be scored the way language models are scored: every checkpoint still needs slow, supervised physical rollouts. This paper argues that learned video world models can stand in for those rollouts only when they preserve the success and failure ranking of real policies, not when they merely produce pretty frames. The authors build WMBench from paired teleoperation and policy trajectories, then run a controlled study of seven world models, four action encodings, and more than 324,000 simulated rollouts against real executions. Three results follow. First, evaluator quality tracks long-horizon action fidelity far more than short-term visual realism; metrics that reward static or action-ignorant videos actively mislead ranking. Second, pretraining helps only when broad physical knowledge is kept in balance with robot-specific controllability. Third, design choices—pixel-aligned action maps, hierarchical memory with a first-frame anchor, and evaluator-focused post-training—decide whether simulated outcomes match real ones. The authors turn those rules into GigaWorld-1, trained on roughly 13,000 hours of mixed data, which raises the core evaluator-alignment average by 14.9 percent over the strongest matched general-purpose baseline.","feed_headline":"Action fidelity beats pretty frames for robot evaluators","feed_subtitle":"Long-horizon consistency, not photorealism, predicts which world models rank policies like real robots do","key_machinery":"WMBench: a paired real/world-model rollout benchmark whose primary target is ranking correlation between real-world and world-model success rates (and the related ordinal WMES score), used to isolate which metrics, data mixes, and architectures actually predict real policy outcomes.","core_discovery":"A world model is a reliable robot-policy evaluator only when its closed-loop rollouts stay action-faithful over long horizons and therefore reproduce the same success/failure ranking that real robots produce; short-term photorealism is secondary, and the decisive levers are balanced physical-plus-robot data, spatially aligned action control, persistent multi-scale memory, and post-training aimed at evaluator agreement rather than generic video quality.","pith_inferences":["The same long-horizon action-faithfulness test could become a filter for world models used as data engines or planners, not only as offline evaluators.","If VLM outcome labeling stays within a few percent of human WMES at method level, most future evaluator leaderboards can run without exhaustive human annotation.","Contact-rich failure modes still show optimistic bias; hybrid world models that add explicit contact or force state may be the next necessary step beyond pure video.","Once ranking correlation is the accepted target, sim-to-real gaps in classical simulators can be measured against the same yardstick rather than against visual fidelity alone."],"forward_implications":["Policy developers can replace a large fraction of hardware rollouts with closed-loop world-model evaluation once the model is scored by ranking agreement rather than frame beauty.","Metric suites that reward static or action-ignorant videos will systematically promote weak evaluators and should be dropped from evaluator leaderboards.","Training recipes must mix broad physical video with robot data; robot-only fine-tuning improves embodiment look but can erase the priors needed for reliable evaluation.","Spatially aligned control maps plus hierarchical memory become standard requirements for any video world model intended as a policy surrogate.","Open release of the benchmark, models, and annotation toolkit lets the community iterate on evaluator design the way language-model groups iterate on digital suites."],"fun_headline_variants":["Long-horizon action fidelity defines reliable robot world models","Action-faithful rollouts beat visual realism for policy evaluators","World models match real robot rankings via long action consistency","Action fidelity over horizons outranks photorealism in eval","Robot policy eval hinges on long-horizon action-faithful rollouts"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"Agreement on WMBench’s eight held-out manipulation families, still drawn from the same platforms, cameras, and task distribution, is enough to claim that a world model will rank novel policies reliably under broader initial states, embodiments, and contact-rich failures.","fun_headline_variants_meta":{"raw":{"variants":["Long-horizon action fidelity defines reliable robot world models","Action-faithful rollouts beat visual realism for policy evaluators","World models match real robot rankings via long action consistency","Action fidelity over horizons outranks photorealism in eval","Robot policy eval hinges on long-horizon action-faithful rollouts"]},"model":"grok-4.5","effort":"low","cost_usd":0.003826,"raw_usage":{"total_tokens":1281,"prompt_tokens":869,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":38260000,"prompt_tokens_details":{"text_tokens":869,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":345,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":869,"tokens_out":67,"duration_ms":14484,"temperature":1.0,"reasoning_tokens":345,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T08:03:22.970962+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Take a set of policies whose real-robot success ranking is known on tasks outside WMBench’s eight families or on a different embodiment; if the world model’s closed-loop ranking correlation collapses while short-horizon visual metrics remain high, the central claim that long-horizon action faithfulness on this benchmark is sufficient fails.","supporting_citations":[],"review_version":1}