Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

No interactive world model yet holds action, vision, physics, and memory together over long open-world roam sessions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 10:01 UTC pith:KWMJVSH3

load-bearing objection A real long-horizon IWM diagnostic: four axes that fix known blind spots, broad multi-model evidence that none dominate, with physics/memory ranks partly judge-mediated. the 3 major comments →

arxiv 2606.31672 v3 pith:KWMJVSH3 submitted 2026-06-30 cs.CV cs.AI

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

classification cs.CV cs.AI
keywords interactive world modelslong-horizon stabilityaction followingvisual driftinteraction physicsscene memorysubject memoryopen-world benchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Interactive world models let a user steer a generated scene with keyboard actions, but short trajectory-only tests hide whether that control stays faithful for tens of seconds. WorldRoamBench is an open-world suite of more than 600 continuous WASD/IJKL rollouts lasting 10–60 seconds across Nature, Urban, and Indoor scenes in first- and third-person views. It scores four long-horizon axes with purpose-built metrics: keystroke-level action accuracy that is independent of each model’s motion scale, segment-based visual drift that catches mid-rollout collapse, physics that is scored only when the model actually follows the command, and memory that localizes the true observation–revisit transition before comparing 3D scene geometry and third-person subject identity. Ten-plus open and closed models are ranked; none dominates every axis, and even the leaders land only in the moderate range. The paper’s claim is that long-horizon multi-axis stability, not short-clip aesthetics, is the right bar for deployable interactive worlds.

Core claim

Across 600+ long open-world interaction cases, no evaluated interactive world model reliably satisfies all four stability dimensions—per-frame action following, visual drift, interaction physics, and action-decoupled scene/subject memory—and the strongest models still reach only moderate overall scores.

What carries the argument

WorldRoamBench’s four-axis long-horizon protocol: latent-stride per-frame action accuracy plus adaptive TrajScore; segment best-vs-worst visual drift; controllability-gated mechanics/optics/3D physics; and transition-localized 3D point-cloud scene memory plus tracking-plus-VLM subject memory.

Load-bearing premise

The automated pose estimators, depth reconstructions, and vision-language judges used for action, physics, and memory are treated as faithful enough to rank models even when human agreement is only partial.

What would settle it

A model that posts high scores on all four WorldRoamBench dimensions under the same 10–60 s WASD protocol, while independent human raters confirm the physics and subject-memory labels, would overturn the claim that current systems remain only moderately stable.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. WorldRoamBench proposes an open-world, long-horizon (10–60 s WASD/IJKL) benchmark for interactive world models spanning four dimensions—per-frame and trajectory action following, segment-based visual drift, controllability-gated interaction physics (mechanics/optics/3D), and action-decoupled scene/subject memory—over 600+ Nature/Urban/Indoor FPV/TPV cases. The paper defines tailored metrics (latent-stride action discretization, adaptive TrajScore, best-vs-worst segment drift, transition-localized point-cloud F1, tracking+VLM subject memory), provides adapters for heterogeneous model interfaces, and evaluates 10+ open- and closed-source IWMs. The central empirical claim is that no model reliably satisfies all four dimensions and even the best overall scores remain only moderate (e.g., Genie 3 FPV overall 73.81/100), with diagnostic findings such as trajectory score ≠ per-frame correctness and memory confounded by action imprecision.

Significance. If the evaluation protocol is accepted as sufficiently faithful, this is a timely and useful contribution: existing IWM benchmarks are largely short-horizon and trajectory-centric and under-treat physics and memory under imperfect control. The four-axis design, open-domain long rollouts, explicit cross-interface adapters (Tables 2/5), and public-facing leaderboard/dataset framing address a real measurement gap. Strengths include carefully motivated metric innovations (adaptive GT trajectories that isolate direction/shape from semantic scale; segment-based drift for non-monotonic collapse; action-aware memory localization), extensive appendices with algorithms and prompts, and an initial human check of VLM judges. The multi-model results and qualitative failure cases make a credible case that current IWMs are uneven and not yet deployably stable.

major comments (3)
  1. [Appendix J.2; Table 3; §3.4–3.5] Appendix J.2 reports only 73.2% agreement between calibrated Qwen3-VL-PLUS and human majority on 100 mechanics videos, and 88.3% ±1 concordance on 60 TPV subject-memory rollouts. Physics and memory are load-bearing for absolute S_physics/S_memory and for several pairwise orderings in Table 3 (e.g., Happy Oyster vs Genie 3 FPV physics 72.33 vs 68.95; memory 55.42 vs 73.24). The manuscript does not report inter-rater reliability among humans, bootstrap/sensitivity of ranks under plausible label flips, or human re-scoring of the full physics/memory suites. Please add rank-stability analysis (or expanded human evaluation) and state clearly which leaderboard claims are robust to residual judge error versus which are provisional.
  2. [§3.2.1; Appendix C; Algorithms 13–14] §3.2.1 and Appendix C select the latent-stride translation threshold per video by maximizing exact-match accuracy against the GT action sequence (Θ={0.002,0.005,0.01}, scaled by K=4). This improves robustness to scale variation but makes Acc_strict partly an optimized match rather than a fixed decoder, which can inflate action scores and complicate fair cross-model comparison if different models systematically prefer different operating points. Please report scores under a single fixed threshold (or a held-out threshold schedule), quantify sensitivity of Table 3 action rankings to the sweep, and clarify whether the reported Acc_strict is best-of-sweep or post-hoc selected.
  3. [§3.4.1; §3.6 Eq. (22); Table 3(b); Table 4] The validity gate (§3.4.1; Appendix E.1) excludes uncontrolled rollouts from physics aggregates (score −1), and TPV overall is further multiplied by control rate (§3.6, Eq. 22). This is methodologically motivated but means S_physics is conditioned on successful action response, so models that fail control are under-represented rather than penalized inside the physics domain. Given large TPV control-rate gaps (Table 4: Happy Oyster 84.43 vs HY-World 1.5 21.23) and only four TPV models, please report physics both gated and ungated (or with failures scored as 0), and discuss how conditioning affects the claim that closed-source models lead physics.
minor comments (5)
  1. [§3.6 Eq. (21); Table 3(c)] Equal default weights w1=…=w4=0.25 for S_ROAM (§3.6, Eq. 21) are reasonable for a first leaderboard but should be accompanied by a short sensitivity table (e.g., leave-one-dimension-out or reweighting) so readers can see whether overall ranking is weight-stable.
  2. [§3.2; §3.5.1; Appendix B.2.2; Appendix F.1] ViPE pose estimation and monocular depth for TrajScore and scene memory are central; a brief failure-mode discussion (textureless regions, pure rotation, closed-source RAFT start-frame alignment in Appendix B.2.2) would help readers interpret residual error floors.
  3. [Figure 9; Figures 7–8] Figure 9 caption states averages over 120 videos per model for the first 300 frames; reconcile this with the per-dimension case counts in Figures 7–8 so the sampling frame for drift curves is unambiguous.
  4. [Table 1; Table 3 caption; §3.4.2] Table 1 footnote on TPV coverage is helpful; consider also stating explicitly in the main text how many physics 3D-consistency cases exist only in FPV, since TPV Physics Score omits 3D (Table 3 caption).
  5. [Figure 3; References] Minor polish: ensure consistent naming (WorldRoamBench vs WorldROAM in Fig. 3; TrajScore vs Trajectory Acc) and that all arXiv-style references for concurrent work are complete where DOIs/venues are available.

Circularity Check

0 steps flagged

No significant circularity: WorldRoamBench is a measurement protocol and empirical leaderboard, not a fitted or self-justifying derivation of a physical/theoretical claim.

full rationale

The paper’s load-bearing claim is empirical—that 10+ IWMs evaluated under a fixed four-axis protocol over 600+ long-horizon WASD cases none dominate all dimensions and even the best remain only moderate (e.g., Genie 3 FPV overall 73.81). Action, visual, physics, and memory scores are defined as evaluation metrics (latent-stride discretization vs GT keys; segment best–worst drift; controllability-gated VLM/mechanics protocols; transition-localized point-cloud F1 and subject VLM scores) and then applied uniformly to model rollouts. They are not derived from, or fitted to, the ranking they produce. Adaptive GT trajectory construction deliberately reuses each model’s own displacement magnitude so TrajScore measures shape/direction rather than scale; that is an explicit fairness design, not a hidden reduction of a prediction to its inputs. Per-video threshold sweep over three move thresholds chooses the operating point that maximizes exact-match accuracy on that video, which can mildly optimistically bias Acc_strict, but still reports a calibrated match rate against independent GT action programs rather than predicting a new quantity from a fitted parameter. Physics and memory rely on external tools (ViPE, monocular depth, SAM2, Qwen3-VL-PLUS) whose reliability is a correctness/judge-error concern, not circular self-justification. Citations (CompassReward/WorldCompass, Helios, WBench, etc.) supply methodological building blocks; none import a uniqueness theorem or ansatz that forces the central ‘none dominates’ result. No self-definitional loop, no fitted-input-as-prediction of a scientific law, and no load-bearing self-citation chain. Score 0 is appropriate for a self-contained benchmark paper.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 5 invented entities

As a benchmark paper, load-bearing content is mostly evaluation design choices and external tooling assumptions rather than physical postulates. Free parameters are thresholds and aggregation weights that can move rankings. Axioms are domain assumptions that pose estimators, monocular reconstruction, and VLM judges are good enough proxies for human notions of action, physics, and memory. Invented entities are the named protocols/metrics themselves; they have independent operational definitions but no existence claim beyond the evaluation pipeline.

free parameters (5)
  • Latent-stride action thresholds (τ_move base set, τ_rot, τ_axis; K=4)
    Discretization thresholds are swept per video over {0.002,0.005,0.01} (scaled by K) and the best exact-match threshold is kept; operating point is not fixed a priori.
  • Visual drift segment count N=10
    Best-vs-worst segment formulation depends on the chosen number of equal segments.
  • TrajScore resampling and floors (N_pts=60, ℓ_min=0.5, ϕ_min=10°)
    Normalization denominators and fixed resample count are hand-set design constants affecting trajectory scores.
  • Physics/memory gates (τ_iou=0.4, τ_ctrl=0.9; point-cloud τ_d, registration fitness)
    Controllability and geometric match thresholds decide which cases enter aggregate scores.
  • Equal dimension weights w1..w4=0.25 and TPV control-rate multiplier
    Overall ROAM ranking depends on equal weighting and TPV control-rate scaling, not an empirically derived utility.
axioms (5)
  • domain assumption ViPE-estimated camera trajectories plus latent-stride directional thresholding recover discrete WASD/IJKL actions well enough for cross-model ranking.
    Action metrics rest on pose estimation and CompassReward-style discretization (Sec. 3.2, App. C).
  • ad hoc to paper Adaptive GT trajectories that reuse each model’s own displacement magnitude fairly isolate direction/shape from semantic scale.
    Core design choice of TrajScore (Sec. 3.2.2, App. D); fairness claim is definitional rather than externally validated.
  • domain assumption Qwen3-VL-PLUS binary/ordinal judgments are adequate proxies for mechanics, optics, and subject-memory human labels after prompt calibration.
    Physics and subject-memory pipelines are VLM-centered (Sec. 3.4–3.5; Apps. E–G; human study App. J.2).
  • domain assumption Monocular depth reconstruction, semantic filtering, and quality-aware registration yield scene geometry comparable enough for revisit F1 memory.
    Scene memory score is defined on aligned observation/revisit point clouds (Sec. 3.5.1, App. F.1).
  • domain assumption Obstacle-free action/memory paths and validity gates successfully isolate target dimensions from confounding physics or static generations.
    Test-suite design principle in Sec. 4 and physics validity gate in Sec. 3.4.1.
invented entities (5)
  • WorldRoamBench four-axis long-horizon stability score (S_ROAM) independent evidence
    purpose: Aggregate ranking of interactive world models under continuous keyboard control.
    Composite score and leaderboard are paper-defined evaluation constructs.
  • Latent-stride per-frame action accuracy + adaptive TrajScore independent evidence
    purpose: Measure keystroke fidelity without unfair cross-model motion-scale comparison and without trajectory-only masking.
    New action-evaluation protocol built on ViPE/CompassReward ideas.
  • Segment-based aesthetic/imaging drift metric independent evidence
    purpose: Capture non-monotonic mid-rollout visual collapse missed by averages or start-vs-end comparisons.
    Defined in Sec. 3.3 as a Helios-inspired but best-vs-worst segment score.
  • Controllability-gated interaction-physics suite (mechanics/optics/3D) independent evidence
    purpose: Score physical plausibility only when the model actually executes the commanded interaction.
    Paper-specific protocol stack with VLM judges and validity gates.
  • Action-decoupled scene/subject memory protocol independent evidence
    purpose: Separate memory degradation from imperfect return trajectories via transition localization, 3D F1, and tracked subject VLM scoring.
    Explicit alternative to symmetric frame-pair memory (Sec. 3.5, App. H).

pith-pipeline@v1.1.0-grok45 · 48253 in / 4022 out tokens · 45061 ms · 2026-07-12T10:01:18.558196+00:00 · methodology

0 comments
read the original abstract

Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We introduce WorldRoamBench, an open-world benchmark for long-horizon stability across four dimensions, each with tailored innovations: (i) Action: per-frame action metric bypassing cross-model semantic scale disparity and exposing failures hidden by trajectory; (ii) Vision: segment-based drift metric capturing non-monotonic mid-sequence collapse missed by start-vs-end comparisons; (iii) Physics: controllability-gated evaluation over mechanics, optics, and 3D consistency, scoring plausibility under faithful action execution; (iv) Memory: action-decoupled protocol evaluating scene memory via transition-localized 3D point-cloud reconstruction and subject memory via tracking-plus-VLM reasoning. The benchmark comprises 600+ test cases across Nature, Urban, and Indoor scenes in first/third-person views with WASD 10-60s continuous interaction. Evaluating 10+ open/closed-source models reveals none reliably satisfies all dimensions; even the best achieves only moderate scores. Advances on WorldRoamBench are steps toward IWMs that are stable, physically grounded, memory-faithful, and deployable in real-world applications.

Figures

Figures reproduced from arXiv: 2606.31672 by Baoquan Chen, Fan Jiang, Hongyu Pan, Jiacheng Sui, Kewei Shi, Mingchao Sun, Mu Xu, Qi Fan, Ting-Bing Xu, Wenjin Yang, Yang Gao, Yong Li, Zhaoxu Sun, Zhe Gao, Zhicheng Liu.

Figure 1
Figure 1. Figure 1: Illustration of WorldRoamBench. WorldRoamBench evaluates interactive world models across action following, visual quality, memory, and interaction physics, covering open-source and closed-source models in diverse scenarios and viewing conditions. Abstract Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at tra￾jectory level and ignore memory and … view at source ↗
Figure 2
Figure 2. Figure 2: Failures revealed by long-horizon interaction. Ex￾tended rollouts expose failures often missed by short clips: high trajectory scores hide per-step action mismatches, visual quality degrades, physical constraints are violated, and revisited scenes are regenerated inconsistently. Yet as these models proliferate, a critical question re￾mains unanswered: How well do they respond to user inputs over extended i… view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the WorldRoamBench evaluation pipeline. Given videos generated from shared initial frames and WASD/IJKL action programs, WorldRoamBench evaluates four complementary dimensions: action following through pose-aligned frame scoring, visual quality through frame-level and drift-aware metrics, interaction physics through mechanics, optics, and 3D-consistency tests, and memory through turning-point-a… view at source ↗
Figure 4
Figure 4. Figure 4: Action following evaluation pipeline. A shared ViPE trajectory estimation stage feeds two branches: (left) per-frame ac￾tion accuracy via latent-stride discretization, and (right) TrajScore via adaptive GT construction and arc-length resampling. an atomic action (e.g., forward) or a compound action. Each action is decomposed into a set of atomic sub-actions P(gt) = {p1, p2, . . .}. The atomic vocabulary co… view at source ↗
Figure 5
Figure 5. Figure 5: Memory evaluation pipeline. WorldRoamBench eval￾uates memory with two trajectory-aware tracks. Scene memory localizes the executed observation–revisit transition and compares reconstructed scene geometry across the two segments. Subject memory tracks the third-person protagonist and evaluates identity, structure, and appearance preservation with a holistic Qwen3-VL￾PLUS judgment. The full pipeline is provi… view at source ↗
Figure 6
Figure 6. Figure 6: Action key distribution across the three evaluation di￾mensions. Each pie chart shows the proportion and count of per￾frame key presses aggregated over all test cases in that dimension. 0 50 100 150 200 Number of Test Cases Memory Action Physics 99 113 212 118 99 217 96 75 171 (a) Case Counts FPV TPV 0 300 300 400 400 500 500 1000 1000 1100 0 100 200 300 400 Number of Cases 115 349 107 19 10 (b) Action Seq… view at source ↗
Figure 7
Figure 7. Figure 7: Test case count and action-sequence length distribu￾tion. (a) Test cases by dimension, split by 1st and 3rd perspective. (b) Distribution of action-sequence lengths in frames. related artifacts. 4.3. Test Suite Statistics [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 9
Figure 9. Figure 9: Per-frame visual quality curves and drift comparison (first-person). Top: mean imaging and aesthetic scores at each frame index (averaged over 120 videos per model, first 300 frames). Bottom: drift scores quantifying quality degradation from the best to the worst segment of the rollout (lower = more stable); bar height is the mean drift across videos, and black error bars span from mean−1 std (lower cap) t… view at source ↗
Figure 10
Figure 10. Figure 10: Per-frame action accuracy by difficulty (first-person). Strict accuracy (left) and partial accuracy (right) grouped by difficulty level (easy = constant action, medium = 1 action switch, hard = 2 action switches). Bar height is the mean across test cases. Most models degrade from easy to hard, but the magnitude of degradation varies substantially. stantially stronger third-person control than the open mod… view at source ↗
Figure 11
Figure 11. Figure 11: Test suite gallery. Each panel shows a representative test case with the first frame, trajectory schematic, and action sequence. The gallery is organized by scene category from top to bottom (Indoor, Urban, Nature) and by perspective from left to right (first-person, third-person), yielding six panels. checkpoint, which generates frames conditioned on both past and future context within each chunk, and ev… view at source ↗
Figure 12
Figure 12. Figure 12: Closed-source automated interaction pipeline for Happy Oyster and Genie 3. The system reads benchmark test cases (action [PITH_FULL_IMAGE:figures/full_fig_p021_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Latent-stride single-frame discretization. [PITH_FULL_IMAGE:figures/full_fig_p022_13.png] view at source ↗
Figure 15
Figure 15. Figure 15: Detailed interaction-physics evaluation pipeline. The full pipeline includes action-responsiveness validation, me￾chanics protocols, optics protocols, 3D-consistency evaluation, Qwen3-VL-PLUS-based VLM scoring, fallback queries, and domain-level aggregation. execute the prescribed action rather than remaining static or unresponsive. This prevents action-following failures from being conflated with physica… view at source ↗
Figure 16
Figure 16. Figure 16: Detailed memory evaluation pipeline. The full pipeline includes action-aware transition localization, point-cloud reconstruc￾tion, registration, filtering, scene-memory scoring, subject tracking, controllability gating, and Qwen3-VL-PLUS-based VLM subject￾memory evaluation. camera-pose chain. Pose drift can introduce a rigid dis￾placement unrelated to memory, artificially increasing both forgetting and ha… view at source ↗
Figure 17
Figure 17. Figure 17: Failure mode of frame-pair memory evaluation. Frame-pair metrics assume that an observation frame and its re￾visit counterpart depict the same spatial location. In long-horizon interactive rollouts, however, imperfect action execution can shift the revisit trajectory, so the paired frames correspond to different viewpoints or even different scene regions. Image-level discrep￾ancies in such pairs therefore… view at source ↗
Figure 18
Figure 18. Figure 18: Action gap in closed-source models. Top: Genie 3. When the D key (rightward translation) is issued, the model rotates the camera to the right instead of translating the character rightward within a stable scene. The character remains centered while the entire scene rotates, making controlled navigation impossible. Bottom: Happy Oyster. When the D key (rightward translation) is issued, the model faithfully… view at source ↗
Figure 19
Figure 19. Figure 19: Qualitative examples of WorldRoamBench scoring across physics, memory, and action. Each block contrasts a positive and a negative rollout that share the same initial frame and action schedule. Physics: W+A in a corridor—Happy Oyster (No Clipping) vs. Genie 3 (Clipping). Memory: J→L Observation/Revisit in an alley—Matrix-Game 3.0 (Good Memory) vs. SANA-WM (Bad Memory). Action: S→A (backward→left)—Genie 3 (… view at source ↗
Figure 20
Figure 20. Figure 20: Collision evaluation prompts. Full VLM prompts used for approach detection and collision-response verification. Deformation Evaluation Prompt — FPV: Trace Detection You are analyzing a video clip ({duration}s total). Showing {N} frames at timestamps: [{t1, t2, ..., tN}]. Think step by step about the following question. Question: In these frames from the later part of a first-person walking video, look car… view at source ↗
Figure 21
Figure 21. Figure 21: Deformation evaluation prompts. Full VLM prompts used for first-person trace detection, third-person reference comparison, and third-person real-time interaction checks. 35 [PITH_FULL_IMAGE:figures/full_fig_p035_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: Clipping evaluation prompts. Full VLM prompts used by the staged clipping-detection cascade. 36 [PITH_FULL_IMAGE:figures/full_fig_p036_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: Gravity evaluation prompts. Full VLM prompts used for scenario-specific gravity checks and the third-person general gravity check. 37 [PITH_FULL_IMAGE:figures/full_fig_p037_23.png] view at source ↗
Figure 24
Figure 24. Figure 24: Terrain-following evaluation prompts. Full VLM prompts used for naturalness, third-person ground contact, and boundary checks. 38 [PITH_FULL_IMAGE:figures/full_fig_p038_24.png] view at source ↗
Figure 25
Figure 25. Figure 25: Reflection evaluation prompts. Full VLM prompts used for coarse reflection plausibility checks and per-frame reflection scoring. 39 [PITH_FULL_IMAGE:figures/full_fig_p039_25.png] view at source ↗
Figure 26
Figure 26. Figure 26: Occlusion and shadow evaluation prompts. Full VLM prompts used for expected-shadow, fallback shadow, and generic shadow checks. 40 [PITH_FULL_IMAGE:figures/full_fig_p040_26.png] view at source ↗
Figure 27
Figure 27. Figure 27: Subject memory evaluation prompt. Full VLM prompt used for holistic third-person subject-memory scoring and diagnostic flag extraction. 41 [PITH_FULL_IMAGE:figures/full_fig_p041_27.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

    cs.CV 2026-08 conditional novelty 6.0

    Across 1,474 cases and 20 models, WorldExam shows that video world models split along paradigm lines — camera-, action-, and language-driven models each dominate one capability, and none combines strong reactivity wit...

Reference graph

Works this paper leans on

60 extracted references · 11 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Genie: Generative interactive environments

    Bruce, J., Dennis, M., Edwards, A., et al. Genie: Generative interactive environments. InICML, 2024

  2. [2]

    Genie 3: A new frontier for world models.https://deepmind.google/discover/ blog/genie- 3- a- new- frontier- for- world- models/, 2025

    Google DeepMind. Genie 3: A new frontier for world models.https://deepmind.google/discover/ blog/genie- 3- a- new- frontier- for- world- models/, 2025

  3. [3]

    Happy Oyster: An open-ended world model for real-time world creation and interaction.https:// happyoyster.cn/, 2026

    Alibaba Group. Happy Oyster: An open-ended world model for real-time world creation and interaction.https:// happyoyster.cn/, 2026

  4. [4]

    Wan 2.1: A comprehensive and unified video generation model.arXiv preprint arXiv:2503.20314, 2025

    Wan Team. Wan 2.1: A comprehensive and unified video generation model.arXiv preprint arXiv:2503.20314, 2025

  5. [5]

    Kling 3.0: Next-generation AI video generation

    Kuaishou. Kling 3.0: Next-generation AI video generation. https://klingai.com/, 2025

  6. [6]

    Sora 2: A large-scale video generation model

    OpenAI. Sora 2: A large-scale video generation model. https://openai.com/sora/, 2025

  7. [7]

    Veo 3: State-of-the-art video generation

    Google DeepMind. Veo 3: State-of-the-art video generation. https : / / deepmind . google / technologies / veo/, 2025

  8. [8]

    Seedance 2.0: Advancing video generation for world complexity.arXiv preprint arXiv:2506.05218, 2025

    ByteDance Seed Team. Seedance 2.0: Advancing video generation for world complexity.arXiv preprint arXiv:2506.05218, 2025

  9. [9]

    Matrix-Game 2.0: An open-source, real-time, and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025

    Matrix-Game Team. Matrix-Game 2.0: An open-source, real-time, and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025

  10. [10]

    Matrix-Game 3.0: Real-time and streaming interactive world model with long-horizon mem- ory.arXiv preprint, 2026

    Matrix-Game Team. Matrix-Game 3.0: Real-time and streaming interactive world model with long-horizon mem- ory.arXiv preprint, 2026

  11. [11]

    HY-World 1.5: A systematic framework for interactive world modeling with real-time latency and ge- ometric consistency.arXiv preprint, 2025

    HY-World Team. HY-World 1.5: A systematic framework for interactive world modeling with real-time latency and ge- ometric consistency.arXiv preprint, 2025

  12. [12]

    Yume 1.5: A text-controlled interactive world generation model.arXiv preprint, 2026

    Yume Team. Yume 1.5: A text-controlled interactive world generation model.arXiv preprint, 2026

  13. [13]

    Advancing open-source world models.arXiv preprint, 2026

    LingBot Team. Advancing open-source world models.arXiv preprint, 2026. 16

  14. [14]

    MIND: Benchmarking memory consis- tency and action following in world models.arXiv preprint arXiv:2602.08025, 2026

    Ye, H., Lu, J., et al. MIND: Benchmarking memory consis- tency and action following in world models.arXiv preprint arXiv:2602.08025, 2026

  15. [15]

    WorldMark: A unified benchmark suite for interactive video world models.arXiv preprint arXiv:2604.21686, 2026

    Alaya Studio. WorldMark: A unified benchmark suite for interactive video world models.arXiv preprint arXiv:2604.21686, 2026

  16. [16]

    iWorld-Bench: A benchmark for interactive world models with a unified action generation framework

    Li, Y ., et al. iWorld-Bench: A benchmark for interactive world models with a unified action generation framework. InICML, 2026

  17. [17]

    WildWorld: A large-scale dataset for dynamic world modeling with actions and explicit state toward gener- ative ARPG.arXiv preprint arXiv:2603.23497, 2026

    Shanda AI. WildWorld: A large-scale dataset for dynamic world modeling with actions and explicit state toward gener- ative ARPG.arXiv preprint arXiv:2603.23497, 2026

  18. [18]

    VBench: Comprehensive benchmark suite for video generative mod- els

    Huang, Z., He, Y ., Yu, J., Zhang, F., Si, C., et al. VBench: Comprehensive benchmark suite for video generative mod- els. InCVPR, 2024

  19. [19]

    VBench++: Comprehensive and versatile benchmark suite for video gen- erative models.arXiv preprint arXiv:2411.13503, 2024

    Huang, Z., Zhang, F., Xu, X., He, Y ., Yu, J., et al. VBench++: Comprehensive and versatile benchmark suite for video gen- erative models.arXiv preprint arXiv:2411.13503, 2024

  20. [20]

    VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025

    Zheng, D., Huang, Z., Liu, H., Zou, K., He, Y ., et al. VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness.arXiv preprint arXiv:2503.21755, 2025

  21. [21]

    WorldScore: A unified evaluation benchmark for world generation

    Stanford. WorldScore: A unified evaluation benchmark for world generation. InICCV, 2025

  22. [22]

    WorldModelBench: Judging video generation models as world models

    UC Berkeley. WorldModelBench: Judging video generation models as world models. InNeurIPS, 2025

  23. [23]

    Towards world simulator: Crafting physical commonsense-based benchmark for video generation

    SJTU, et al. Towards world simulator: Crafting physical commonsense-based benchmark for video generation. In ICML, 2025

  24. [24]

    WorldBench: Disambiguating physics for di- agnostic evaluation of world models.arXiv preprint arXiv:2601.21282, 2026

    UCLA. WorldBench: Disambiguating physics for di- agnostic evaluation of world models.arXiv preprint arXiv:2601.21282, 2026

  25. [25]

    Video PreTraining (VPT): Learning to act by watching unlabeled online videos

    Baker, B., et al. Video PreTraining (VPT): Learning to act by watching unlabeled online videos. InNeurIPS, 2022

  26. [26]

    MineWorld: A real-time and open-source interactive world model on Minecraft.arXiv preprint arXiv:2504.08388, 2025

    Microsoft. MineWorld: A real-time and open-source interactive world model on Minecraft.arXiv preprint arXiv:2504.08388, 2025

  27. [27]

    SANA-WM: Efficient minute-scale world model- ing with hybrid linear diffusion transformer.arXiv preprint, 2026

    NVIDIA. SANA-WM: Efficient minute-scale world model- ing with hybrid linear diffusion transformer.arXiv preprint, 2026

  28. [28]

    Lyra 2.0: Explorable generative 3D worlds.arXiv preprint, 2026

    NVIDIA. Lyra 2.0: Explorable generative 3D worlds.arXiv preprint, 2026

  29. [29]

    minWM: A full-stack open-source frame- work for real-time interactive video world models.arXiv preprint, 2026

    minWM Team. minWM: A full-stack open-source frame- work for real-time interactive video world models.arXiv preprint, 2026

  30. [30]

    LAION- 5B: An open large-scale dataset for training next generation image-text models

    Schuhmann, C., Beaumont, R., Vencu, R., et al. LAION- 5B: An open large-scale dataset for training next generation image-text models. InNeurIPS, 2022

  31. [31]

    MUSIQ: Multi-scale image quality transformer

    Ke, J., Wang, Q., Wang, Y ., Milanfar, P., and Yang, F. MUSIQ: Multi-scale image quality transformer. InICCV, 2021

  32. [32]

    Helios: A comprehensive benchmark for video generative models.arXiv preprint, 2025

    Helios Team. Helios: A comprehensive benchmark for video generative models.arXiv preprint, 2025

  33. [33]

    WorldCompass: Reinforcement learning for long-horizon world models.arXiv preprint, 2026

    WorldCompass Team. WorldCompass: Reinforcement learning for long-horizon world models.arXiv preprint, 2026

  34. [34]

    ViPE: Visual pose estimation for camera trajectory recovery

  35. [35]

    and Deng, J

    Teed, Z. and Deng, J. DROID-SLAM: Deep visual SLAM for monocular, stereo, and RGB-D cameras. InNeurIPS, 2021

  36. [36]

    WBench: A comprehensive benchmark for evaluating world models via action-conditioned video gener- ation.arXiv preprint, 2026

    Ying, Z., et al. WBench: A comprehensive benchmark for evaluating world models via action-conditioned video gener- ation.arXiv preprint, 2026

  37. [37]

    and Medioni, G

    Chen, Y . and Medioni, G. Object modelling by registra- tion of multiple range images.Image and Vision Computing, 10(3):145–155, 1992

  38. [38]

    B., Blodow, N., and Beetz, M

    Rusu, R. B., Blodow, N., and Beetz, M. Fast Point Feature Histograms (FPFH) for 3D registration. InICRA, 2009

  39. [39]

    Point Transformer V3: Simpler, faster, stronger

    Wu, X., Jiang, L., Wang, P.-S., et al. Point Transformer V3: Simpler, faster, stronger. InCVPR, 2024

  40. [40]

    SAM 2: Seg- ment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024

    Ravi, N., Gabeur, V ., Hu, Y .-T., et al. SAM 2: Seg- ment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024

  41. [41]

    Two-frame motion estimation based on poly- nomial expansion

    Farneb ¨ack, G. Two-frame motion estimation based on poly- nomial expansion. InScandinavian Conference on Image Analysis (SCIA), 2003

  42. [42]

    What limits virtual agent appli- cation? OmniBench: A scalable multi-dimensional bench- mark for essential virtual agent capabilities

    Bu, W., Wu, Y ., Yu, Q., et al. What limits virtual agent appli- cation? OmniBench: A scalable multi-dimensional bench- mark for essential virtual agent capabilities. InICML, 2025

  43. [43]

    Omni-WorldBench: Towards a comprehensive interaction-centric evaluation for world models.arXiv preprint arXiv:2603.22212, 2026

    Wu, M., Cai, Z., Zhao, F., et al. Omni-WorldBench: Towards a comprehensive interaction-centric evaluation for world models.arXiv preprint arXiv:2603.22212, 2026

  44. [44]

    Do vision-language models have internal world models? Towards an atomic evaluation

    Gao, Q., Pi, X., Liu, K., et al. Do vision-language models have internal world models? Towards an atomic evaluation. InACL, 2025

  45. [45]

    How far is video generation from world model: A physical law perspective

    Kang, B., Yue, Y ., Lu, R., et al. How far is video generation from world model: A physical law perspective. InICML, 2025

  46. [46]

    WorldLens: Full-spectrum evaluations of driving world models in real world.arXiv preprint arXiv:2512.10958, 2025

    Liang, A., Kong, L., Yan, T., et al. WorldLens: Full-spectrum evaluations of driving world models in real world.arXiv preprint arXiv:2512.10958, 2025

  47. [47]

    WorldArena: A unified benchmark for evaluating perception and functional utility of embodied world models.arXiv preprint arXiv:2602.08971, 2026

    Shang, Y ., Li, Z., Ma, Y ., et al. WorldArena: A unified benchmark for evaluating perception and functional utility of embodied world models.arXiv preprint arXiv:2602.08971, 2026

  48. [48]

    ACT-Bench: Towards action controllable world models for autonomous driving.arXiv preprint arXiv:2412.05337, 2024

    Arai, H., Ishihara, K., Takahashi, T., and Yamaguchi, Y . ACT-Bench: Towards action controllable world models for autonomous driving.arXiv preprint arXiv:2412.05337, 2024

  49. [49]

    EWMBench: Evaluat- ing scene, motion, and semantic quality in embodied world models

    Hu, Y ., Huang, S., Liao, Y ., et al. EWMBench: Evaluat- ing scene, motion, and semantic quality in embodied world models. InBMVC, 2025

  50. [50]

    Worl- dOlympiad: Can Your World Model Survive a Triathlon? arXiv preprint arXiv:2606.11129, 2026

    Zhao, Y ., Zhao, W., Wang, W., Zhang, Z., An, D., Liu, A., Yu, Y ., Tang, J., Wang, F., Wang, W., and Zhuang, B. Worl- dOlympiad: Can Your World Model Survive a Triathlon? arXiv preprint arXiv:2606.11129, 2026

  51. [51]

    Person moves forward; Camera turns left

    Teed, Z. and Deng, J. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow.arXiv preprint arXiv:2003.12039, 2020. 17 A. Test Suite Gallery To provide a qualitative overview of the visual and ac- tion coverage in WorldRoamBench, we include a gallery of representative test cases in Figure 11. Each panel shows the first-frame image together with an ac...

  52. [52]

    camera-pose chain

    Depth-percentile Filteringretain nearest p%Far Near Cross-Segment Registration Coarse:PTY3 / FPFH+ RANSACFine:point-to-planeICPQuality-awareacceptance:fitness ≥ τᵩand Chamferimprovest … Video + GTActions⋯WW(Explore) SS(Revisit)⋯ Observation Segment(t ≤ t*) Frame-LevelPointClouds t Revisit Segment(t > t*) t Frame-LevelPointClouds t*Memory Evaluation Pipeli...

  53. [53]

    Holistic scoring accommodates smooth viewpoint and illumination changes that can confound frame-level com- parisons, while producing a single benchmark-compatible scalar without requiring an external reference-feature li- brary. G. Subject Memory Prompt For third-person memory evaluation, the Subject Memory Evaluation Prompt in Figure 27 scores video-leve...

  54. [54]

    Is there a shadow visible on the wall or floor?

  55. [55]

    If yes, does the shadow move or change in a way that is physically consistent with the camera movement and the light source position? Answer ’yes’ if the shadow appears and behaves correctly, ’no’ if the shadow is missing or behaves incorrectly. After your reasoning, conclude with exactly one line: Answer: yes or Answer: no Figure 26.Occlusion and shadow ...

  56. [56]

    Identity change: the subject becomes a different individual or category

  57. [57]

    Structural distortion: the body, anatomy, proportions, limbs, head, face, or key parts become deformed or implausible

  58. [58]

    Appearance drift: color, texture, clothing, hair, material, or style is truly rewritten

  59. [59]

    Subject disappearance: the subject becomes partly or fully invisible

  60. [60]

    subject memory score

    Quality degradation: severe blur, low resolution, diffused edges, or loss of key details makes the subject hard to identify. Important calibration rules: - Viewpoint change alone is not inconsistency. Back-to-front, front-to-back, side-to-front, or far-to-close changes are acceptable if the subject can reasonably be the same individual under the new view....