Pith. sign in

REVIEW 23 cited by

WorldSimBench: Towards Video Generation Models as World Simulators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.18072 v1 pith:6LJX6CDU submitted 2024-10-23 cs.CV

classification cs.CV
keywords embodiedevaluationmodelssimulatorsworldhumanpredictivevideo
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress of predictive model development. Additionally, existing benchmarks are unable to effectively evaluate higher-capability, highly embodied predictive models from an embodied perspective. In this work, we classify the functionalities of predictive models into a hierarchy and take the first step in evaluating World Simulators by proposing a dual evaluation framework called WorldSimBench. WorldSimBench includes Explicit Perceptual Evaluation and Implicit Manipulative Evaluation, encompassing human preference assessments from the visual perspective and action-level evaluations in embodied tasks, covering three representative embodied scenarios: Open-Ended Embodied Environment, Autonomous, Driving, and Robot Manipulation. In the Explicit Perceptual Evaluation, we introduce the HF-Embodied Dataset, a video assessment dataset based on fine-grained human feedback, which we use to train a Human Preference Evaluator that aligns with human perception and explicitly assesses the visual fidelity of World Simulators. In the Implicit Manipulative Evaluation, we assess the video-action consistency of World Simulators by evaluating whether the generated situation-aware video can be accurately translated into the correct control signals in dynamic environments. Our comprehensive evaluation offers key insights that can drive further innovation in video generation models, positioning World Simulators as a pivotal advancement toward embodied artificial intelligence.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation

    cs.RO 2026-08 conditional novelty 7.0 of 10

    WorldSimProbe is a five-suite benchmark showing that six action-conditioned world models systematically degrade in action-to-motion fidelity and interaction grounding across RoboTwin, ManiSkill, and LIBERO.

  2. KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding

    cs.RO 2026-07 conditional novelty 7.0 of 10

    KineBench evaluates embodied world models by extracting 6D robot-pose trajectories from generated videos, executing them in a physics simulator, and scoring them with SPARC and manipulability metrics.

  3. MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    None of ten tested video-generation models reliably remembers objects after occlusion in dynamic scenes; static-camera videos inflate consistency scores.

  4. What if? Emulative Simulation with World Models for Situated Reasoning

    cs.CV 2026-03 conditional novelty 6.5 of 10

    WanderDream supplies 15.8K panoramic mental-exploration videos and 158K QA pairs showing that world-model imagination measurably improves situated spatial reasoning without active exploration.

  5. WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Across 1,474 cases and 20 models, WorldExam shows that video world models split along paradigm lines — camera-, action-, and language-driven models each dominate one capability, and none combines strong reactivity wit...

  6. DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

    cs.LG 2026-07 accept novelty 6.0 of 10

    Schema-guided interleaved state-transition pretraining with selective attention and reweighted loss improves hierarchical visual dynamics modeling for narrative generation and world simulation.

  7. GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Long-horizon action-faithful consistency, not short-term visual realism, dominates world-model reliability for robot policy evaluation; GigaWorld-1 implements that roadmap and gains 14.9% on evaluator-alignment metrics.

  8. Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A 3D-aware VLM, RoboTracer, generates metric-grounded spatial traces for robot manipulation using scale supervision and metric-sensitive reinforcement rewards.

  9. BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

    cs.CL 2025-07 conditional novelty 6.0 of 10

    BMMR provides a 110k-question bilingual, multimodal, college-level dataset across 300 subjects where state-of-the-art models score at most about 50%.

  10. CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Causal Diffusion Policy adds historical action conditioning and attention cache sharing to diffusion-based robot policies, improving success rates on most tested manipulation tasks under degraded observations.

  11. From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Frozen CogVideoX1.5, adapted with LoRA on 3 to 30 input-output videos, performs segmentation, pose estimation, and abstract reasoning (ARC-AGI 16.75%) with modest but real generalization.

  12. Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Using an inverse dynamics model to label and verify frame transitions lets vision-language models learn forward dynamics and outperform specialized image editors on Aurora-Bench.

  13. Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Context-as-Memory conditions video generation on selected historical frames chosen by camera FOV overlap, improving scene consistency in long generated videos.

  14. MOVi: Training-free Text-conditioned Multi-Object Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MOVi improves multi-object video generation without retraining by using LLM-planned trajectories to reinitialize the diffusion noise and by reweighting attention to stop objects from mixing together.

  15. DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion model fuses per-person masks with pose keypoints to generate identity-preserving, two-person interactive videos from a single reference image, outperforming prior single-person-animation pipelines.

  16. Generative Physical AI in Vision: A Survey

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.

  17. Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new story-completion benchmark, StoryEval, shows that 11 current text-to-video models complete fewer than half of the consecutive events in short story prompts.

  18. Owl-1: Omni World Model for Consistent Long Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Owl-1 generates long, multi-scene videos by using a language model to maintain a latent state and predict text dynamics, then rendering each clip with a video diffusion model.

  19. Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A VLM-generated code monitor tracks geometric elements and evaluates spatio-temporal constraints, enabling real-time reactive and proactive failure detection for robotic manipulation.

  20. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  21. A Definition and Roadmap for World Models

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A perspective article defining world models as finite-resource compression of physical state transitions and outlining a roadmap toward physical AGI via unified representations and interactive simulators.

  22. Towards Conscious Service Robots

    cs.RO 2025-01 unverdicted novelty 2.0 of 10

    Service robots should be built with a two-system cognitive architecture, including a global workspace and metacognitive monitoring, to generalize to novel situations.

  23. A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

    cs.CV 2025-01 reject novelty 2.0 of 10

    A survey that catalogs large vision-language models, their alignment methods, benchmarks, and challenges, but is compromised by inconsistent counts and misclassified entries.

Pith tools