Pith. sign in

REVIEW 16 cited by

PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.09595 v1 pith:ITLQ6KES submitted 2025-03-12 cs.CV

classification cs.CV
keywords modelingmodelsvideopost-trainingtaskaccurategenerationlarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale pre-trained video generation models excel in content creation but are not reliable as physically accurate world simulators out of the box. This work studies the process of post-training these models for accurate world modeling through the lens of the simple, yet fundamental, physics task of modeling object freefall. We show state-of-the-art video generation models struggle with this basic task, despite their visually impressive outputs. To remedy this problem, we find that fine-tuning on a relatively small amount of simulated videos is effective in inducing the dropping behavior in the model, and we can further improve results through a novel reward modeling procedure we introduce. Our study also reveals key limitations of post-training in generalization and distribution modeling. Additionally, we release a benchmark for this task that may serve as a useful diagnostic tool for tracking physical accuracy in large-scale video generative model development.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    PhysEditWorld is a new dataset of over 60 million frames from 12 UE5 cinematic scenes with synchronized multimodal signals and explicit gravity labels, built via replay to support physics-editable world models.

  2. Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    PQSG is a hierarchical question-graph pipeline for fine-grained physical plausibility evaluation in text-to-video generation that achieves higher correlation with human judgments than prior methods on the FinePhyEval dataset.

  3. YoCausal: How Far is Video Generation from World Model? A Causality Perspective

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    YoCausal benchmark shows video diffusion models detect the arrow of time but lack genuine causal understanding relative to humans.

  4. ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    ACWM-Phys is a controllable simulator benchmark with in- and out-of-distribution protocols for evaluating action-conditioned world models across rigid, kinematic, deformable, and particle dynamics.

  5. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Using executable Blender code as an intermediate simulation draft improves physical consistency in text-to-video generation, lifting OmniWeaving from 0.475 to 0.558 on PhyGenBench and from 52.18% to 77.88% on VBench-2.0.

  6. Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Explicit instance-level physical parameter conditioning with routing attention improves physical-law consistency in image-to-video generation, as measured on the authors' new simulator-based benchmark.

  7. Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    A multi-agent video world model using simplex rotary agent encoding and sparse hub attention achieves better fidelity, controllability, and consistency than baselines while generalizing from 2 to 4 players.

  8. ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    ACWM-Phys benchmark shows action-conditioned world models generalize on simple geometric interactions but drop sharply on deformable contacts, high-dimensional control, and complex articulated motion, indicating relia...

  9. VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A vision-language model writes a simulation program—grounded by segmentation and 3D tools—that predicts physically plausible futures from an image and caption, outperforming video generators on modified benchmarks.

  10. Epipolar Geometry Improves Video Generation Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Ranking generated videos by their epipolar (Sampson) error and fine-tuning Wan2.1 with Flow-DPO cuts epipolar error 31% and raises human-rated 3D consistency from 54% to 72%.

  11. Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility

    cs.CV 2025-09 unverdicted novelty 6.0 of 10

    A training-free framework uses physics-violating counterfactual prompts and Synchronized Decoupled Guidance to suppress implausible motions in diffusion-based video generation while preserving photorealism.

  12. PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    PhysRAG curates 7K videos from WISA-80K, builds a physical video database, and injects knowledge via learnable queries into a diffusion model to reach SOTA visual quality and physical compliance on PhyGenBench and VBench.

  13. PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    PhysEditWorld supplies 12 UE5 scenes, 60+ million frames, and explicit gravity labels via a replay paradigm to support gravity-faithful and physically editable world models.

  14. Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models

    cs.LG 2025-10 unverdicted novelty 5.0 of 10

    GRWM uses temporal contrastive learning to geometrically regularize latent spaces in world models for high-fidelity cloning of deterministic 3D worlds.

  15. From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    An explainable model using BLEURT, CometKiwi, pause features, and Chinese phraseological diversity predicts human-rated quality dimensions in English-Chinese consecutive interpreting, with SHAP identifying the stronge...

  16. Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A hierarchical direct preference optimization with four alignment levels plus automated data selection improves physical plausibility of text-to-video models.

Pith tools