Pith. sign in

REVIEW 22 cited by

Intuitive physics understanding emerges from self-supervised pretraining on natural videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11831 v1 pith:CXGD6G4V submitted 2025-02-17 cs.CV cs.AI

Intuitive physics understanding emerges from self-supervised pretraining on natural videos

classification cs.CV cs.AI
keywords intuitivephysicsunderstandingmodelsspacetrainedvideoachieve
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We investigate the emergence of intuitive physics understanding in general-purpose deep neural network models trained to predict masked regions in natural videos. Leveraging the violation-of-expectation framework, we find that video prediction models trained to predict outcomes in a learned representation space demonstrate an understanding of various intuitive physics properties, such as object permanence and shape consistency. In contrast, video prediction in pixel space and multimodal large language models, which reason through text, achieve performance closer to chance. Our comparisons of these architectures reveal that jointly learning an abstract representation space while predicting missing parts of sensory input, akin to predictive coding, is sufficient to acquire an understanding of intuitive physics, and that even models trained on one week of unique video achieve above chance performance. This challenges the idea that core knowledge -- a set of innate systems to help understand the world -- needs to be hardwired to develop an understanding of intuitive physics.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

    cs.LG 2026-07 conditional novelty 7.0

    Prediction targets, not fused inputs, decide which physical parameters enter a latent world model; slow ratio-type parameters like drag remain largely unacquired despite high recoverability certificates.

  2. What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

    cs.LG 2026-07 conditional novelty 7.0

    In latent world models, prediction targets—not input sensors or data volume—determine which physical parameters the learned representation contains; drag remains systematically unlearned by deterministic prediction ob...

  3. Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

    cs.CV 2026-07 conditional novelty 7.0

    A law-grounded benchmark, Apple-PI, grades video models stage-by-stage on physics reasoning and finds they top out at 0.473, well short of reliable simulation.

  4. LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")

    cs.LG 2026-06 unverdicted novelty 7.0

    VLMs show partial alignment with children's performance on six cognitive tasks, with stronger models matching better at task and item levels but struggling on matrix reasoning and mental rotation.

  5. YoCausal: How Far is Video Generation from World Model? A Causality Perspective

    cs.CV 2026-05 unverdicted novelty 7.0

    YoCausal benchmark shows video diffusion models detect the arrow of time but lack genuine causal understanding relative to humans.

  6. Emergent Compositional Communication for Latent World Properties

    cs.MA 2026-03 conditional novelty 7.0

    Multi-agent iterated learning produces emergent positionally disentangled communication protocols for latent physical properties from unsupervised video features.

  7. What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

    cs.LG 2026-07 conditional novelty 6.0

    Prediction targets, not inputs, decide which physical parameters a latent world model acquires; a certified-recoverable drag parameter stays unlearned under every deterministic prediction objective tested.

  8. Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection

    cs.CV 2026-06 unverdicted novelty 6.0

    A self-supervised framework learns implicit 3D physics by lifting V-JEPA features into voxels and performing volumetric feature advection conditioned on actions.

  9. Asymmetric physics enables efficient learning in quadrupedal robot swarms

    cs.RO 2026-06 unverdicted novelty 6.0

    Asymmetric physics (high-fidelity non-diff simulator plus differentiable surrogates) enables end-to-end training of decentralized vision-based policies for up to 512 quadrupeds that transfer zero-shot to real hardware.

  10. LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    LaMo adds self-supervised latent motion priors via a motion drift loss during training and motion prior guidance during sampling to boost physical fidelity in video diffusion models like CogVideoX.

  11. Latent Video Prediction Learns Better World Models

    cs.CV 2026-05 unverdicted novelty 6.0

    Latent prediction video models exhibit a distinct robustness profile across corruption, occlusion, fine-grained discrimination, and temporal sensitivity compared to other self-supervised video models when used as worl...

  12. LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

    cs.LG 2026-03 unverdicted novelty 6.0

    LeWM is a ~15M-parameter JEPA world model that trains end-to-end from pixels with only next-embedding prediction plus a Gaussian latent regularizer, cutting loss hyperparameters to one.

  13. LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

    cs.LG 2026-03 unverdicted novelty 6.0

    LeWM is the first end-to-end trainable JEPA from pixels that uses only two loss terms for stable training and fast planning on 2D/3D control tasks.

  14. VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction

    cs.CV 2026-02 unverdicted novelty 6.0

    VisPhyWorld evaluates MLLMs' physical reasoning via executable code generation for video reconstruction, with VisPhyBench showing strong semantics but weak parameter inference and dynamics simulation.

  15. Cambrian-S: Towards Spatial Supersensing in Video

    cs.CV 2025-11 unverdicted novelty 6.0

    Cambrian-S introduces VSI-SUPER benchmarks for long-horizon spatial recall and counting, shows data scaling yields 30% gains on existing tests, and demonstrates a self-supervised next-latent predictor using surprise o...

  16. Video models are zero-shot learners and reasoners

    cs.LG 2025-09 unverdicted novelty 6.0

    Generative video models exhibit emergent zero-shot capabilities across perception, manipulation, and basic reasoning tasks.

  17. PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

    cs.CV 2026-07 conditional novelty 5.0

    PAVXploreRL post-trains action-conditioned world models with VJEPA-2 latent rewards and perturbed 'OOD' actions, reporting a 5.6% average gain and lowered policy-overestimation bias.

  18. Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

    cs.CV 2026-06 unverdicted novelty 5.0

    Video foundation models encode intuitive physics knowledge that is strongest in V-JEPA at intermediate-to-late layers and depends on pretraining type and probe design.

  19. NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results

    cs.CV 2026-04 unverdicted novelty 5.0

    The NTIRE 2026 Challenge released a public dataset of 2,000 videos with crowdsourced saliency maps and reported results from participating teams using standard quality metrics.

  20. Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

    cs.CV 2026-04 unverdicted novelty 5.0

    Phantom jointly models visual content and latent physical dynamics via a physics-aware video representation to generate physically consistent videos.

  21. Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics

    cs.CV 2026-04 unverdicted novelty 5.0

    Phantom generates visually realistic and physically consistent videos by jointly modeling visual content and latent physical dynamics via an abstract physics-aware representation.

  22. How VLAs (Really) Work In Open-World Environments

    cs.RO 2026-04 unverdicted novelty 4.0

    Standard success metrics for VLAs on complex chores overlook safety violations and intermediate failures, leading to exaggerated claims; new evaluation protocols are proposed to measure robustness and safety.