Using an inverse dynamics model to label and verify frame transitions lets vision-language models learn forward dynamics and outperform specialized image editors on Aurora-Bench.
Video generation models as world simulators
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics
Using an inverse dynamics model to label and verify frame transitions lets vision-language models learn forward dynamics and outperform specialized image editors on Aurora-Bench.