REVIEW 4 major objections 5 minor 9 cited by
PhysMotion: Physics-Grounded Dynamics From a Single Image
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read PhysMotion turns a single image into a physically plausible video by simulating the object's 3D shape with continuum mechanics, then polishing the coarse simulation with a diffusion model.
desk verdict Plausible single-image physics I2V pipeline, but the paper never checks whether the diffusion stage preserves the simulated motion, so the headline claim outruns the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the physics-integrated 3D Gaussian representation: each Gaussian carries a deformation gradient F, so the MPM simulation updates the Gaussian centers and covariances through the continuum momentum equation, making the rendering itself differentiable. The second stage leverages DDIM+ inversion, ControlNet depth and edge conditioning, and TokenFlow-style cross-frame attention to propagate enhanced key-frame features through the whole video, which is what keeps the final frames visually consistent with the input image and with each other.
What would settle it
Take one generated scene, such as a falling and bouncing object, and track a few surface points in the final enhanced video; compare those trajectories to the trajectories of the corresponding MPM particles. If the enhanced video's motion diverges from the simulation beyond a small tolerance, or if the video fails to respond appropriately to a changed applied force, the paper's physics-grounded claim is not borne out.
Extended reading notes
Core claim
The central claim is that physics-grounded 3D dynamics can be generated from a single image by first reconstructing the foreground as a set of 3D Gaussians, time-stepping that representation with a differentiable Material Point Method using elastoplastic material models, and then using a text-to-image diffusion model with cross-frame attention to enhance the coarse simulated frames. The paper reports that this pipeline produces videos that score highest among compared image-to-video baselines on physical commonsense and semantic adherence metrics, and that it supports rigid motion, elastic deformation, viscoplastic flow, granular particles, and fracture from a single RGB image plus an applied force or torque.
Load-bearing premise
The diffusion-based enhancement stage must preserve the motion produced by the MPM simulation; if the enhancement noticeably changes trajectories, contact timing, or deformation, the resulting video is no longer physics-grounded.
Editorial extensions
If this is right
- A user can specify an applied force or torque and watch the object respond with deformable, rigid, granular, or fracturing motion, rather than merely interpolating pixels.
- Text- or trajectory-conditioned video models that produce physically implausible motion could be replaced or augmented by this simulation-first recipe in applications like visual effects and game asset animation.
- Because the pipeline starts from a single image, it extends physics-based video synthesis beyond the 3D-model or multi-view inputs that earlier physics-grounded methods require.
- The quantitative results position the method as a baseline for evaluating physical commonsense in image-to-video generation tasks.
Reading between the lines
- A direct testable extension is to measure how much the final enhanced video deviates from the coarse MPM simulation; if the deviation is small, the pipeline can double as a controllable physical simulator with photorealistic output.
- The same coarse-to-fine recipe could be applied to other physics simulators, such as rigid-body or fluid solvers, and to other single-image 3D representations, since the enhancement stage only needs a sequence of coarse frames.
- If the physics is preserved, PhysMotion's outputs could serve as pseudo-ground-truth for training or benchmarking video diffusion models on physical plausibility, an application the paper does not discuss.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PhysMotion, a framework that turns a single input image into a video of physics-driven object dynamics. The method first segments the foreground, reconstructs a 3D Gaussian representation using LGM with depth and color refinement, and then steps that representation forward with a Material Point Method (MPM) simulator using continuum-mechanics elastoplasticity models to produce a coarse video of the object. A diffusion-based enhancement stage, built on DDIM+ inversion, ControlNet depth/edge conditioning, and cross-frame attention, blends the coarse dynamics with the background and restores fine texture and visual quality. The authors evaluate against several image-to-video diffusion baselines using VideoPhy's Physical Commonsense (PC) and Semantic Adherence (SA) z-scores, a user study, and qualitative comparisons, reporting advantages on all quantitative criteria and demonstrating a range of materials and fracture effects. The central claim is that PhysMotion is the first single-image-to-video framework that combines 3D geometry awareness with physics-grounded dynamics, where the physics comes from the MPM simulation and the diffusion stage acts as a detail refiner.
Significance. If the central claim is fully supported, the paper would make a useful step toward controllable, physically plausible video synthesis from a single image, an area where purely data-driven video diffusion models often violate physical laws. The use of standard continuum mechanics and an off-the-shelf MPM solver gives the physics component a solid foundation, and the qualitative results, especially for fracture and elastoplastic deformation, are visually compelling and demonstrate versatility across materials. The method builds on a sensible combination of existing components (PhysGaussian-style physics-integrated 3DGS, LGM reconstruction, TokenFlow-style cross-frame attention, and ControlNet), and the paper is generally clearly written. However, the current manuscript does not release code or the 13-image test set, the quantitative evaluation is thin, and no measurement is provided that the diffusion enhancement preserves the simulated motion. Those gaps directly affect the strength of the 'physics-grounded' claim, and they can be addressed within the manuscript's scope.
major comments (4)
- [Section 3.4 and Section 4.2] The paper's central claim (Abstract and Section 1) is that the output video is physics-grounded because it is guided by MPM simulation. However, no experiment measures whether the diffusion enhancement stage preserves the simulated motion. The feature/attention injection timesteps tau_f and tau_A (Section 3.4.2) are free hyperparameters that control how much of the coarse simulation survives, and the quantitative metrics (VideoPhy PC/SA and the user study) measure plausibility and semantic adherence, not fidelity to the MPM trajectory. Please add a quantitative comparison between coarse and enhanced videos, for example by tracking keypoints or by comparing rendered foreground masks or optical flow of the enhanced video against the MPM output, and include an ablation that varies tau_f and tau_A and reports the resulting motion deviation. Without this, the claim that the enhancement is a detail refiner rather than a motion rewriter is unsupported.
- [Table 1 and Section 4.2] The user study results are not reported for the proposed method: the 'Ours' row of Table 1 contains only em-dashes for TC, PP, and Overall, yet the text states that 'our proposed method consistently outperforms all baselines across every evaluation criterion in the user preference results.' Please report the actual preference percentages and their statistical significance (for example, bootstrap confidence intervals or a paired test), and clarify whether the baseline percentages in Table 1 denote the preference rate for the proposed method over each baseline or the preference rate for each baseline over the proposed method. The current presentation does not support the claimed superiority.
- [Section 4.2 and Appendix A.3] The quantitative evaluation uses only 13 scenes, with per-scene z-score normalization across the six methods, and no error bars or significance tests are provided for the reported average z-scores. With 13 scenes, the reported differences (for example, PC z-score 0.5142 vs. 0.1853) may be within the noise, especially because z-score normalization forces each scene's mean to zero and discards the absolute score scale. Please report the per-scene scores (or include them as a table in the appendix), provide error bars or confidence intervals, and perform a significance test. In addition, the paper should state whether the 13-image test set and the evaluation code will be released, since independent verification is not currently possible.
- [Section 4.4 and Section 5] The ablation studies are qualitative only. In particular, the key claim that the enhancement stage is necessary and that it does not degrade physical motion is supported by a single visual example (the bread-tearing case in Fig. 6), while Section 5 concedes that the diffusion stage may introduce artifacts and color distortions. Please provide quantitative ablations for the enhancement stage, such as VideoPhy scores with and without enhancement as well as the motion-deviation metric described above, and for the hard-depth loss, ideally including reconstruction metrics on the input view in addition to the qualitative Fig. 8.
minor comments (5)
- [Figure 5] The caption cites 'I2VGen-XL [72]' but I2VGen-XL is reference [105] in the bibliography; the reference numbering is inconsistent between the text (Section 4.2 uses [105]) and the figure caption.
- [Section 3.4.2] The sentence 'The timesteps for feature and attention injection is controlled by two hyperparameters' has a subject-verb agreement error and should read 'are controlled.' Also, Eq. (12) introduces notation Q_{jk,c} and K_{j1,c} without explicitly connecting it to the previously defined Q_c and K_c, which makes the equational flow hard to follow.
- [Appendix A.2.2] The phrase 'we randomly choose key-frames every 5 frames' is ambiguous: if key-frames are selected every 5 frames, the selection is deterministic, not random, and the word 'randomly' should be removed or the random mechanism specified.
- [Section 4.1 and Appendix A.2] The implementation appendix gives training epochs and learning rates but does not specify the material model parameters (for example, Young's modulus, yield stress, or hardening parameters) used for each showcased material; please either list these parameters or state explicitly that they are inherited from the MPM solver or from PhysGaussian [91].
- [Section 2.1] The phrase 'simulation-based methods produce generative dynamics that are physically grounded [4]' cites VideoPhy [4] as support, but VideoPhy is an evaluation benchmark for physical commonsense, not a method that produces physically grounded dynamics; consider rephrasing to avoid conflating evaluation with generation.
Circularity Check
No significant circularity: the physics backbone is standard continuum mechanics and the diffusion enhancement is a separate stage; the unquantified motion-fidelity gap is a validation risk, not a circular reduction.
full rationale
PhysMotion's derivation chain is: single image -> LGM 3DGS reconstruction plus depth/color refinement -> MPM simulation using standard continuum mechanics (Eq. 3) -> coarse frames -> DDIM+ inversion with ControlNet and cross-frame attention (Sec. 3.4) -> final enhanced video. Each stage is defined externally to the paper's claims: the momentum equation is textbook conservation of momentum, the 3DGS deformation coupling follows prior published work (PhysGaussian [91]), and the MPM solver [113] is an open-sourced implementation. The enhancement stage is a separate diffusion process; the paper does not claim the final pixels equal the MPM output, and Eq. (12)-(15) describe feature/attention injection that steers the diffusion rather than a fitted identity. The absence of a quantitative fidelity metric between the enhanced video and the simulated motion is a genuine validation gap (a correctness risk), but it is not circularity: 'physics-grounded' is not defined as 'identical to MPM output,' and no parameter is fitted to the VideoPhy benchmark. VideoPhy [4] is a public external evaluator; author overlap does not make its scores an input to the method. No equation in the paper reduces a prediction to a fitted parameter, to a self-citation, or to its own input by construction. Therefore no circular steps are present.
Assumptions & free parameters
free parameters (4)
- Material model and parameters per scene =
not specified
- Applied force/torque per scene =
not specified
- Depth loss weighting lambda_depth =
not specified
- Feature/attention injection timesteps tau_f, tau_A =
not specified
assumptions (4)
- domain assumption The MPM with continuum elastoplasticity models conserves mass and momentum and is a valid prior for object dynamics.
- domain assumption LGM's single-view 3DGS reconstruction is sufficiently accurate for physics simulation.
- domain assumption VideoPhy's Physical Commonsense score is a valid measure of physical plausibility.
- ad hoc to paper The diffusion enhancement does not break the simulated motion.
Cite this review
Pith. "Pith review of PhysMotion: Physics-Grounded Dynamics From a Single Image." pith.science (2026). https://pith.science/paper/GJLGBLSQ
@misc{pith2026241117189,
author = {Pith},
title = {Pith review of: PhysMotion: Physics-Grounded Dynamics From a Single Image},
year = {2026},
howpublished = {\url{https://pith.science/paper/GJLGBLSQ}},
note = {Machine review of arXiv:2411.17189}
}
read the original abstract
We introduce PhysMotion, a novel framework that leverages principled physics-based simulations to guide intermediate 3D representations generated from a single image and input conditions (e.g., applied force and torque), producing high-quality, physically plausible video generation. By utilizing continuum mechanics-based simulations as a prior knowledge, our approach addresses the limitations of traditional data-driven generative models and result in more consistent physically plausible motions. Our framework begins by reconstructing a feed-forward 3D Gaussian from a single image through geometry optimization. This representation is then time-stepped using a differentiable Material Point Method (MPM) with continuum mechanics-based elastoplasticity models, which provides a strong foundation for realistic dynamics, albeit at a coarse level of detail. To enhance the geometry, appearance and ensure spatiotemporal consistency, we refine the initial simulation using a text-to-image (T2I) diffusion model with cross-frame attention, resulting in a physically plausible video that retains intricate details comparable to the input image. We conduct comprehensive qualitative and quantitative evaluations to validate the efficacy of our method. Our project page is available at: https://supertan0204.github.io/physmotion_website/.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 9 Pith papers
-
SOPHY: Learning to Generate Simulation-Ready Objects with Physical Materials
A diffusion-based generative model jointly predicts shape, texture, and physics material parameters for 3D objects, using a new VLM-and-expert annotated dataset of 3,004 objects.
-
Learning Explicit Physical Parameter Control and Benchmarking for Video Generation
Explicit instance-level physical parameter conditioning with routing attention improves physical-law consistency in image-to-video generation, as measured on the authors' new simulator-based benchmark.
-
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
A two-stage framework that predicts per-part material properties from a single image and uses editable physics simulation to guide video generation.
-
Learning an Implicit Physics Model for Image-based Fluid Simulation
Using a simplified physics loss and 3D Gaussians, a neural network animates a single fluid image into videos with novel views, beating earlier methods on quality and motion accuracy.
-
Generative Physical AI in Vision: A Survey
A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.
-
Distilling Physical Priors into Streaming World Models
PhyS adds physics-aware video data, teacher distillation, and windowed reward routing to make streaming world models generate more physically plausible long rollouts.
-
RoboScape: Physics-informed Embodied World Model
RoboScape jointly learns RGB video, depth, and keypoint-token consistency in one autoregressive world model, improving video quality, geometry, action control, synthetic-data policy training, and policy evaluation for...
-
PhysAnimator: Physics-Guided Generative Cartoon Animation
PhysAnimator combines 2D deformable-body physics simulation with a sketch-guided video diffusion model to animate static anime illustrations with controllable, physically plausible motion.
-
Grounding Creativity in Physics: A Brief Survey of Physical Priors in AIGC
A survey that organizes physics-aware 3D and 4D generation methods into a taxonomy and compares several on a synthetic benchmark.
Reference graph
Works this paper leans on
-
[50]
Instant3d: Fast text-to-3d with sparse- view generation and large reconstruction model. arXiv:2311.06214 (2023). 2
arXiv 2023
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Ale- man, Diogo Almeida, Janko Altenschmidt, Sam Alt- man, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv:2303.08774 (2023). 5, 7
arXiv 2023
-
[2]
Tomer Amit, Tal Shaharbany, Eliya Nachmani, and Lior Wolf. 2021. Segdiff: Image segmentation with diffusion probabilistic models. arXiv:2112.00390 (2021). 2
arXiv 2021
-
[3]
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, et al
-
[4]
Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Michal Yarom, Yonatan Bitton, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang, and Aditya Grover
-
[5]
Victor Blomqvist. 2023. Pymunk. https:// pymunk.org. 2
2023
-
[6]
G. Bradski. 2000. The OpenCV Library. Dr. Dobb’s Journal of Software Tools(2000). 15
2000
-
[7]
Junhao Cai, Yuji Yang, Weihao Yuan, Yisheng He, Zilong Dong, Liefeng Bo, Hui Cheng, and Qifeng Chen. 2024. Gaussian-Informed Continuum for Physical Property Identification and Simulation. arXiv:2406.14927 (2024). 2
arXiv 2024
Show all 132 references
-
[8]
Duygu Ceylan, Chun-Hao Huang, and Niloy J. Mi- tra. 2023. Pix2Video: Video Editing using Image Diffusion. In International Conference on Computer Vision (ICCV). 3
2023
-
[9]
Wenhao Chai, Xun Guo, Gaoang Wang, and Yan Lu
-
[10]
Pascal Chang, Jingwei Tang, Markus Gross, and Vinicius C Azevedo. 2024. How I Warped Your Noise: a Temporally-Correlated Noise Prior for Dif- fusion Models. In The Twelfth International Confer- ence on Learning Representations. 3
2024
-
[11]
David Charatan, Sizhe Lester Li, Andrea Tagliasac- chi, and Vincent Sitzmann. 2024. pixelsplat: 3d gaussian splats from image pairs for scalable gener- alizable 3d reconstruction. In Computer Vision and Pattern Recognition (CVPR). 19457–19467. 2
2024
-
[12]
Honghua Chen, Chen Change Loy, and Xingang Pan
-
[13]
Haoxin Chen, Yong Zhang, Xiaodong Cun, Meng- han Xia, Xintao Wang, Chao Weng, and Ying Shan
-
[14]
Weifeng Chen, Yatai Ji, Jie Wu, Hefeng Wu, Pan Xie, Jiashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin
-
[15]
Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bo- han Zhuang, Marc Pollefeys, Andreas Geiger, Tat- 9 Jen Cham, and Jianfei Cai. 2024. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. arXiv:2403.14627 (2024). 2
2024 arXiv
-
[16]
In Computer Vi- sion and Pattern Recognition (CVPR)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior. In Computer Vi- sion and Pattern Recognition (CVPR). 5344–5353. 2
-
[17]
Nathaniel Cohen, Vladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas, and Tomer Michaeli
-
[18]
InComputer Vision and Pattern Recognition (CVPR)
Videocrafter2: Overcoming data limitations for high-quality video diffusion models. InComputer Vision and Pattern Recognition (CVPR). 7310–7320. 2
-
[19]
Prafulla Dhariwal and Alexander Nichol. 2021. Dif- fusion models beat gans on image synthesis. Ad- vances in neural information processing systems 34 (2021), 8780–8794. 2
2021
-
[20]
arXiv:2305.13840 (2023)
Control-a-video: Controllable text-to-video generation with diffusion models. arXiv:2305.13840 (2023). 1
2023 arXiv
-
[21]
Theodore F Gast, Craig Schroeder, Alexey Stom- akhin, Chenfanfu Jiang, and Joseph M Teran. 2015. Optimization integrator for large time steps. IEEE transactions on visualization and computer graphics 21, 10 (2015), 1103–1115. 3
2015
-
[22]
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. 2023. Diffusion policy: Visuomotor policy learning via action diffusion. The International Journal of Robotics Research (2023), 02783649241273668. 2
2023
-
[23]
Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. 2023. TokenFlow: Consistent Dif- fusion Features for Consistent Video Editing. arXiv:2307.10373 (2023). 2, 3, 5, 6, 8, 15
2023 arXiv
-
[24]
arXiv:2405.12211 (2024)
Slicedit: Zero-Shot Video Editing With Text- to-Image Diffusion Models Using Spatio-Temporal Slices. arXiv:2405.12211 (2024). 3
2024 arXiv
-
[25]
Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ramanan. 2022. Depth-supervised nerf: Fewer views and faster training for free. InComputer Vision and Pattern Recognition (CVPR). 12882–12891. 2
2022
-
[26]
Sai Sree Harsha, Ambareesh Revanur, Dhwanit Agarwal, and Shradha Agrawal. 2024. GenVideo: One-shot target-image and shape aware video edit- ing using T2I diffusion models. In Computer Vision and Pattern Recognition (CVPR). 7559–7568. 3
2024
-
[27]
Ruiqi Gao, Aleksander Holynski, Philipp Hen- zler, Arthur Brussee, Ricardo Martin-Brualla, Pratul Srinivasan, Jonathan T Barron, and Ben Poole. 2024. Cat3d: Create anything in 3d with multi-view diffu- sion models. arXiv:2405.10314 (2024). 2
2024 arXiv
-
[28]
Jonathan Ho and Tim Salimans. 2022. Classifier- Free Diffusion Guidance. arXiv:2207.12598 [cs.LG] 15
2022 arXiv
-
[29]
Songwei Ge, Taesung Park, Jun-Yan Zhu, and Jia- Bin Huang. 2023. Expressive text-to-image genera- tion with rich text. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 7545–
2023
-
[30]
Tianyu Huang, Yihan Zeng, Hui Li, Wangmeng Zuo, and Rynson WH Lau. 2024. DreamPhysics: Learning Physical Properties of Dynamic 3D Gaus- sians with Video Diffusion Priors.arXiv:2406.01476 (2024). 2, 3
2024 arXiv
-
[31]
Yuchao Gu, Yipin Zhou, Bichen Wu, Licheng Yu, Jia-Wei Liu, Rui Zhao, Jay Zhangjie Wu, David Jun- hao Zhang, Mike Zheng Shou, and Kevin Tang
-
[32]
In Computer Vision and Pattern Recognition (CVPR)
Videoswap: Customized video subject swapping with interactive semantic point correspon- dence. In Computer Vision and Pattern Recognition (CVPR). 7621–7630. 2, 3
-
[33]
Yuwei Guo, Ceyuan Yang, Anyi Rao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2025. Sparsec- trl: Adding sparse controls to text-to-video diffusion models. In European Conference on Computer Vi- sion. Springer, 330–348. 1
2025
-
[34]
Ying Jiang, Chang Yu, Tianyi Xie, Xuan Li, Yu- tao Feng, Huamin Wang, Minchen Li, Henry Lau, Feng Gao, Yin Yang, et al. 2024. Vr-gs: A physical dynamics-aware interactive gaussian splatting sys- tem in virtual reality. InACM SIGGRAPH 2024 Con- ference Papers. 1–1. 2
2024
-
[35]
Yingqing He, Menghan Xia, Haoxin Chen, Xiaodong Cun, Yuan Gong, Jinbo Xing, Yong Zhang, Xintao Wang, Chao Weng, Ying Shan, et al. 2023. Animate- a-story: Storytelling with retrieval-augmented video generation. arXiv:2307.06940 (2023). 1
2023 arXiv
-
[36]
Tsung-Wei Ke, Nikolaos Gkanatsios, and Katerina Fragkiadaki. 2024. 3d diffuser actor: Policy diffusion 10 with 3d scene representations. arXiv:2402.10885 (2024). 2
2024 arXiv
-
[37]
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In International Con- ference on Learning Representations . https:// openreview.net/forum?id=nZeVKeeFYf9 6
2022
-
[38]
Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll´ar, and Ross Girshick. 2023. Segment Anything. arXiv:2304.02643 (2023). 4
2023 arXiv
-
[39]
Ajay Jain, Matthew Tancik, and Pieter Abbeel. 2021. Putting nerf on a diet: Semantically consistent few- shot view synthesis. InProceedings of the IEEE/CVF International Conference on Computer Vision. 5885–
2021
-
[40]
Hyeonho Jeong, Geon Yeong Park, and Jong Chul Ye. 2024. Vmc: Video motion customization using temporal attention adaption for text-to-video diffu- sion models. In Computer Vision and Pattern Recog- nition (CVPR). 9212–9221. 1, 2
2024
-
[41]
Chenfanfu Jiang, Craig Schroeder, Joseph Teran, Alexey Stomakhin, and Andrew Selle. 2016. The material point method for simulating continuum ma- terials. In Acm siggraph 2016 courses. 1–52. 2, 3
2016
-
[42]
Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xin Ning, Jun Zhou, and Lin Gu. 2024. DNGaus- sian: Optimizing Sparse-View 3D Gaussian Radi- ance Fields with Global-Local Depth Normalization. arXiv:2403.06912 (2024). 2, 4, 5, 15
2024 arXiv
-
[43]
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. 2023. Imagic: Text-based real image editing with diffusion models. In Computer Vision and Pattern Recognition (CVPR). 6007–6017. 2
2023
-
[44]
Mingxiao Li, Bo Wan, Marie-Francine Moens, and Tinne Tuytelaars. 2024. Animate Your Mo- tion: Turning Still Images into Dynamic Videos. arXiv:2403.10179 (2024). 2
2024 arXiv
-
[45]
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk¨uhler, and George Drettakis. 2023. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics 42, 4 (July 2023). https://repo-sam.inria.fr/ fungraph/3d-gaussian-splatting/ 2, 3, 4, 5
2023
-
[46]
Xuan Li, Yi-Ling Qiao, Peter Yichen Chen, Kr- ishna Murthy Jatavallabhula, Ming Lin, Chen- fanfu Jiang, and Chuang Gan. 2023. Pac- nerf: Physics augmented continuum neural radiance fields for geometry-agnostic system identification. arXiv:2303.05512 (2023). 2
2023 arXiv
-
[47]
Nicholas Kolkin, Jason Salavon, and Gregory Shakhnarovich. 2019. Style transfer by relaxed opti- mal transport and self-similarity. In Computer Vision and Pattern Recognition (CVPR). 10051–10060. 8
2019
-
[48]
Max Ku, Cong Wei, Weiming Ren, Harry Yang, and Wenhu Chen. 2024. AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks. arXiv:2403.14468 (2024). 2, 3
2024 arXiv
-
[49]
Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, and Sai Bi
-
[51]
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl V ondrick
-
[52]
Minchen Li, Chenfanfu Jiang, and Zhaofeng Luo
-
[53]
https:// phys-sim-book.github.io/ 3
Physics-Based Simulation . https:// phys-sim-book.github.io/ 3
-
[54]
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song- Hai Zhang, Marc Habermann, Christian Theobalt, et al. 2024. Wonder3d: Single image to 3d using cross-domain diffusion. In Computer Vision and Pat- tern Recognition (CVPR). 9970–9980. 2
2024
-
[55]
Xuan Li, Minchen Li, and Chenfanfu Jiang. 2022. Energetically consistent inelasticity for optimization time integration. ACM Transactions on Graphics (TOG) 41, 4 (2022), 1–16. 3
2022
-
[56]
Miles Macklin, Matthias M ¨uller, and Nuttapong Chentanez. 2016. XPBD: position-based simulation of compliant constrained dynamics. In Proceedings of the 9th International Conference on Motion in Games. 49–54. 2
2016
-
[57]
Zhengqi Li, Richard Tucker, Noah Snavely, and Aleksander Holynski. 2024. Generative image dy- namics. In Computer Vision and Pattern Recognition (CVPR). 24142–24153. 2
2024
-
[58]
Jun Hao Liew, Hanshu Yan, Jianfeng Zhang, Zhong- cong Xu, and Jiashi Feng. 2023. Magicedit: High-fidelity and temporally coherent video editing. arXiv:2308.14749 (2023). 2
2023 arXiv
-
[59]
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. 2023. Magic3d: High-resolution text-to-3d content cre- ation. In Computer Vision and Pattern Recognition (CVPR). 300–309. 2
2023
-
[60]
Jiajing Lin, Zhenzhong Wang, Yongjie Hou, Yuzhou Tang, and Min Jiang. 2024. Phy124: Fast Physics- Driven 4D Content Generation from a Single Image. arXiv:2409.07179 (2024). 2, 3
2024 arXiv
-
[61]
Haomiao Ni, Changhao Shi, Kai Li, Sharon X Huang, and Martin Renqiang Min. 2023. Conditional image-to-video generation with latent flow diffusion models. In Computer Vision and Pattern Recognition (CVPR). 18444–18455. 3
2023
-
[62]
In Proceedings of the IEEE/CVF international conference on computer vision
Zero-1-to-3: Zero-shot one image to 3d ob- ject. In Proceedings of the IEEE/CVF international conference on computer vision. 9298–9309. 2
-
[63]
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, and Shenlong Wang. 2025. PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation. In European Conference on Computer Vision. Springer, 360–378. 2
2025
-
[64]
Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin, and Jiaya Jia. 2024. Video-p2p: Video editing with cross-attention control. In Computer Vision and Pat- tern Recognition (CVPR). 8599–8608. 2
2024
-
[65]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M¨uller, Joe Penna, and Robin Rombach. 2023. Sdxl: Improving latent diffusion models for high-resolution image synthe- sis. arXiv:2307.01952 (2023). 2
2023 arXiv
-
[66]
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022. Repaint: Inpainting using denoising diffusion probabilistic models. In Computer Vision and Pattern Recognition (CVPR). 11461–11471. 2
2022
-
[67]
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. 2023. Fatezero: Fusing attentions for zero- shot text-based video editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 15932–15942. 2, 3
2023
-
[68]
Willi Menapace, Aliaksandr Siarohin, Ivan Sko- rokhodov, Ekaterina Deyneka, Tsai-Shien Chen, Anil Kag, Yuwei Fang, Aleksei Stoliar, Elisa Ricci, Jian Ren, et al. 2024. Snap video: Scaled spa- tiotemporal transformers for text-to-video synthe- sis. In Computer Vision and Patter...
2024
-
[69]
Chenlin Meng, Robin Rombach, Ruiqi Gao, Diederik Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. 2023. On distillation of guided diffu- sion models. In Computer Vision and Pattern Recog- nition (CVPR). 14297–14306. 2
2023
-
[70]
Fanqing Meng, Jiaqi Liao, Xinyu Tan, Wenqi Shao, Quanfeng Lu, Kaipeng Zhang, Yu Cheng, Dianqi Li, Yu Qiao, and Ping Luo. 2024. Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.arXiv:2410.05363 (2024). 2
2024 arXiv
-
[71]
Ben Mildenhall, Pratul P Srinivasan, Matthew Tan- cik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radi- ance fields for view synthesis. Commun. ACM 65, 1 (2021), 99–106. 2, 3
2021
-
[72]
Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian, Dasong Li, Yi Zhang, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, et al. 2024. Motion-i2v: Consistent and controllable image-to-video generation with explicit motion mod- eling. In ACM SIGGRAPH 2024 Conference Pa...
2024
-
[73]
Michael Niemeyer, Jonathan T Barron, Ben Milden- hall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. 2022. Regnerf: Regularizing neural radi- ance fields for view synthesis from sparse inputs. In Computer Vision and Pattern Recognition (CVPR) . 5480–5490. 2
2022
-
[74]
Avinash Paliwal, Wei Ye, Jinhui Xiong, Dmytro Kotovenko, Rakesh Ranjan, Vikas Chandra, and Nima Khademi Kalantari. 2024. CoherentGS: Sparse novel view synthesis with coherent 3D Gaussians. arXiv:2403.19495 2 (2024). 2
2024 arXiv
-
[75]
Taesung Park, Jun-Yan Zhu, Oliver Wang, Jing- wan Lu, Eli Shechtman, Alexei Efros, and Richard Zhang. 2020. Swapping autoencoder for deep image manipulation. Advances in Neural Information Pro- cessing Systems 33 (2020), 7198–7211. 8
2020
-
[76]
Gabriela Ben Melech Stan, Diana Wofk, Scottie Fox, Alex Redden, Will Saxton, Jean Yu, Estelle Aflalo, Shao-Yen Tseng, Fabio Nonato, Matthias Muller, et al. 2023. LDM3D: Latent Diffusion Model for 3D. arXiv:2305.10853 (2023). 6, 15
2023 arXiv
-
[77]
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv:2209.14988 (2022). 2
2022 arXiv
-
[78]
Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. 2024. Splatter Image: Ultra-Fast 12 Single-View 3D Reconstruction. In Computer Vision and Pattern Recognition (CVPR). 2
2024
-
[79]
Jiawei Ren, Liang Pan, Jiaxiang Tang, Chi Zhang, Ang Cao, Gang Zeng, and Ziwei Liu. 2023. Dream- Gaussian4D: Generative 4D Gaussian Splatting. arXiv:2312.17142 (2023). 3
2023 arXiv
-
[80]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Computer Vision and Pattern Recognition (CVPR). 10684–10695. 2
2022
-
[81]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. 2022. High-Resolution Image Synthesis With Latent Diffu- sion Models. In Computer Vision and Pattern Recog- nition (CVPR). 10684–10695. 6
2022
-
[82]
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman
-
[83]
arXiv:2208.12242 (2022)
DreamBooth: Fine Tuning Text-to-image Diffusion Models for Subject-Driven Generation. arXiv:2208.12242 (2022). 5
2022 arXiv
-
[84]
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. 2024. Motionctrl: A unified and flexi- ble motion controller for video generation. In ACM SIGGRAPH 2024 Conference Papers. 1–11. 1, 2
2024
-
[85]
Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. 2023. MV- Dream: Multi-view Diffusion for 3D Generation. arXiv:2308.16512 (2023). 4
2023 arXiv
-
[86]
Yujun Shi, Chuhui Xue, Jun Hao Liew, Jiachun Pan, Hanshu Yan, Wenqing Zhang, Vincent YF Tan, and Song Bai. 2024. Dragdiffusion: Harnessing diffusion models for interactive point-based image editing. In Computer Vision and Pattern Recognition (CVPR) . 8839–8849. 2
2024
-
[87]
Jascha Sohl-Dickstein, Eric Weiss, Niru Mah- eswaranathan, and Surya Ganguli. 2015. Deep un- supervised learning using nonequilibrium thermo- dynamics. In International conference on machine learning. PMLR, 2256–2265. 2
2015
-
[88]
Rundi Wu, Ben Mildenhall, Philipp Henzler, Keun- hong Park, Ruiqi Gao, Daniel Watson, Pratul P Srini- vasan, Dor Verbin, Jonathan T Barron, Ben Poole, et al. 2024. Reconfusion: 3d reconstruction with dif- fusion priors. InComputer Vision and Pattern Recog- nition (CVPR). 21551...
2024
-
[89]
Alexey Stomakhin, Craig Schroeder, Lawrence Chai, Joseph Teran, and Andrew Selle. 2013. A mate- rial point method for snow simulation. ACM Trans. Graph. 32, 4, Article 102 (July 2013), 10 pages. https : / / doi . org / 10 . 1145 / 2461912 . 2461948 3
2013
-
[90]
Weijia Wu, Zhuang Li, Yuchao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, and Di Zhang. 2025. Draganything: Motion control for anything using entity representa- tion. In European Conference on Computer Vision . Springer, 331–348. 2, 7, 8, 16, 17
2025
-
[91]
Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. 2024. LGM: Large Multi-View Gaussian Model for High- Resolution 3D Content Creation. arXiv:2402.05054 (2024). 4
2024 arXiv
-
[92]
Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. 2023. DreamGaussian: Genera- tive Gaussian Splatting for Efficient 3D Content Cre- ation. arXiv:2309.16653 (2023). 2
2023 arXiv
-
[93]
Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. 2023. Plug-and-Play Diffusion Fea- tures for Text-Driven Image-to-Image Translation. In Computer Vision and Pattern Recognition (CVPR) . 1921–1930. 5, 15
2023
-
[94]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st Inter- national Conference on Neural Information Pro- cessing Systems (Long Beach, Californ...
2017
-
[95]
Guangcong Wang, Zhaoxi Chen, Chen Change Loy, and Ziwei Liu. 2023. Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. In Pro- ceedings of the IEEE/CVF International Conference on Computer Vision. 9065–9076. 2
2023
-
[96]
Shuai Yang, Yifan Zhou, Ziwei Liu, and Chen Change Loy. 2023. Rerender a video: Zero-shot text-guided video-to-video translation. In SIGGRAPH Asia 2023 Conference Papers. 1–11. 2, 3
2023
-
[97]
Yujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan, Zhiheng Liu, Yu Liu, Yingya Zhang, Jingren Zhou, and Hongming Shan. 2024. Dreamvideo: Composing your dream videos with customized sub- ject and motion. In Computer Vision and Pattern Recognition (CVPR). 6537–6549. 2
2024
-
[98]
Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen. 2024. latentsplat: Au- toencoding variational gaussians for fast generaliz- able 3d reconstruction. arXiv:2403.16292 (2024). 2
2024 arXiv
-
[99]
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 2024. 4D Gaussian Splat- ting for Real-Time Dynamic Scene Rendering. In Computer Vision and Pattern Recognition (CVPR) . 20310–20320. 3
2024
-
[100]
Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Gong Ming, and Nan Duan. 2023. Dragnuwa: Fine-grained control in video gener- ation by integrating text, image, and trajectory. arXiv:2308.08089 (2023). 1
2023 arXiv
-
[101]
Ronghuan Wu, Wanchao Su, Kede Ma, and Jing Liao. 2024. AniClipart: Clipart Animation with Text-to-Video Priors. arXiv:2404.12347 (2024). 2
2024 arXiv
-
[102]
Hong-Xing Yu, Haoyi Duan, Charles Herrmann, William T Freeman, and Jiajun Wu. 2024. Wonder- World: Interactive 3D Scene Generation from a Sin- gle Image. arXiv:2406.09394 (2024). 2
2024 arXiv
-
[103]
Tianyi Xie, Zeshun Zong, Yuxing Qiu, Xuan Li, Yutao Feng, Yin Yang, and Chenfanfu Jiang. 2023. PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics. arXiv:2311.12198 (2023). 2, 3, 4, 8
2023 arXiv
-
[104]
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, and Tien-Tsin Wong
-
[105]
Shiwei Zhang, Jiayu Wang, Yingya Zhang, Kang Zhao, Hangjie Yuan, Zhiwu Qin, Xiang Wang, Deli Zhao, and Jingren Zhou. 2023. I2vgen-xl: High- quality image-to-video synthesis via cascaded diffu- sion models. arXiv:2311.04145 (2023). 2, 7, 8, 17
2023 arXiv
-
[106]
Haolin Xiong, Sairisheek Muttukuru, Rishi Upad- hyay, Pradyumna Chari, and Achuta Kadambi. 2023. Sparsegs: Real-time 360 {\deg} sparse view syn- thesis using gaussian splatting. arXiv:2312.00206 (2023). 2
2023 arXiv
-
[107]
Jiawei Yang, Marco Pavone, and Yue Wang. 2023. Freenerf: Improving few-shot neural rendering with free frequency regularization. In Computer Vision and Pattern Recognition (CVPR). 8254–8263. 2
2023
-
[108]
Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xi- aodong Chen, and Jing Liao. 2024. Direct-a-video: Customized video generation with user-directed camera movement and object motion. In ACM SIGGRAPH 2024 Conference Papers. 1–12. 2
2024
-
[109]
Rui Zhao, Yuchao Gu, Jay Zhangjie Wu, David Jun- hao Zhang, Jiawei Liu, Weijia Wu, Jussi Keppo, and Mike Zheng Shou. 2023. Motiondirector: Mo- tion customization of text-to-video diffusion models. arXiv:2310.08465 (2023). 2
2023 arXiv
-
[110]
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al
-
[111]
arXiv:2408.06072 (2024)
Cogvideox: Text-to-video diffusion mod- els with an expert transformer. arXiv:2408.06072 (2024). 2, 7, 8, 16, 17
2024 arXiv
-
[112]
Danah Yatim, Rafail Fridman, Omer Bar-Tal, Yoni Kasten, and Tali Dekel. 2024. Space-time diffusion features for zero-shot text-driven motion transfer. In 13 Computer Vision and Pattern Recognition (CVPR) . 8466–8476. 3
2024
-
[113]
Jiraphon Yenphraphai, Xichen Pan, Sainan Liu, Daniele Panozzo, and Saining Xie. 2024. Image Sculpting: Precise Object Editing with 3D Geome- try Control. arXiv:2401.01702 (2024). 2, 5
2024 arXiv
-
[115]
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. 2021. pixelnerf: Neural radiance fields from one or few images. InComputer Vision and Pat- tern Recognition (CVPR). 4578–4587. 2
2021
-
[117]
Yanjie Ze, Gu Zhang, Kangning Zhang, Chenyuan Hu, Muhan Wang, and Huazhe Xu. 2024. 3d diffu- sion policy: Generalizable visuomotor policy learn- ing via simple 3d representations. In ICRA 2024 Workshop on 3D Visual Representations for Robot Manipulation. 2
2024
-
[118]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala
-
[119]
Adding Conditional Control to Text-to-Image Diffusion Models. 2, 5, 6
-
[121]
Tianyuan Zhang, Hong-Xing Yu, Rundi Wu, Bran- don Y Feng, Changxi Zheng, Noah Snavely, Jiajun Wu, and William T Freeman. 2025. Physdreamer: Physics-based interaction with 3d objects via video generation. In European Conference on Computer Vision. Springer, 388–406. 2, 3, 8
2025
-
[122]
Yuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang, Junqi Cheng, Yuefeng Zhu, and Fangyuan Zou
-
[123]
arXiv:2406.19680 (2024)
Mimicmotion: High-quality human motion video generation with confidence-aware pose guid- ance. arXiv:2406.19680 (2024). 1
2024 arXiv
-
[124]
Zhenghao Zhang, Junchao Liao, Menghao Li, Long Qin, and Weizhi Wang. 2024. Tora: Trajectory- oriented Diffusion Transformer for Video Genera- tion. arXiv:2407.21705 (2024). 2
2024 arXiv
-
[126]
Licheng Zhong, Hong-Xing Yu, Jiajun Wu, and Yun- zhu Li. 2025. Reconstruction and simulation of elas- tic objects with spring-mass 3D Gaussians. In Eu- ropean Conference on Computer Vision . Springer, 407–423. 2
2025
-
[127]
Zehao Zhu, Zhiwen Fan, Yifan Jiang, and Zhangyang Wang. 2025. Fsgs: Real-time few-shot view synthe- sis using gaussian splatting. In European Conference on Computer Vision. Springer, 145–163. 2
2025
-
[128]
Zeshun Zong, Chenfanfu Jiang, and Xuchen Han
-
[129]
arXiv:2403.13783 (2024)
A Convex Formulation of Frictional Con- tact for the Material Point Method and Rigid Bodies. arXiv:2403.13783 (2024). 3
2024 arXiv
-
[130]
Zeshun Zong, Xuan Li, Minchen Li, Maurizio M Chiaramonte, Wojciech Matusik, Eitan Grinspun, Kevin Carlberg, Chenfanfu Jiang, and Peter Yichen Chen. 2023. Neural stress fields for reduced-order elastoplasticity and fracture. In SIGGRAPH Asia 2023 Conference Papers. 1–11. 3, 5 1...
2023
-
[132]
A basketball is bouncing up and down on the court
We apply deterministic DDIM+ sampling with 50 steps. We notice that in general, enhanced video achieves bet- ter temporal consistency with higher number of key-frames sampled. For most of out experiments, we randomly choose key-frames every 5 frames, and we set the guidance sc...
-
[500]
We do not apply the soft-depth supervision as indi- cated in [42] since we do not observe significant change of quality in reconstruction output under our settings. A.2.2 Generative Video Enhancement We set the deterministic DDIM+ inversion total steps as 1000, and we set the ...
-
[2022]
arXiv:2211.01324 (2022)
ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv:2211.01324 (2022). 2
2022 arXiv
-
[2023]
In International Conference on Computer Vision (ICCV)
Stablevideo: Text-driven consistency-aware diffusion video editing. In International Conference on Computer Vision (ICCV). 23040–23050. 2
-
[2024]
arXiv:2406.03520 (2024)
VideoPhy: Evaluating Physical Commonsense for Video Generation. arXiv:2406.03520 (2024). 2, 7, 8, 15
2024 arXiv
-
[2025]
In European Con- ference on Computer Vision
Dynamicrafter: Animating open-domain im- ages with video diffusion priors. In European Con- ference on Computer Vision. Springer, 399–417. 2, 7, 8, 16, 17
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.