REVIEW 4 major objections 5 minor 2 cited by
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that large video generation models can act as implicit physics simulators, predicting full videos of interactive object dynamics from one image and a control signal that specifies only the driving motion.
desk verdict Solid controllable-video paper whose implicit-physics claim is ahead of the evidence, but the gap is fixable with identity-aware counterfactual checks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a temporal ControlNet-style conditioning branch attached to a frozen Stable Video Diffusion (SVD) U-Net. The control signal is a per-frame, pixel-aligned map, usually a sequence of binary hand masks, encoded by a small CNN, processed by the branch's interleaved convolution, spatial, and temporal blocks, and injected into the frozen decoder through zero-initialized skip connections. This design lets InterDyn interpret noisy or coarse control sequences, such as masks from SAM2 with motion blur, while keeping the dynamics prior intact. The key design choice is that only the driving entity is conditioned; uncontrolled objects' trajectories are entirely generated, which is what turns video generation into a physics probe.
What would settle it
A quantitative failure test would be: on CLEVRER-style scenes with annotated collision events, measure whether the generated post-collision trajectories of uncontrolled objects match the momentum-conservation predictions, and whether counterfactual futures diverge at the moment of the controlled intervention. A model expressing dataset statistics rather than physical understanding should fail under varied masses, shapes, and collision angles.
Extended reading notes
Core claim
InterDyn's central claim is that a large video diffusion model pre-trained on web-scale video already possesses a workable implicit model of interactive dynamics, and that this capacity can be unlocked by conditioning the generation process on the motion of a driving entity. The paper extends Stable Video Diffusion with a trainable control branch in the style of ControlNet, fed by a pixel-wise control signal (in most experiments a sequence of binary hand masks), while the main U-Net stays frozen to preserve the learned dynamics prior. The model conditions only the driving agent; all other objects are unconditioned, so the video generator must produce their consequential motion on its own. The CLEVRER experiments show generated force propagation across uncontrolled objects and counterfactually different futures for the same input frame under different control signals, while the Something-Something-v2 experiments show realistic hand-object interactions such as pouring, squeezing, dropping, and stacking. The authors' conclusion is that the video generator acts as both neural renderer and implicit physics engine, with no 3D reconstruction and no explicit simulation anywhere in the loop.
Load-bearing premise
The load-bearing premise is that the generated motion reflects learned physical understanding rather than memorized action-class patterns, since the model is fine-tuned and evaluated on the same hand-action video collection and its physics probes are qualitative.
Editorial extensions
If this is right
- Because the control signal is a generic pixel-aligned map, the same framework can extend beyond hand masks to skeletons, meshes, or depth maps, letting a user steer scene dynamics without retraining the backbone.
- Continuous, post-contact dynamics, such as objects rolling, water levels rising, and deformable objects compressing, become generated as video rather than as discrete future states, changing how interactive prediction can be evaluated and used.
- The frozen-backbone design preserves the dynamics prior of the base video model, so fine-tuning on a new interaction dataset mostly teaches the control branch, reducing the data needed to adapt to new scenarios.
- Counterfactual generation on a single input image offers a way to sample diverse physically plausible futures, which could serve as training data or test cases for reasoning about interventions.
Reading between the lines
- Beyond the paper: the force-propagation and counterfactual results suggest the model could be used as a proposal generator for physical prediction tasks, but those results should be stress-tested with out-of-distribution objects, lighting, and camera angles to separate genuine physics from dataset priors.
- Beyond the paper: because InterDyn generates entire videos, its output could be used to bootstrap training of video-based world models or manipulation policies, treating generated trajectories as cheap synthetic experience, an application the paper does not pursue.
- Beyond the paper: a natural testable extension is to quantify, rather than visualize, the implicit physics by measuring whether generated trajectories conserve momentum or energy; the current evidence is qualitative, and a trajectory-level metric would settle the question.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes InterDyn, a framework for controllable interactive video generation. It extends Stable Video Diffusion (SVD) with a trainable ControlNet-like branch that takes an input image and a control signal (e.g., a sequence of binary hand masks) and generates a video of consequent object dynamics. The authors fine-tune the model on CLEVRER and Something-Something-v2 (SSV2), compare against CosHand, Seer, and DynamiCrafter, and argue that large video diffusion models can act as implicit physics simulators by demonstrating force propagation and counterfactual futures on CLEVRER and realistic human-object interactions on SSV2. The paper reports quantitative improvements in image quality, FVD, and a Motion Fidelity metric.
Significance. If the central claim holds, this work is a step toward using video diffusion models as generalizable dynamics engines without explicit physics simulation or 3D reconstruction. The architecture is simple and elegant: the SVD prior is frozen, the control branch is lightweight, and the method does not require fitted physical parameters or an explicit simulator. The paper also provides useful ablations on control signal type and robustness to noisy masks. However, the significance is currently limited by the evaluation: the only quantitative dynamics-related metric is not identity-aware, and the physics probes on CLEVRER are purely qualitative. With stronger evaluation, the contribution could be valuable to the video generation and intuitive physics communities.
major comments (4)
- [Section 4.2] The Motion Fidelity metric defined in Eqs. (1)-(2) does not measure physical interaction fidelity. The bidirectional maximum matching in Eq. (1) allows a generated tracklet to be matched to a ground-truth tracklet of any object rather than the same physical object, and the correlation in Eq. (2) is a cosine similarity of per-frame 2D displacements, which is invariant to speed and absolute position. A video with a correctly moving hand and uncontrolled objects drifting in a similar broad direction as the ground-truth object can therefore score highly without any causal collision or force propagation. Since this is the only quantitative dynamics-related metric in the paper, it cannot support the implicit-physics claim. Please adopt an identity-aware matching (e.g., by initial mask overlap) and report absolute trajectory errors and confidence intervals.
- [Section 4.3] The CLEVRER experiments are purely qualitative. There is no quantitative evaluation of collision outcomes, no identity-aware tracking of uncontrolled objects, no error bars, and no baseline comparing against a no-interaction condition or a model without the frozen SVD prior. The single counterfactual example (Fig. 4b) is illustrative but does not establish distribution-level counterfactual consistency. Because these probes are the core evidence for 'force propagation' and 'counterfactual dynamics,' the paper should include quantitative metrics (e.g., object-wise trajectory divergence between with-control and without-control generations, collision accuracy) and a control condition with no driving signal.
- [Section 4.4, Table 1] The quantitative comparisons do not report error bars, confidence intervals, or repeated-seed statistics, and the claimed improvements are relative to the best baseline per metric; additionally, DynamiCrafter was not trained on SSV2, making cross-method comparison less direct. The class-wise breakdown in Table S3 shows that Motion Fidelity is lowest precisely for complex interactive classes such as spinning, poking, and burying, which is consistent with the weaker interpretation that the model reproduces per-class kinematic priors rather than causal physics. Please report variance across seeds and discuss the class-wise pattern in the main text.
- [Abstract and Section 4.4] The claim that InterDyn 'generalizes to unseen objects' is not systematically evaluated. All quantitative results are on the SSV2 validation split, which comes from the same action-class distribution used for fine-tuning; the few zero-shot examples are qualitative. Please add a held-out action-class or held-out object-category evaluation, or explicitly scope the generalization claim in the abstract.
minor comments (5)
- [Figure 4 caption] The caption contains a typo: '/searcZoom in for details.' should read 'Zoom in for details.'
- [Table S4] The reported KVD values for InterDyn are negative (-0.131 and -0.450); since KVD is an unbiased kernel distance estimate that can be negative in small samples, please clarify the estimator or provide standard errors.
- [Introduction] The introduction calls CosHand 'the previous SOTA' without noting that CosHand solves a different task (single-frame state transition); clarify the task relationship to avoid overstating the comparison.
- [Section 3] The control signal is introduced as c ∈ R^{N×H×W×3} but the experiments use binary masks; please clarify how the three channels are used for mask inputs.
- [Section 3] The paper would benefit from a brief discussion of the relationship between the 'motion ID' conditioning in SVD and the learned control branch; currently the choice of motion ID 40 is only justified by alignment with the frozen prior.
Circularity Check
No significant circularity found: InterDyn's contributions are empirically evaluated against external baselines, and the self-citations in related work are not load-bearing.
full rationale
InterDyn's derivation chain is not circular. The method is SVD plus a ControlNet-style branch (Sec. 3) with frozen SVD weights, fine-tuned on CLEVRER and SSV2; the claimed dynamics capability is then evaluated by comparing generated videos against ground truth using standard image/video metrics and an adapted Motion Fidelity metric (Eqs. 1-2). No parameter is fitted to the test quantity and then reported as a prediction; the motion fidelity score is computed on independently generated samples, not optimized. The self-citations in Related Work (e.g., ARCTIC, HOLD, GRAB, GRIP, MANO) are contextual references to prior hand-object datasets and models, and they do not carry the load of the central claim. The CLEVRER analysis (Sec. 4.3) is qualitative and therefore weak evidence for implicit physics, and the bidirectional matching in Motion Fidelity may not isolate causal force propagation, but these are evaluation-validity limitations, not circularity. Fine-tuning and evaluating on SSV2 uses separate train/validation splits and external baselines, so same-distribution training does not reduce the result to its inputs. Appendix C candidly documents classes where the model fails, which is inconsistent with a forced-by-construction outcome. Overall circularity: none.
Assumptions & free parameters
free parameters (3)
- motion ID =
40
- video subsampling rate =
7 FPS
- denoising steps =
50
assumptions (4)
- domain assumption SVD has already learned interactive dynamics from large-scale video data
- domain assumption Binary hand masks are a sufficient control signal to specify the driving motion
- domain assumption The SSV2/CLEVRER ground-truth videos provide a correct target for interactive dynamics
- domain assumption Motion Fidelity (Eq. 1-2) measures physical plausibility of generated dynamics
Cite this review
Pith. "Pith review of InterDyn: Controllable Interactive Dynamics with Video Diffusion Models." pith.science (2026). https://pith.science/paper/ZTYOCSQM
@misc{pith2026241211785,
author = {Pith},
title = {Pith review of: InterDyn: Controllable Interactive Dynamics with Video Diffusion Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZTYOCSQM}},
note = {Machine review of arXiv:2412.11785}
}
read the original abstract
Predicting the dynamics of interacting objects is essential for both humans and intelligent systems. However, existing approaches are limited to simplified, toy settings and lack generalizability to complex, real-world environments. Recent advances in generative models have enabled the prediction of state transitions based on interventions, but focus on generating a single future state which neglects the continuous dynamics resulting from the interaction. To address this gap, we propose InterDyn, a novel framework that generates videos of interactive dynamics given an initial frame and a control signal encoding the motion of a driving object or actor. Our key insight is that large video generation models can act as both neural renderers and implicit physics ``simulators'', having learned interactive dynamics from large-scale video data. To effectively harness this capability, we introduce an interactive control mechanism that conditions the video generation process on the motion of the driving entity. Qualitative results demonstrate that InterDyn generates plausible, temporally consistent videos of complex object interactions while generalizing to unseen objects. Quantitative evaluations show that InterDyn outperforms baselines that focus on static state transitions. This work highlights the potential of leveraging video generative models as implicit physics engines. Project page: https://interdyn.is.tue.mpg.de/
Figures
Figures from the paper (4 more)
Forward citations
Cited by 2 Pith papers
-
Precise Action-to-Video Generation Through Visual Action Prompts
Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.
-
Generative Physical AI in Vision: A Survey
A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.
Reference graph
Works this paper leans on
-
[1]
Lumiere: A space-time dif- fusion model for video generation
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Her- rmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, et al. Lumiere: A space-time dif- fusion model for video generation. In International Con- ference on Computer Graphics and Interactive Techniques (SIGGRAPH), pages 1–11, 2024. 2, 3
2024
-
[2]
CoPhy: Counterfactual learning of physical dynamics
Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, and Christian Wolf. CoPhy: Counterfactual learning of physical dynamics. In International Conference on Learn- ing Representations (ICLR), 2020. 2, 3
2020
-
[3]
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al. Interaction networks for learning about objects, relations and physics. In Conference on Neu- ral Information Processing Systems (NeurIPS), 2016. 3
2016
-
[4]
Sutherland, Michael Arbel, and Arthur Gretton
Mikołaj Bi ´nkowski, Danica J. Sutherland, Michael Arbel, and Arthur Gretton. Demystifying MMD GANs. In Inter- national Conference on Learning Representations (ICLR) ,
-
[5]
Stable Video Diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable Video Diffusion: Scaling latent video diffusion models to large datasets. arXiv:2311.15127, 2023. 2, 3, 4, 5
arXiv 2023
-
[6]
Align your latents: High-resolution video synthe- sis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthe- sis with latent diffusion models. In Computer Vision and Pattern Recognition (CVPR), pages 22563–22575, 2023. 3
2023
-
[7]
Tim Brooks, Aleksander Holynski, and Alexei A. Efros. InstructPix2Pix: Learning to follow image editing instruc- tions. In Computer Vision and Pattern Recognition (CVPR), pages 18392–18402, 2023. 14
2023
-
[8]
Generative rendering: Controllable 4D-guided video generation with 2D diffusion models
Shengqu Cai, Duygu Ceylan, Matheus Gadelha, Chun- Hao Paul Huang, Tuanfeng Yang Wang, and Gordon Wet- zstein. Generative rendering: Controllable 4D-guided video generation with 2D diffusion models. In Computer Vision and Pattern Recognition (CVPR), pages 7611–7620, 2024. 3
2024
Show all 106 references
-
[9]
Realtime multi-person 2D pose estimation using part affin- ity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Realtime multi-person 2D pose estimation using part affin- ity fields. In Computer Vision and Pattern Recognition (CVPR), pages 7291–7299, 2017. 14
2017
-
[10]
Text2HOI: Text-guided 3D motion generation for hand-object interaction
Junuk Cha, Jihyeon Kim, Jae Shin Yoon, and Seungryul Baek. Text2HOI: Text-guided 3D motion generation for hand-object interaction. In Computer Vision and Pattern Recognition (CVPR), pages 1577–1585, 2024. 3
2024
-
[11]
Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, et al
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, et al. DexYCB: A benchmark for capturing hand grasping of objects. In Computer Vision and Pattern Recognition (CVPR) , pages 9044–905...
2021
-
[12]
VideoCrafter1: Open diffusion models for high-quality video generation
Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, et al. VideoCrafter1: Open diffusion models for high-quality video generation. arXiv:2310.19512, 2023. 3
-
[13]
Control-A- Video: Controllable text-to-video generation with diffusion models
Weifeng Chen, Yatai Ji, Jie Wu, Hefeng Wu, Pan Xie, Ji- ashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin. Control-A- Video: Controllable text-to-video generation with diffusion models. arXiv:2305.13840, 2023. 3
2023 arXiv
-
[14]
GanHand: Predicting human grasp affordances in multi-object scenes
Enric Corona, Albert Pumarola, Guillem Alenya, Francesc Moreno-Noguer, and Gr´egory Rogez. GanHand: Predicting human grasp affordances in multi-object scenes. In Com- puter Vision and Pattern Recognition (CVPR), pages 5031– 5041, 2020. 3
2020
-
[15]
CG-HOI: Contact-guided 3D human-object interaction generation
Christian Diller and Angela Dai. CG-HOI: Contact-guided 3D human-object interaction generation. In Computer Vi- sion and Pattern Recognition (CVPR), pages 19888–19901,
-
[16]
Scaling recti- fied flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling recti- fied flow transformers for high-resolution image synthesis. In International Conference on Machine Learning (ICML),
-
[17]
Black, and Ot- mar Hilliges
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Ot- mar Hilliges. ARCTIC: A dataset for dexterous bimanual hand-object manipulation. In Computer Vision and Pattern Recognition (CVPR), pages 12943–12954, 2023. 3
2023
-
[18]
Black, and Otmar Hilliges
Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J. Black, and Otmar Hilliges. HOLD: Category-agnostic 3D reconstruction of interacting hands and objects from video. In Computer Vision and Pattern Recognition (CVPR) , pages 494–504,
-
[19]
Phys- ically grounded vision-language models for robotic manip- ulation
Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao, Jiajun Wu, Brian Ichter, Anirudha Majumdar, and Dorsa Sadigh. Phys- ically grounded vision-language models for robotic manip- ulation. In International Conference on Robotics and Au- tomation (ICRA), pages 12462–12469, 2024. 3
2024
-
[20]
Emu Video: Factor- izing text-to-video generation by explicit image condition- ing
Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Du- val, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. Emu Video: Factor- izing text-to-video generation by explicit image condition- ing. In European Conference on Computer Vision (ECCV),
-
[21]
The something something video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michal- ski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The something something video database for learning and evaluating visual common sense. In In...
2017
-
[22]
Fuchs, Ingmar Posner, and Andrea Vedaldi
Oliver Groth, Fabian B. Fuchs, Ingmar Posner, and Andrea Vedaldi. ShapeStacks: Learning vision-based physical in- tuition for generalised object stacking. In European Con- ference on Computer Vision (ECCV), pages 702–717, 2018. 3
2018
-
[23]
Seer: Language instructed video prediction with latent diffusion models
Xianfan Gu, Chuan Wen, Weirui Ye, Jiaming Song, and Yang Gao. Seer: Language instructed video prediction with latent diffusion models. In International Conference on Learning Representations (ICLR), 2024. 2, 3, 8, 17
2024
-
[24]
Disentangling physi- cal dynamics from unknown factors for unsupervised video prediction
Vincent Le Guen and Nicolas Thome. Disentangling physi- cal dynamics from unknown factors for unsupervised video prediction. In Computer Vision and Pattern Recognition (CVPR), pages 11471–11481, 2020. 2, 3
2020
-
[25]
I2V-Adapter: A general image-to- video adapter for diffusion models
Xun Guo, Mingwu Zheng, Liang Hou, Yuan Gao, Yufan Deng, Pengfei Wan, Di Zhang, Yufan Liu, Weiming Hu, Zhengjun Zha, et al. I2V-Adapter: A general image-to- video adapter for diffusion models. In International Con- ference on Computer Graphics and Interactive Techniques (SIGGRA...
2024
-
[26]
AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv:2307.04725, 2023. 3
2023 arXiv
-
[27]
Black, Ivan Laptev, and Cordelia Schmid
Yana Hasson, G ¨ul Varol, Dimitrios Tzionas, Igor Kale- vatykh, Michael J. Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manip- ulated objects. In Computer Vision and Pattern Recognition (CVPR), pages 11807–11816, 2019. 3
2019
-
[28]
Towards unconstrained joint hand-object reconstruction from RGB videos
Yana Hasson, G ¨ul Varol, Cordelia Schmid, and Ivan Laptev. Towards unconstrained joint hand-object reconstruction from RGB videos. In International Conference on 3D Vi- sion (3DV), pages 659–668, 2021. 3
2021
-
[29]
CameraC- trl: Enabling camera control for text-to-video generation
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. CameraC- trl: Enabling camera control for text-to-video generation. arXiv:2404.02101, 2024. 3
2024 arXiv
-
[30]
GANs trained by a two time-scale update rule converge to a local nash equi- librium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule converge to a local nash equi- librium. In Conference on Neural Information Processing Systems (NeurIPS), pages 6629–6640, 2017. 5
2017
-
[31]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv:2207.12598, 2022. 5
2022 arXiv
-
[32]
Kingma, Ben Poole, Mohammad Norouzi, David J
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, et al. Ima- gen Video: High definition video generation with diffusion models. arXiv:2210.02303, 2022. 3
-
[33]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video dif- fusion models. In Conference on Neural Information Pro- cessing Systems (NeurIPS), pages 8633–8646, 2022. 3
2022
-
[34]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language mod- els. In International Conference on Learning Representa- tions (ICLR), 2022. 3
2022
-
[35]
Animate Anyone: Consistent and controllable image-to-video synthesis for character animation
Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. Animate Anyone: Consistent and controllable image-to-video synthesis for character animation. In Com- puter Vision and Pattern Recognition (CVPR), pages 8153– 8163, 2024. 3
2024
-
[36]
DepthCrafter: Generating consistent long depth sequences for open-world videos
Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xi- aodong Cun, Yong Zhang, Long Quan, and Ying Shan. DepthCrafter: Generating consistent long depth sequences for open-world videos. arXiv:2409.02095, 2024. 4
2024 arXiv
-
[37]
Make It Move: Controllable image-to-video generation with text descriptions
Yaosi Hu, Chong Luo, and Zhenzhong Chen. Make It Move: Controllable image-to-video generation with text descriptions. In Computer Vision and Pattern Recognition (CVPR), pages 18219–18228, 2022. 3
2022
-
[38]
VideoControlNet: A motion- guided video-to-video translation framework by using dif- fusion model with ControlNet
Zhihao Hu and Dong Xu. VideoControlNet: A motion- guided video-to-video translation framework by using dif- fusion model with ControlNet. arXiv:2307.14073, 2023. 3
2023 arXiv
-
[39]
Filtered-CoPhy: Unsupervised learning of counterfactual physics in pixel space
Steeven Janny, Fabien Baradel, Natalia Neverova, Madiha Nadri, Greg Mori, and Christian Wolf. Filtered-CoPhy: Unsupervised learning of counterfactual physics in pixel space. In International Conference on Learning Represen- tations (ICLR), 2022. 2, 3
2022
-
[40]
Hand-object contact consistency reasoning for hu- man grasps generation
Hanwen Jiang, Shaowei Liu, Jiashun Wang, and Xiaolong Wang. Hand-object contact consistency reasoning for hu- man grasps generation. In International Conference on Computer Vision (ICCV), pages 11107–11116, 2021. 3
2021
-
[41]
Co- Tracker3: Simpler and better point tracking by pseudo- labelling real videos
Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- Tracker3: Simpler and better point tracking by pseudo- labelling real videos. InEuropean Conference on Computer Vision (ECCV), pages 18–35, 2024. 5
2024
-
[42]
DreamPose: Fashion image-to-video synthesis via stable diffusion
Johanna Karras, Aleksander Holynski, Ting-Chun Wang, and Ira Kemelmacher-Shlizerman. DreamPose: Fashion image-to-video synthesis via stable diffusion. In Inter- national Conference on Computer Vision (ICCV) , pages 22623–22633, 2023. 3
2023
-
[43]
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. Conference on Neural Information Processing Sys- tems (NeurIPS), 35:26565–26577, 2022. 5
2022
-
[44]
Be- yond the Contact: Discovering comprehensive affordance for 3D objects from pre-trained 2D diffusion models
Hyeonwoo Kim, Sookwan Han, Patrick Kwon, et al. Be- yond the Contact: Discovering comprehensive affordance for 3D objects from pre-trained 2D diffusion models. In European Conference on Computer Vision (ECCV) , pages 400–419, 2024. 3
2024
-
[45]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015. 5
2015
-
[46]
Neural relational inference for interacting systems
Thomas Kipf, Ethan Fetaya, Kuan-Chieh Wang, Max Welling, and Richard Zemel. Neural relational inference for interacting systems. InInternational Conference on Ma- chine Learning (ICML), pages 2693–2702, 2018. 2
2018
-
[47]
Efros, and Krishna Kumar Singh
Sumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu, Jimei Yang, Jingwan Lu, Alexei A. Efros, and Krishna Kumar Singh. Putting people in their place: Affordance-aware hu- man insertion into scenes. In Computer Vision and Pattern Recognition (CVPR), pages 17089–17099, 2023. 3
2023
-
[48]
Learning phys- ical intuition of block towers by example
Adam Lerer, Sam Gross, and Rob Fergus. Learning phys- ical intuition of block towers by example. In International Conference on Machine Learning (ICML), pages 430–438,
-
[49]
Causal discovery in physical sys- tems from videos
Yunzhu Li, Antonio Torralba, Anima Anandkumar, Dieter Fox, and Animesh Garg. Causal discovery in physical sys- tems from videos. Conference on Neural Information Pro- cessing Systems (NeurIPS), 2020. 2
2020
-
[50]
Freeman, Fr ´edo Du- rand, and Edward H
Ce Liu, Antonio Torralba, William T. Freeman, Fr ´edo Du- rand, and Edward H. Adelson. Motion magnification. Transactions on Graphics (TOG), 24(3):519–526, 2005. 5
2005
-
[51]
Mind’s Eye: Grounded language model reasoning through simulation
Ruibo Liu, Jason Wei, Shixiang Shane Gu, Te-Yen Wu, Soroush V osoughi, Claire Cui, Denny Zhou, and Andrew M Dai. Mind’s Eye: Grounded language model reasoning through simulation. In International Conference on Learn- ing Representations (ICLR), 2023. 3
2023
-
[52]
Grounding DINO: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding DINO: Marrying dino with grounded pre-training for open-set object detection. arXiv:2303.05499, 2023. 14
2023 arXiv
-
[53]
PhysGen: Rigid-body physics-grounded image-to-video generation
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, and Shen- long Wang. PhysGen: Rigid-body physics-grounded image-to-video generation. In European Conference on Computer Vision (ECCV), pages 360–378, 2024. 3
2024
-
[54]
Sora: A review on background, technol- ogy, limitations, and opportunities of large vision models
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, et al. Sora: A review on background, technol- ogy, limitations, and opportunities of large vision models. arXiv:2402.17177, 2024. 2, 4
2024 arXiv
-
[55]
Freeman, and Michael Rubinstein
Erika Lu, Forrester Cole, Tali Dekel, Weidi Xie, An- drew Zisserman, David Salesin, William T. Freeman, and Michael Rubinstein. Layered neural rendering for retiming people in video. Transactions on Graphics (TOG), 39(6): 1–14, 2020. 3
2020
-
[56]
Freeman, and Michael Rubinstein
Erika Lu, Forrester Cole, Tali Dekel, Andrew Zisserman, William T. Freeman, and Michael Rubinstein. Omnimatte: Associating objects and their effects in video. In Computer Vision and Pattern Recognition (CVPR), pages 4507–4515,
-
[57]
UGG: Unified generative grasping
Jiaxin Lu, Hao Kang, Haoxiang Li, Bo Liu, Yiding Yang, Qixing Huang, and Gang Hua. UGG: Unified generative grasping. In European Conference on Computer Vision (ECCV), pages 414–433, 2024. 3
2024
-
[58]
Bastiaan Kleijn
Wan-Duo Kurt Ma, John P Lewis, and W. Bastiaan Kleijn. TrailBlazer: Trajectory control for diffusion-based video generation. In International Conference on Computer Graphics and Interactive Techniques (SIGGRAPH) , pages 1–11, 2023. 3
2023
-
[59]
Something-Else: Compositional action recognition with spatial-temporal in- teraction networks
Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang, and Trevor Darrell. Something-Else: Compositional action recognition with spatial-temporal in- teraction networks. In Computer Vision and Pattern Recog- nition (CVPR), pages 1049–1059, 2020. 6
2020
-
[60]
T2I-Adapter: Learn- ing adapters to dig out more controllable ability for text-to- image diffusion models
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2I-Adapter: Learn- ing adapters to dig out more controllable ability for text-to- image diffusion models. In AAAI Conference on Artificial Intelligence, pages 4296–4304, 2024. 3
2024
-
[61]
Otaduy, Dan Casas, and Christian Theobalt
Franziska Mueller, Micah Davis, Florian Bernard, Olek- sandr Sotnychenko, Mickeal Verschoor, Miguel A. Otaduy, Dan Casas, and Christian Theobalt. Real-time pose and shape reconstruction of two interacting hands with a sin- gle depth camera. Transactions on Graphics (TOG), 38(4...
2019
-
[62]
Guibas, and Jimei Yang
Boxiao Pan, Zhan Xu, Chun-Hao Paul Huang, Krishna Ku- mar Singh, Yang Zhou, Leonidas J. Guibas, and Jimei Yang. ActAnywhere: Subject-aware video background genera- tion. In Conference on Neural Information Processing Sys- tems (NeurIPS), 2024. 3
2024
-
[63]
3D whole-body grasp synthesis with directional controllability
Georgios Paschalidis, Romana Wilschut, Dimitrije Anti ´c, Omid Taheri, and Dimitrios Tzionas. 3D whole-body grasp synthesis with directional controllability. In International Conference on 3D Vision (3DV), 2025. 3
2025
-
[64]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International Conference on Machine Learning...
2021
-
[65]
Girshick, Piotr Doll ´ar, and Christoph Feichtenhofer
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chlo´e Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross B. Girshick, Piotr Doll ´ar, and Christoph...
2024 arXiv
-
[66]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Computer Vi- sion and Pattern Recognition (CVPR), pages 10684–10695,
-
[67]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bod- ies together. Transactions on Graphics (TOG), 36(6), 2022. 14
2022
-
[68]
Make-A-Video: Text-to-video genera- tion without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-A-Video: Text-to-video genera- tion without text-video data. In International Conference on Learning Representations (ICLR), 2023. 3
2023
-
[69]
StyleGAN-V: A continuous video generator with the price, image quality and perks of StyleGAN2
Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elho- seiny. StyleGAN-V: A continuous video generator with the price, image quality and perks of StyleGAN2. In Computer Vision and Pattern Recognition (CVPR), pages 3626–3636,
-
[70]
Controlling the world by sleight of hand
Sruthi Sudhakar, Ruoshi Liu, Basile Van Hoorick, Carl V ondrick, and Richard Zemel. Controlling the world by sleight of hand. In European Conference on Computer Vi- sion (ECCV), pages 414–430, 2024. 2, 3, 6, 8, 14, 17
2024
-
[71]
Black, and Dim- itrios Tzionas
Omid Taheri, Nima Ghorbani, Michael J. Black, and Dim- itrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. In European Conference on Computer Vision (ECCV), pages 581–600, 2020. 3
2020
-
[72]
Omid Taheri, Yi Zhou, Dimitrios Tzionas, Yang Zhou, Duygu Ceylan, Soren Pirk, and Michael J. Black. GRIP: Generating interaction poses using latent consistency and spatial cues. In International Conference on 3D Vision (3DV), pages 933–943, 2024. 3
2024
-
[73]
Collaborative learning for hand and object reconstruction with attention-guided graph convolu- tion
Tze Ho Elden Tse, Kwang In Kim, Ales Leonardis, and Hyung Jin Chang. Collaborative learning for hand and object reconstruction with attention-guided graph convolu- tion. In Computer Vision and Pattern Recognition (CVPR), pages 1664–1674, 2022. 3
2022
-
[74]
FVD: A new metric for video generation
Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Rapha¨el Marinier, Marcin Michalski, and Sylvain Gelly. FVD: A new metric for video generation. In Interna- tional Conference on Learning Representations Workshops (ICLRw), 2019. 5
2019
-
[75]
SV3D: Novel multi-view synthesis and 3D generation from a single image using la- tent video diffusion
Vikram V oleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, and Varun Jampani. SV3D: Novel multi-view synthesis and 3D generation from a single image using la- tent video diffusion. In European Conference on Computer...
2025
-
[76]
Modelscope text-to-video technical report
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv:2308.06571, 2023. 3
2023 arXiv
-
[77]
Boximator: Gen- erating rich and controllable motions for video synthesis
Jiawei Wang, Yuchen Zhang, Jiaxin Zou, Yan Zeng, Guo- qiang Wei, Liping Yuan, and Hang Li. Boximator: Gen- erating rich and controllable motions for video synthesis. In International Conference on Machine Learning (ICML),
-
[78]
VideoComposer: Compositional video syn- thesis with motion controllability
Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. VideoComposer: Compositional video syn- thesis with motion controllability. In Conference on Neural Information Processing Systems (NeurIPS), 2024. 3
2024
-
[79]
Bovik, Hamid R
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. Image Quality Assessment: From error visibil- ity to structural similarity. Transactions on Image Process- ing (TIP), 13(4):600–612, 2004. 3, 5
2004
-
[80]
MotionCtrl: A unified and flexible motion controller for video generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. MotionCtrl: A unified and flexible motion controller for video generation. In Transactions on Graphics (TOG) , pages 1–11, 2024. 3
2024
-
[81]
Visual Interaction Networks: Learning a physics simulator from video
Nicholas Watters, Daniel Zoran, Theophane Weber, Peter Battaglia, Razvan Pascanu, and Andrea Tacchetti. Visual Interaction Networks: Learning a physics simulator from video. In Conference on Neural Information Processing Systems (NeurIPS), pages 4539–4547, 2017. 3
2017
-
[82]
Lim, Bill Freeman, and Joshua B
Jiajun Wu, Ilker Yildirim, Joseph J. Lim, Bill Freeman, and Joshua B. Tenenbaum. Galileo: Perceiving physical ob- ject properties by integrating a physics engine with deep learning. In Conference on Neural Information Processing Systems (NeurIPS), pages 127–135, 2015. 3
2015
-
[83]
Lim, Hongyi Zhang, Joshua B
Jiajun Wu, Joseph J. Lim, Hongyi Zhang, Joshua B. Tenen- baum, and William T. Freeman. Physics 101: Learning physical object properties from unlabeled videos. InBritish Machine Vision Conference (BMVC), 2016
2016
-
[84]
Learning to see physics via visual de- animation
Jiajun Wu, Erika Lu, Pushmeet Kohli, Bill Freeman, and Josh Tenenbaum. Learning to see physics via visual de- animation. In Conference on Neural Information Process- ing Systems (NeurIPS), pages 153–164, 2017. 2, 3
2017
-
[85]
Tune-A-Video: One-shot tun- ing of image diffusion models for text-to-video generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-A-Video: One-shot tun- ing of image diffusion models for text-to-video generation. In International Conference on Computer Vision (ICCV)...
2023
-
[86]
THOR: Text to human-object interac- tion diffusion via relation intervention
Qianyang Wu, Ye Shi, Xiaoshui Huang, Jingyi Yu, Lan Xu, and Jingya Wang. THOR: Text to human-object interac- tion diffusion via relation intervention. arXiv:2403.11208,
-
[87]
DragAnything: Motion control for any- thing using entity representation
Weijia Wu, Zhuang Li, Yuchao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, and Di Zhang. DragAnything: Motion control for any- thing using entity representation. In European Conference on Computer Vision (ECCV), pages 331–348, 2025. 3
2025
-
[88]
Template free reconstruction of human- object interaction with procedural interaction generation
Xianghui Xie, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. Template free reconstruction of human- object interaction with procedural interaction generation. In Computer Vision and Pattern Recognition (CVPR) , pages 10003–10015, 2024. 3
2024
-
[89]
DynamiCrafter: Animating open-domain images with video diffusion priors
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang, Ying Shan, and Tien-Tsin Wong. DynamiCrafter: Animating open-domain images with video diffusion priors. In Euro- pean Conference on Computer Vision (ECCV), pages 399– 417, 2025. ...
2025
-
[90]
InterDiff: Generating 3D human-object interactions with physics-informed diffusion
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui. InterDiff: Generating 3D human-object interactions with physics-informed diffusion. In International Con- ference on Computer Vision (ICCV) , pages 14928–14940,
-
[91]
MagicAnimate: Temporally consistent human image animation using diffusion model
Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Han- shu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. MagicAnimate: Temporally consistent human image animation using diffusion model. In Com- puter Vision and Pattern Recognition (CVPR), pages 1481– 1490, 2024. 3
2024
-
[92]
CogVideoX: Text- to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. CogVideoX: Text- to-video diffusion models with an expert transformer. arXiv:2408.06072, 2024. 3
2024 arXiv
-
[93]
Space-time diffusion features for zero-shot text-driven motion transfer
Danah Yatim, Rafail Fridman, Omer Bar-Tal, Yoni Kasten, and Tali Dekel. Space-time diffusion features for zero-shot text-driven motion transfer. InComputer Vision and Pattern Recognition (CVPR), pages 8466–8476, 2024. 5, 8, 14, 17
2024
-
[94]
Diffusion-guided reconstruction of everyday hand-object interaction clips
Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shubham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. In International Conference on Computer Vision (ICCV), pages 19717–19728, 2023. 3
2023
-
[95]
Affordance Diffusion: Synthesizing hand-object in- teractions
Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello, Stan Birchfield, Jiaming Song, Shubham Tulsiani, and Sifei Liu. Affordance Diffusion: Synthesizing hand-object in- teractions. In Computer Vision and Pattern Recognition (CVPR), pages 22479–22489, 2023. 3
2023
-
[96]
G-HOP: Generative hand-object prior for interaction reconstruction and grasp synthesis
Yufei Ye, Abhinav Gupta, Kris Kitani, and Shubham Tul- siani. G-HOP: Generative hand-object prior for interaction reconstruction and grasp synthesis. In Computer Vision and Pattern Recognition (CVPR), pages 1911–1920, 2024. 3
1911
-
[97]
Tenenbaum
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Ji- ajun Wu, Antonio Torralba, and Joshua B. Tenenbaum. CLEVRER: Collision events for video representation and reasoning. In International Conference on Learning Repre- sentations (ICLR), 2020. 2, 4, 6
2020
-
[98]
DragNUW A: Fine-grained control in video generation by integrating text, image, and trajectory
Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Gong Ming, and Nan Duan. DragNUW A: Fine-grained control in video generation by integrating text, image, and trajectory. arXiv:2308.08089, 2023. 3
2023 arXiv
-
[99]
Video probabilistic diffusion models in projected latent space
Sihyun Yu, Kihyuk Sohn, Subin Kim, and Jinwoo Shin. Video probabilistic diffusion models in projected latent space. In Computer Vision and Pattern Recognition (CVPR), pages 18456–18466, 2023. 3
2023
-
[100]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. InIn- ternational Conference on Computer Vision (ICCV), pages 3836–3847, 2023. 3, 4
2023
-
[101]
HOIDiffusion: Generating real- istic 3D hand-object interaction data
Mengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu, Zhuowen Tu, and Xiaolong Wang. HOIDiffusion: Generating real- istic 3D hand-object interaction data. In Computer Vision and Pattern Recognition (CVPR), pages 8521–8531, 2024. 3
2024
-
[102]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Computer Vision and Pattern Recognition (CVPR), pages 586–595, 2018. 5
2018
-
[103]
CAMS: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthe- sis
Juntian Zheng, Qingyuan Zheng, Lixing Fang, Yun Liu, and Li Yi. CAMS: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthe- sis. In Computer Vision and Pattern Recognition (CVPR), pages 585–594, 2023. 3
2023
-
[104]
MagicVideo: Efficient video generation with latent diffusion models.arXiv:2211.11018,
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. MagicVideo: Efficient video generation with latent diffusion models.arXiv:2211.11018,
-
[105]
GEARS: Local geometry-aware hand- object interaction synthesis
Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. GEARS: Local geometry-aware hand- object interaction synthesis. In Computer Vision and Pat- tern Recognition (CVPR), pages 20634–20643, 2024. 3
2024
-
[106]
pre- tending
Tianqiang Zhu, Rina Wu, Xiangbo Lin, and Yi Sun. To- ward human-like grasp: Dexterous grasping via semantic representation of object-hand. In International Conference on Computer Vision (ICCV), pages 15741–15751, 2021. 3 InterDyn: Controllable Interactive Dynamics with Video D...
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.