Pith. sign in

REVIEW 3 major objections 93 references

GIRAF: Towards Generalizable Human Interactions with Articulated Objects

T0 review · 3 major / 0 minor · reviewed 2026-07-10 · grok-4.5

Pith's one-line read A text-conditioned diffusion model generates full-body human motion that approaches, contacts, and actuates articulated objects, generalizing to placements never seen in training.

desk verdict Clean systems paper that finally joins full-body locomotion with two-DoF articulated manipulation under text; contact numbers improve, but the generalization claim outruns the narrow reconstruction eval. read the letter →

arxiv 2607.07880 v1 pith:534QAA3N submitted 2026-07-08 cs.CV

classification cs.CV
keywords human-objectinteractionarticulatedobjectsmotiondiffusionfull-bodysynthesistext-conditionedgenerationcontactmodelinglocomotion-to-manipulation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Prior work either handles simple full-body actions with static objects or restricts itself to hand-only grasping, leaving open the harder problem of coordinated sequences that walk up to an articulated object, make fine contact, and move its parts. This paper claims that one diffusion model can close that gap when three ingredients are combined: an object-centric representation that places contact labels, hand end-effectors, and object surfaces in the same shared space; mixed-domain training that balances free locomotion with interaction; and contact-preserving relocation of training examples to expand spatial diversity. The resulting sequences are physically plausible over long horizons and adapt body reach, grasp, and object articulation to new positions and scales. If the claim holds, synthetic motion of this kind becomes a practical source of training data for robots and virtual agents that must navigate and manipulate the everyday world.

What carries the argument

Dynamic basis point sets (dynamic BPS) canonicalized to the articulated object part: a fixed cloud of points that jointly encode object-surface distances, distances to hand end-effectors, and binary contact labels, rendering contact shape-agnostic and transferable across geometries.

What would settle it

Evaluate the trained model on object placements far outside the 0.1 m augmentation grid and on articulated mechanisms absent from training (for example multi-link doors); if contact distance or penetration then exceeds the baselines reported for the original test set, the generalization claim is falsified.

Watch

Extended reading notes

Core claim

Unifying hand–object contact, hand end-effectors, and object surfaces inside a dynamic object-centric basis-point representation, trained with mixed locomotion–interaction batches and contact-based data relocation, lets a single text-conditioned diffusion model synthesize seamless full-body sequences that approach, grasp, and actuate articulated objects while generalizing to unseen configurations and outperforming prior methods on contact, penetration, pose, and text-alignment metrics.

Load-bearing premise

That relocating contacts onto a coarse 3-D grid and re-solving human pose with inverse kinematics plus joint limits still yields training examples realistic enough for the model to generalize to truly novel placements and scales.

Editorial extensions

If this is right

  • Long-horizon locomotion-to-manipulation can be generated by one model without separate navigation and grasp stages.
  • Scarce paired human–scene data can be expanded by contact-preserving object relocation rather than new motion capture.
  • Text prompts can steer not only the action but also contact strategy (one hand versus both).
  • Synthetic sequences become usable training data for embodied agents that must both navigate and actuate everyday articulated objects.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same shared-point voting scheme for contact could extend to multi-object clutter or continuous multi-DoF articulations beyond the two-DoF hinges used here.
  • Coupling the diffusion prior more tightly with a physics simulator might remove residual foot sliding and joint jitter without a separate noise-optimization stage.
  • Text control over grasp laterality suggests a route to interactive virtual agents whose contact preferences can be specified on the fly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. GIRAF proposes a text-conditioned diffusion model for synthesizing long-horizon full-body human motion that approaches, contacts, and actuates two-part articulated household objects (drawer, microwave, refrigerator, washing machine). The method combines three components: a dynamic object-centric basis-point-set (BPS) representation that jointly encodes object surface distances, hand end-effector distances and binary contact labels; a mixed-domain training strategy that uses FiLM conditioning and an annealing schedule over homogeneous locomotion/interaction batches; and a contact-preserving augmentation that relocates and rescales objects on a discrete grid while re-solving human motion via CCD inverse kinematics. After generation, a two-stage scene-aware DDIM noise optimization reduces contact error, penetration and jitter. Experiments on an augmented ParaHome+Babel corpus (2100 sequences, ~3 h) report consistent gains over retrained LINGO and CHOIS baselines on contact distance/F1, penetration, MPJPE/hand/object error, R-precision and FID, together with qualitative examples of height variation, multi-door closets and left/right/both-hand strategies.

Significance. If the claimed generalization holds, the work would supply a practical, text-driven generator of coordinated locomotion-plus-articulated-manipulation sequences that current HOI and hand-only pipelines do not jointly address. The object-centric BPS encoding and the explicit mixed-domain schedule are concrete, reusable design choices for the growing literature on diffusion-based human-scene interaction. The paper is purely empirical; it does not ship code, proofs or parameter-free derivations, but the quantitative tables and qualitative figures already demonstrate measurable improvements on standard contact, reconstruction and text-to-motion metrics within the four ParaHome categories.

major comments (3)
  1. The central claim of 'strong generalization to unseen object configurations' (abstract, §1, §5.6) is not supported by a quantitative held-out protocol. Tables 1–3 report reconstruction-style metrics conditioned on the exact initial pose, object state and text of test sequences drawn from the same four ParaHome categories used for training. Section 5.1 states that after contact-based augmentation the corpus contains only 2100 sequences; no split isolating novel placements, scales or joint types is reported, and the conclusion itself concedes failure on unseen rotational mechanisms. Figures 4–6 supply only qualitative anecdotes. Without a quantitative OOD evaluation (e.g., held-out grid cells, novel object scales, or a fifth articulated category), the numerical superiority may largely reflect better in-distribution fitting plus post-hoc noise optimization rather than the claimed generaliza
  2. Section 4.3’s contact-based augmentation (0.1 m grid, random rescaling, CCD-IK with rotational limits) is load-bearing for the diversity claim, yet no ablation quantifies residual kinematic artifacts or their effect on the diffusion prior. Because the later noise-optimization stage (§4.4) can mask many of these artifacts, it remains unclear how much of the reported contact and penetration gains (Table 1) are attributable to the learned model versus the post-processing. An ablation that reports Tables 1–2 both with and without the augmentation (and with/without noise optimization) is required to substantiate the contribution.
  3. The mixed-domain training claim (§4.2) that 'homogeneous batches are imperative' is asserted without external controls. The paper reports neither an ablation that replaces homogeneous batches by mixed batches of equal size nor a comparison against a pure interaction-only baseline that still receives the locomotion mask and FiLM embeddings. Given that the annealing schedule and the 0.5 m locomotion threshold are free parameters, the necessity of the proposed schedule remains unproven.

Circularity Check

0 steps flagged · score 0.0 of 10

Empirical diffusion method paper with no circular derivation; claims rest on standard held-out reconstruction metrics and qualitative generalization figures.

full rationale

GIRAF is a standard computer-vision method paper: it defines an object-centric dynamic BPS representation, a mixed-domain training schedule with FiLM, contact-based augmentation via grid relocation + CCD IK, and post-hoc DDIM noise optimization, then reports reconstruction-style metrics (contact distance, MPJPE, FID, etc.) against two re-trained baselines on ParaHome-derived test sequences plus qualitative examples of height/door variation. None of the load-bearing steps reduce by construction to their inputs. There is no fitted scalar re-labeled as a prediction, no self-definitional identity (e.g., Eq. X ≡ Eq. Y), no uniqueness theorem imported from overlapping authors, and no ansatz smuggled via self-citation. The remark that homogeneous batches are “imperative” is an empirical observation, not a circular premise. Evaluation uses external HOI and text-to-motion metrics on held-out sequences; any over-claim about the breadth of “unseen configurations” is a separate correctness issue, not circularity. Score 0 is therefore the correct, proportionate finding.

Assumptions & free parameters 5 free parameters · 4 assumptions · 1 invented entities

The central empirical claim rests on standard diffusion and body-model machinery plus three engineering choices (dynamic BPS, mixed-domain FiLM+annealing, contact grid augmentation) whose free parameters and domain assumptions are only partially justified by the limited ParaHome+Babel data. No new physical entities are postulated; the invented pieces are representational and training heuristics.

free parameters (5)
  • locomotion distance threshold = 0.5 m
    Frames >0.5 m from object are masked as locomotion (Section 4.2); chosen by hand and directly controls the mixed-domain signal.
  • augmentation grid extents and resolution = 0.1×0.2×0.7 m @ 0.1 m
    Objects relocated on a 0.1×0.2×0.7 m grid at 0.1 m resolution plus random rescaling (Section 4.3); ad-hoc spatial prior that determines training diversity.
  • FiLM locomotion embedding initialization = N(0,0.02)
    Trainable embeddings drawn from N(0,0.02) (Section 5.2); scale chosen by hand.
  • annealing schedule for batch mixing = linear (unspecified exact breakpoints)
    Linear annealing from balanced to interaction-heavy batches (Section 4.2); free schedule that the authors state is 'imperative'.
  • BPS cardinality K and unit-sphere sampling
    Number and distribution of basis points that define the object-centric features; not reported numerically yet controls representation capacity.
assumptions (4)
  • domain assumption SMPL-X body model with 6-D rotations and DistilBERT text embeddings are adequate and fixed representations for the task.
    Adopted without ablation in Section 3; all subsequent claims inherit their biases.
  • domain assumption ParaHome objects are two-part articulated mechanisms with one rotational and one translational DoF; contact can be voted onto a fixed BPS.
    Stated in Section 3 and used to define dynamic BPS; fails for multi-DoF or continuous mechanisms noted in the conclusion.
  • ad hoc to paper Homogeneous batches plus FiLM conditioning suffice to learn seamless locomotion-interaction transitions without mode collapse.
    Asserted as 'imperative' in Section 4.2; supported only by the authors' internal observation.
  • domain assumption CCD inverse-kinematics with rotational limits after contact relocation yields kinematically valid full-body poses.
    Used for all augmented data (Section 4.3); no quantitative validation of residual IK error.
invented entities (1)
  • dynamic BPS (object-centric basis point set encoding distances, hand end-effectors and binary contact)
    purpose: Unify hand-object contact and geometry in a shape-agnostic space so the diffusion model can generalize across object instances.
    Extension of classical BPS; the dynamic, contact-voting, hand-joint concatenation is paper-specific and has no independent external validation beyond the reported tables.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GIRAF: Towards Generalizable Human Interactions with Articulated Objects." pith.science (2026). https://pith.science/paper/534QAA3N

@misc{pith2026260707880,
  author       = {Pith},
  title        = {Pith review of: GIRAF: Towards Generalizable Human Interactions with Articulated Objects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/534QAA3N}},
  note         = {Machine review of arXiv:2607.07880}
}
read the original abstract

Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications in robotics training and virtual agents. Existing models remain limited: some focus on simple activities with static objects, while others restrict attention to hand-only manipulation. This leaves open the problem of generating coordinated full-body motion that approaches, manipulates, and moves articulated objects in a realistic and generalizable way. The key difficulty lies in reasoning jointly about locomotion, fine-grained contact, and object articulation. Models must capture subtle hand-object correspondences that transfer across object geometries, while also producing seamless transitions from navigation to manipulation. At the same time, the scarcity of large-scale paired motion-scene data makes it difficult to generalize across diverse object positions and shapes. We introduce a text-conditioned diffusion model that addresses these challenges through three core ideas: an object-centric representation that unifies hand-object contact with object surfaces, a mixed-domain training strategy that balances locomotion and interaction, and a contact-based augmentation scheme that expands training diversity. Through experiments, our method demonstrated strong generalization to unseen object configurations, surpassing current state-of-the-art methods.

Figures

Figures reproduced from arXiv: 2607.07880 by the authors.

Figure 1
Figure 1. Given an initial pose of the human and the object, along with a textual instruction, our goal is to synthesize realistic [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Architecture overview. GIRAF leverages a transformer-based diffusion model. Given a current pose of the human and the object, [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Qualitative comparison with baselines. By comparing with the baselines, it can be shown that with our novel human-object [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Our model synthesizes realistic sequences of opening different doors of a closet, despite these configurations not being seen [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: We show synthesized sequences of “open the drawer” at two different drawer heights. The model produces consistent reaching [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: For the same object and instruction, our model can synthesize interactions with different hand contact strategies, such as using [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

93 extracted references · 93 canonical work pages

  1. [1]

    Behave: Dataset and method for tracking human object inter- actions

    Bharat Lal Bhatnagar, Xianghui Xie, Ilya Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Behave: Dataset and method for tracking human object inter- actions. InIEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  2. [2]

    Physically plausible full-body hand-object interaction synthesis

    Jona Braun, Sammy Christen, Muhammed Kocabas, Emre Aksan, and Otmar Hilliges. Physically plausible full-body hand-object interaction synthesis

  3. [3]

    Physically plausible full-body hand-object interaction synthesis

    Jona Braun, Sammy Christen, Muhammed Kocabas, Emre Aksan, and Otmar Hilliges. Physically plausible full-body hand-object interaction synthesis. InInternational Confer- ence on 3D Vision (3DV 2024), 2024

  4. [4]

    Gener- ating human motion in 3d scenes from text descriptions

    Zhi Cen, Huaijin Pi, Sida Peng, Zehong Shen, Minghui Yang, Zhu Shuai, Hujun Bao, and Xiaowei Zhou. Gener- ating human motion in 3d scenes from text descriptions. In CVPR, 2024

  5. [5]

    Text2hoi: Text-guided 3d motion generation for hand- object interaction

    Junuk Cha, Jihyeon Kim, Jae Shin Yoon, and Seungryul Baek. Text2hoi: Text-guided 3d motion generation for hand- object interaction. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 1577–1585, 2024

  6. [6]

    Executing your commands via motion diffusion in latent space

    Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, and Gang Yu. Executing your commands via motion diffusion in latent space. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023

  7. [7]

    D-grasp: Physi- cally plausible dynamic grasp synthesis for hand-object in- teractions

    Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D-grasp: Physi- cally plausible dynamic grasp synthesis for hand-object in- teractions. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  8. [8]

    Diffh2o: Diffusion-based synthesis of hand-object interactions from textual descriptions

    Sammy Christen, Shreyas Hampali, Fadime Sener, Edoardo Remelli, Tomas Hodan, Eric Sauser, Shugao Ma, and Bugra Tekin. Diffh2o: Diffusion-based synthesis of hand-object interactions from textual descriptions. InSIGGRAPH Asia 2024 Conference Papers, 2024

Show all 93 references
  1. [9]

    Laserhuman: Language-guided scene-aware human motion generation in free environment, 2024

    Peishan Cong, Ziyi Wang, Zhiyang Dou, Yiming Ren, Wei Yin, Kai Cheng, Yujing Sun, Xiaoxiao Long, Xinge Zhu, and Yuexin Ma. Laserhuman: Language-guided scene-aware human motion generation in free environment, 2024

  2. [10]

    Anyskill: Learning open- vocabulary physical skill for interactive agents

    Jieming Cui, Tengyu Liu, Nian Liu, Yaodong Yang, Yixin Zhu, and Siyuan Huang. Anyskill: Learning open- vocabulary physical skill for interactive agents. InConfer- ence on Computer Vision and Pattern Recognition(CVPR), year=2024

  3. [11]

    Mofusion: A framework for denoising-diffusion-based motion synthesis

    Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt. Mofusion: A framework for denoising-diffusion-based motion synthesis. InCom- puter Vision and Pattern Recognition (CVPR), 2023

  4. [12]

    Cg-hoi: Contact-guided 3d human-object interaction generation

    Christian Diller and Angela Dai. Cg-hoi: Contact-guided 3d human-object interaction generation. 2024

  5. [13]

    Black, and Otmar Hilliges

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Otmar Hilliges. ARCTIC: A dataset for dexterous bimanual hand- object manipulation. InProceedings IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023

  6. [14]

    Imos: Intent-driven full-body motion synthesis for human-object interactions

    Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Chris- tian Theobalt, and Philipp Slusallek. Imos: Intent-driven full-body motion synthesis for human-object interactions. In Eurographics, 2023

  7. [15]

    Stochastic scene-aware motion prediction

    Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael Black. Stochastic scene-aware motion prediction. InProceedings of the Inter- national Conference on Computer Vision 2021, 2021

  8. [16]

    Black, Sanja Fidler, and Xue Bin Peng

    Mohamed Hassan, Yunrong Guo, Tingwu Wang, Michael J. Black, Sanja Fidler, and Xue Bin Peng. Synthesizing phys- ical character-scene interactions.CoRR, abs/2302.00883, 2023

  9. [17]

    Hoang, Kehong Gong, Chuan Guo, and Michael Bi Mi

    Nhat M. Hoang, Kehong Gong, Chuan Guo, and Michael Bi Mi. Motionmix: Weakly-supervised diffusion for control- lable motion generation. InThirty-Eighth Conference on Ar- tificial Intelligence, AAAI 2024

  10. [18]

    Phase- functioned neural networks for character control.ACM Trans

    Daniel Holden, Taku Komura, and Jun Saito. Phase- functioned neural networks for character control.ACM Trans. Graph., 36(4), 2017

  11. [19]

    Hoigpt: Learning long sequence hand-object interac- tion with language models

    Mingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J Liang, Haoyu Ma, Weiyao Wang, Xingyu Chen, Pierre Gleize, Hongfei Xue, Siwei Lyu, Kris Kitani, Matt Feiszli, and Hao Tang. Hoigpt: Learning long sequence hand-object interac- tion with language models. InIEEE Conference on Com- ...

  12. [20]

    Full-body articulated human-object interaction

    Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. Full-body articulated human-object interaction. InICCV, 2023

  13. [21]

    Autonomous character-scene interaction syn- thesis from text instruction

    Nan Jiang, Zimo He, Hongjie Li, Yixin Chen, Siyuan Huang, and Yixin Zhu. Autonomous character-scene interaction syn- thesis from text instruction. InSIGGRAPH Asia Conference Papers, 2024

  14. [22]

    Scaling up dynamic human-scene interaction mod- eling

    Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction mod- eling. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 1737–1747, 2024

  15. [23]

    Op- timizing diffusion noise can serve as universal motion priors

    Korrawe Karunratanakul, Konpat Preechakul, Emre Aksan, Thabo Beeler, Supasorn Suwajanakorn, and Siyu Tang. Op- timizing diffusion noise can serve as universal motion priors. Inarxiv:2312.11994, 2023

  16. [24]

    Parahome: Parameterizing everyday home activities to- wards 3d generative modeling of human-object interactions, 2024

    Jeonghwan Kim, Jisoo Kim, Jeonghyeon Na, and Hanbyul Joo. Parahome: Parameterizing everyday home activities to- wards 3d generative modeling of human-object interactions, 2024

  17. [25]

    Priority-centric human motion gener- ation in discrete latent space

    Hanyang Kong, Kehong Gong, Dongze Lian, Michael Bi Mi, and Xinchao Wang. Priority-centric human motion gener- ation in discrete latent space. InIEEE/CVF International Conference on Computer Vision, ICCV 2023,

  18. [26]

    Nifty: Neural object interaction fields for guided human mo- tion synthesis, 2023

    Nilesh Kulkarni, Davis Rempe, Kyle Genova, Abhijit Kundu, Justin Johnson, David Fouhey, and Leonidas Guibas. Nifty: Neural object interaction fields for guided human mo- tion synthesis, 2023

  19. [27]

    Locomotion-action- manipulation: Synthesizing human-scene interactions in complex 3d environments.arXiv preprint arXiv:2301.02667, 2023

    Jiye Lee and Hanbyul Joo. Locomotion-action- manipulation: Synthesizing human-scene interactions in complex 3d environments.arXiv preprint arXiv:2301.02667, 2023

  20. [28]

    Karen Liu

    Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C. Karen Liu. Controllable human-object interaction synthesis

  21. [29]

    Karen Liu

    Jiaman Li, Jiajun Wu, and C. Karen Liu. Object motion guided human motion synthesis.ACM Transactions on Graphics, 42(6), 2023

  22. [30]

    Karen Liu

    Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C. Karen Liu. Controllable human-object interaction synthesis. InECCV, 2024

  23. [31]

    Task-oriented human-object interactions generation with im- plicit neural representations

    Quanzhou Li, Jingbo Wang, Chen Change Loy, and Bo Dai. Task-oriented human-object interactions generation with im- plicit neural representations. InIEEE/CVF Winter Confer- ence on Applications of Computer Vision, WACV 2024,

  24. [32]

    Learning physics-based full-body human reaching and grasping from brief walking references

    Yitang Li, Mingxian Lin, Zhuo Lin, Yipeng Deng, Yue Cao, and Li Yi. Learning physics-based full-body human reaching and grasping from brief walking references. InCVPR, 2025

  25. [33]

    Geneoh diffusion: Towards general- izable hand-object interaction denoising via denoising diffu- sion

    Xueyi Liu and Li Yi. Geneoh diffusion: Towards general- izable hand-object interaction denoising via denoising diffu- sion. InThe Twelfth International Conference on Learning Representations, 2024

  26. [34]

    Revisit human-scene interaction via space occu- pancy.arXiv preprint arXiv:2312.02700, 2023

    Xinpeng Liu, Haowen Hou, Yanchao Yang, Yong-Lu Li, and Cewu Lu. Revisit human-scene interaction via space occu- pancy.arXiv preprint arXiv:2312.02700, 2023

  27. [35]

    Hu- mantomato: Text-aligned whole-body motion generation

    Shunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin, Ruimao Zhang, Lei Zhang, and Heung-Yeung Shum. Hu- mantomato: Text-aligned whole-body motion generation. arxiv:2310.12978, 2023

  28. [36]

    Winkler, Kris Ki- tani, and Weipeng Xu

    Zhengyi Luo, Jinkun Cao, Alexander W. Winkler, Kris Ki- tani, and Weipeng Xu. Perpetual humanoid control for real- time simulated avatars. InInternational Conference on Com- puter Vision (ICCV), 2023

  29. [37]

    Omnigrasp: Grasping di- verse objects with simulated humanoids.Advances in Neural Information Processing Systems, 2024

    Zhengyi Luo, Jinkun Cao, Sammy Christen, Alexander Win- kler, Kris Kitani, and Weipeng Xu. Omnigrasp: Grasping di- verse objects with simulated humanoids.Advances in Neural Information Processing Systems, 2024

  30. [38]

    Kitani, and Weipeng Xu

    Zhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler, Jing Huang, Kris M. Kitani, and Weipeng Xu. Universal humanoid motion representations for physics-based control. InThe Twelfth International Conference on Learning Repre- sentations, 2024

  31. [39]

    Contact-aware human motion generation from textual de- scriptions.arXiv preprint arXiv:2403.15709, 2024

    Sihan Ma, Qiong Cao, Jing Zhang, and Dacheng Tao. Contact-aware human motion generation from textual de- scriptions.arXiv preprint arXiv:2403.15709, 2024

  32. [40]

    Catch & carry: reusable neural controllers for vision-guided whole-body tasks.ACM Trans

    Josh Merel, Saran Tunyasuvunakool, Arun Ahuja, Yuval Tassa, Leonard Hasenclever, Vu Pham, Tom Erez, Greg Wayne, and Nicolas Heess. Catch & carry: reusable neural controllers for vision-guided whole-body tasks.ACM Trans. Graph., 2020

  33. [41]

    Generating continual human motion in diverse 3d scenes

    Aymen Mir, Xavier Puig, Angjoo Kanazawa, and Gerard Pons-Moll. Generating continual human motion in diverse 3d scenes. InInternational Conference on 3D Vision (3DV), 2024

  34. [42]

    Synthesizing phys- ically plausible human motions in 3d scenes, 2023

    Liang Pan, jingbo Wang, Buzhen Huang, Junyu Zhang, Hao- fan Wang, Xu Tang, and Yangang Wang. Synthesizing phys- ically plausible human motions in 3d scenes, 2023

  35. [43]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. InProceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2019

  36. [44]

    Hoi-diff: Text-driven synthe- sis of 3d human-object interactions using diffusion models

    Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. Hoi-diff: Text-driven synthe- sis of 3d human-object interactions using diffusion models. arXiv preprint arXiv:2312.06553, 2023

  37. [45]

    Courville

    Ethan Perez, Florian Strub, Harm de Vries, Vincent Du- moulin, and Aaron C. Courville. Film: Visual reasoning with a general conditioning layer. InAAAI, 2018

  38. [46]

    Hierarchical generation of human-object inter- actions with diffusion probabilistic models

    Huaijin Pi, Sida Peng, Minghui Yang, Xiaowei Zhou, and Hujun Bao. Hierarchical generation of human-object inter- actions with diffusion probabilistic models. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), 2023

  39. [47]

    Coda: Coordinated diffusion noise optimization for whole-body manipulation of articulated objects.Advances in Neural In- formation Processing Systems, 2025

    Huaijin Pi, Zhi Cen, Zhiyang Dou, and Taku Komura. Coda: Coordinated diffusion noise optimization for whole-body manipulation of articulated objects.Advances in Neural In- formation Processing Systems, 2025

  40. [48]

    Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J

    Abhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J. Black. BABEL: Bodies, action and behavior with english labels. InProceedings IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), 2021

  41. [49]

    Trace and pace: Controllable pedestrian animation via guided tra- jectory diffusion

    Davis Rempe, Zhengyi Luo, Xue Bin Peng, Ye Yuan, Kris Kitani, Karsten Kreis, Sanja Fidler, and Or Litany. Trace and pace: Controllable pedestrian animation via guided tra- jectory diffusion. InConference on Computer Vision and Pattern Recognition (CVPR), 2023

  42. [50]

    Roey Ron, Guy Tevet, Haim Sawdayee, and Amit H. Bermano. Hoidini: Human-object interaction through diffusion noise optimization.arXiv preprint arXiv::2506.15625, 2025

  43. [51]

    Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter.arXiv preprintarXiv:1910.01108, 2019

    V Sanh. Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter.arXiv preprintarXiv:1910.01108, 2019

  44. [52]

    Human motion diffusion as a generative prior

    Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H Bermano. Human motion diffusion as a generative prior. arXiv preprint arXiv:2303.01418, 2023

  45. [53]

    Denois- ing diffusion implicit models.arXiv:2010.02502, 2020

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models.arXiv:2010.02502, 2020

  46. [54]

    Neural state machine for character-scene interactions.ACM Trans

    Sebastian Starke, He Zhang, Taku Komura, and Jun Saito. Neural state machine for character-scene interactions.ACM Trans. Graph., 38(6), 2019

  47. [55]

    Black, and Dim- itrios Tzionas

    Omid Taheri, Nima Ghorbani, Michael J. Black, and Dim- itrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. InEuropean Conference on Computer Vision (ECCV), 2020

  48. [56]

    Black, and Dim- itrios Tzionas

    Omid Taheri, Vasileios Choutas, Michael J. Black, and Dim- itrios Tzionas. GOAL: Generating 4D whole-body motion for hand-object grasping. InConference on Computer Vision and Pattern Recognition (CVPR), 2022

  49. [57]

    Omid Taheri, Yi Zhou, Dimitrios Tzionas, Yang Zhou, Duygu Ceylan, Soren Pirk, and Michael J. Black. Grip: Gen- erating interaction poses conditioned on object and body mo- tion. InInternational Conference on 3D Vision (3DV 2024), 2024

  50. [58]

    Omid Taheri, Yi Zhou, Dimitrios Tzionas, Yang Zhou, Duygu Ceylan, Soren Pirk, and Michael J. Black. GRIP: Generating interaction poses using latent consistency and spatial cues. InInternational Conference on 3D Vision (3DV), 2024

  51. [59]

    Flex: Full- body grasping without full-body grasps

    Purva Tendulkar, D´ıdac Sur´ıs, and Carl V ondrick. Flex: Full- body grasping without full-body grasps. InConference on Computer Vision and Pattern Recognition (CVPR), 2023

  52. [60]

    Maskedmimic: Unified physics-based char- acter control through masked motion inpainting.ACM Trans- actions on Graphics (TOG), 2024

    Chen Tessler, Yunrong Guo, Ofir Nabati, Gal Chechik, and Xue Bin Peng. Maskedmimic: Unified physics-based char- acter control through masked motion inpainting.ACM Trans- actions on Graphics (TOG), 2024

  53. [61]

    Human motion diffu- sion model

    Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffu- sion model. InThe Eleventh International Conference on Learning Representations, 2023

  54. [62]

    CLoSD: Closing the loop between simulation and diffusion for multi-task character control

    Guy Tevet, Sigal Raab, Setareh Cohan, Daniele Reda, Zhengyi Luo, Xue Bin Peng, Amit Haim Bermano, and Michiel van de Panne. CLoSD: Closing the loop between simulation and diffusion for multi-task character control. In The Thirteenth International Conference on Learning Rep- re...

  55. [63]

    Physhoi: Physics-based imita- tion of dynamic human-object interaction.arXiv preprint arXiv:2312.04393, 2023

    Yinhuai Wang, Jing Lin, Ailing Zeng, Zhengyi Luo, Jian Zhang, and Lei Zhang. Physhoi: Physics-based imita- tion of dynamic human-object interaction.arXiv preprint arXiv:2312.04393, 2023

  56. [64]

    Skillmimic: Learning reusable basketball skills from demonstrations, 2024

    Yinhuai Wang, Qihan Zhao, Runyi Yu, Ailing Zeng, Jing Lin, Zhengyi Luo, Hok Wai Tsui, Jiwen Yu, Xiu Li, Qifeng Chen, Jian Zhang, Lei Zhang, and Ping Tan. Skillmimic: Learning reusable basketball skills from demonstrations, 2024

  57. [65]

    Humanise: Language-conditioned hu- man motion generation in 3d scenes

    Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang. Humanise: Language-conditioned hu- man motion generation in 3d scenes. InAdvances in Neural Information Processing Systems (NeurIPS), 2022

  58. [66]

    Move as you say, interact as you can: Language-guided human motion generation with scene af- fordance

    Zan Wang, Yixin Chen, Baoxiong Jia, Puhao Li, Jinlu Zhang, Jingze Zhang, Tengyu Liu, Yixin Zhu, Wei Liang, and Siyuan Huang. Move as you say, interact as you can: Language-guided human motion generation with scene af- fordance. InProceedings of the IEEE/CVF Conference on Compu...

  59. [67]

    THOR: text to human-object interaction diffusion via relation intervention.CoRR, abs/2403.11208, 2024

    Qianyang Wu, Ye Shi, Xiaoshui Huang, Jingyi Yu, Lan Xu, and Jingya Wang. THOR: text to human-object interaction diffusion via relation intervention.CoRR, abs/2403.11208, 2024

  60. [68]

    Uniphys: Unified planner and controller with diffusion for flexible physics-based character control

    Yan Wu, Korrawe Karunratanakul1and, Zhengyi Luo, and Siyu Tang. Uniphys: Unified planner and controller with diffusion for flexible physics-based character control. InIn- ternational Conference on Computer Vision (ICCV), 2025

  61. [69]

    Human- object interaction from human-level instructions

    Zhen Wu, Jiaman Li, Pei Xu, and C Karen Liu. Human- object interaction from human-level instructions. In IEEE/CVF International Conference on Computer Vision, ICCV, 2025

  62. [70]

    Motionstreamer: Streaming motion genera- tion via diffusion-based autoregressive model in causal latent space.arXiv preprint arXiv:2503.15451, 2025

    Lixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan, Liang Pan, Yueer Zhou, Ziyong Feng, Xiaowei Zhou, Sida Peng, and Jingbo Wang. Motionstreamer: Streaming motion genera- tion via diffusion-based autoregressive model in causal latent space.arXiv preprint arXiv:2503.15451, 2025

  63. [71]

    Interdiff: Generating 3d human-object interactions with physics-informed diffusion

    Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui. Interdiff: Generating 3d human-object interactions with physics-informed diffusion. InICCV, 2023

  64. [72]

    Interdreamer: Zero-shot text to 3d dynamic human-object interaction.arXiv preprint arXiv:2403.19652, 2024

    Sirui Xu, Ziyin Wang, Yu-Xiong Wang, and Liang-Yan Gui. Interdreamer: Zero-shot text to 3d dynamic human-object interaction.arXiv preprint arXiv:2403.19652, 2024

  65. [73]

    Interact: Advancing large-scale versatile 3d human-object interaction generation

    Sirui Xu, Dongting Li, Yucheng Zhang, Xiyan Xu, Qi Long, Ziyin Wang, Yunzhi Lu, Shuchang Dong, Hezi Jiang, Ak- shat Gupta, Yu-Xiong Wang, and Liang-Yan Gui. Interact: Advancing large-scale versatile 3d human-object interaction generation. InCVPR, 2025

  66. [74]

    Intermimic: Towards universal whole-body control for physics-based human-object interactions

    Sirui Xu, Hung Yu Ling, Yu-Xiong Wang, and Liangyan Gui. Intermimic: Towards universal whole-body control for physics-based human-object interactions. InCVPR, 2025

  67. [75]

    F-hoi: Toward fine-grained semantic-aligned 3d human-object interactions.European Conference on Computer Vision, 2024

    Jie Yang, Xuesong Niu, Nan Jiang, Ruimao Zhang, and Huang Siyuan. F-hoi: Toward fine-grained semantic-aligned 3d human-object interactions.European Conference on Computer Vision, 2024

  68. [76]

    Black, Xue Bin Peng, and Davis Rempe

    Hongwei Yi, Justus Thies, Michael J. Black, Xue Bin Peng, and Davis Rempe. Generating human interaction motions in scenes with text control.arXiv:2404.10685, 2024

  69. [77]

    Core4d: A 4d human-object-human interac- tion dataset for collaborative object rearrangement.arXiv preprint arXiv:2406.19353, 2024

    Chengwen Zhang, Yun Liu, Ruofan Xing, Bingda Tang, and Li Yi. Core4d: A 4d human-object-human interac- tion dataset for collaborative object rearrangement.arXiv preprint arXiv:2406.19353, 2024

  70. [78]

    Manipnet: neural manipulation synthesis with a hand-object spatial representation.ACM Trans

    He Zhang, Yuting Ye, Takaaki Shiratori, and Taku Komura. Manipnet: neural manipulation synthesis with a hand-object spatial representation.ACM Trans. Graph., 2021

  71. [79]

    Hoi-m3: Capture multiple humans and objects in- teraction within contextual environment.arXiv preprint arXiv:2404.00299, 2024

    Juze Zhang, Jingyan Zhang, Zining Song, Zhanhe Shi, Chengfeng Zhao, Ye Shi, Jingyi Yu, Lan Xu, and Jingya Wang. Hoi-m3: Capture multiple humans and objects in- teraction within contextual environment.arXiv preprint arXiv:2404.00299, 2024

  72. [80]

    Re- modiffuse: Retrieval-augmented motion diffusion model

    Mingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai, Fangzhou Hong, Huirong Li, Lei Yang, and Ziwei Liu. Re- modiffuse: Retrieval-augmented motion diffusion model. In 2023 IEEE/CVF International Conference on Computer Vi- sion, ICCV

  73. [81]

    Motiondif- fuse: Text-driven human motion generation with diffusion model.arXiv preprint arXiv:2208.15001, 2022

    Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. Motiondif- fuse: Text-driven human motion generation with diffusion model.arXiv preprint arXiv:2208.15001, 2022

  74. [82]

    Roam: Robust and object-aware motion gen- eration using neural pose descriptors.arXiv preprint arXiv:2308.12969, 2023

    Wanyue Zhang, Rishabh Dabral, Thomas Leimk ¨uhler, Vladislav Golyanik, Marc Habermann, and Christian Theobalt. Roam: Robust and object-aware motion gen- eration using neural pose descriptors.arXiv preprint arXiv:2308.12969, 2023

  75. [83]

    Couch: Towards controllable human-chair interactions

    Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Vladimir Guzov, and Gerard Pons-Moll. Couch: Towards controllable human-chair interactions. InEuropean Confer- ence on Computer Vision (ECCV), 2022

  76. [84]

    Petrov, Vladimir Guzov, Helisa Dhamo, Eduardo P´erez Pellitero, and Gerard Pons-Moll

    Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Ilya A. Petrov, Vladimir Guzov, Helisa Dhamo, Eduardo P´erez Pellitero, and Gerard Pons-Moll. Force: Dataset and method for intuitive physics guided human-object interac- tion. InArxiv, 2024

  77. [85]

    Scenic: Scene-aware semantic navigation with instruction- guided control

    Xiaohan Zhang, Sebastian Starke, Vladimir Guzov, Helisa Dhamo, Eduardo P ´erez Pellitero, and Gerard Pons-Moll. Scenic: Scene-aware semantic navigation with instruction- guided control. 2024

  78. [86]

    Diffgrasp: Whole-body grasping synthesis guided by object motion us- ing a diffusion model

    Yonghao Zhang, Qiang He, Yanguang Wan, Yinda Zhang, Xiaoming Deng, Cuixia Ma, and Hongan Wang. Diffgrasp: Whole-body grasping synthesis guided by object motion us- ing a diffusion model. InAAAI, 2025

  79. [87]

    Tedi: Temporally-entangled diffusion for long- term motion synthesis

    Zihan Zhang, Richard Liu, Kfir Aberman, and Rana Hanocka. Tedi: Temporally-entangled diffusion for long- term motion synthesis. InSIGGRAPH, Technical Papers, 2024

  80. [88]

    Synthesizing diverse human motions in 3d indoor scenes

    Kaifeng Zhao, Yan Zhang, Shaofei Wang, Thabo Beeler, , and Siyu Tang. Synthesizing diverse human motions in 3d indoor scenes. InInternational conference on computer vi- sion (ICCV), 2023

  81. [89]

    A diffusion-based au- toregressive motion model for real-time text-driven motion control

    Kaifeng Zhao, Gen Li, and Siyu Tang. A diffusion-based au- toregressive motion model for real-time text-driven motion control. InArxiv, 2024

  82. [90]

    DartControl: A diffusion-based autoregressive motion model for real-time text-driven motion control

    Kaifeng Zhao, Gen Li, and Siyu Tang. DartControl: A diffusion-based autoregressive motion model for real-time text-driven motion control. InThe Thirteenth International Conference on Learning Representations (ICLR), 2025

  83. [91]

    Cams: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthe- sis

    Juntian Zheng, Qingyuan Zheng, Lixing Fang, Yun Liu, and Li Yi. Cams: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthe- sis. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 585– ...

  84. [92]

    Emdm: Efficient mo- tion diffusion model for fast, high-quality motion generation

    Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang, Wenjia Wang, Yuan Liu, Taku Komura, Wenping Wang, and Lingjie Liu. Emdm: Efficient mo- tion diffusion model for fast, high-quality motion generation. arXiv preprint arXiv:2312.02256, 2023

  85. [93]

    On the continuity of rotation representations in neural networks

    Yi Zhou, Connelly Barnes, Lu Jingwan, Yang Jimei, and Li Hao. On the continuity of rotation representations in neural networks. InThe IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.