Pith. sign in

Hoi-diff: Text-driven syn- thesis of 3d human-object interactions using diffusion mod- els

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it
abstract

We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop a dual-branch diffusion model (HOI-DM) to generate both human and object motions conditioned on the input text, and encourage coherent motions by a cross-attention communication module between the human and object motion generation branches. We also develop an affordance prediction diffusion model (APDM) to predict the contacting area between the human and object during the interactions driven by the textual prompt. The APDM is independent of the results by the HOI-DM and thus can correct potential errors by the latter. Moreover, it stochastically generates the contacting points to diversify the generated motions. Finally, we incorporate the estimated contacting points into the classifier-guidance to achieve accurate and close contact between humans and objects. To train and evaluate our approach, we annotate BEHAVE dataset with text descriptions. Experimental results on BEHAVE and OMOMO demonstrate that our approach produces realistic HOIs with various interactions and different types of objects.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 4 cs.RO 1

years

2026 4 2025 1

roles

background 1

polarities

background 1

representative citing papers

GenHSI: Controllable Generation of Human-Scene Interaction Videos

cs.CV · 2025-06-24 · unverdicted · novelty 7.0

GenHSI is a training-free three-stage pipeline that turns a scene image, character image, and complex HSI prompt into long videos with plausible chained interactions by generating atomic actions, 3D keyframes via 2D inpainting plus optimization, and then feeding them to pre-trained video diffusion.

GIRAF: Towards Generalizable Human Interactions with Articulated Objects

cs.CV · 2026-07-08 · conditional · novelty 6.0

A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat adapted baselines on contact and pose metrics.

citing papers explorer

Showing 5 of 5 citing papers.

  • Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics cs.CV · 2026-06-22 · unverdicted · none · ref 39

    A framework called Policy-as-Data generates task-oriented synthetic HOI data via RL policies in physics simulators, retargets it, and trains diffusion models that generalize to unseen objects and long horizons.

  • Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation cs.CV · 2026-05-13 · conditional · none · ref 31

    CMC decouples trajectory control and text-conditioned motion completion with selective inpainting to achieve state-of-the-art accuracy and quality in multimodal human motion generation.

  • GenHSI: Controllable Generation of Human-Scene Interaction Videos cs.CV · 2025-06-24 · unverdicted · none · ref 70

    GenHSI is a training-free three-stage pipeline that turns a scene image, character image, and complex HSI prompt into long videos with plausible chained interactions by generating atomic actions, 3D keyframes via 2D inpainting plus optimization, and then feeding them to pre-trained video diffusion.

  • ContactMimic: Humanoid Object Interaction via Contact Control cs.RO · 2026-07-09 · conditional · none · ref 38 · internal anchor

    A humanoid tracking policy is trained with contact-following rewards and trajectory augmentation to decouple physical contact from keypoint geometry, enabling runtime contact control.

  • GIRAF: Towards Generalizable Human Interactions with Articulated Objects cs.CV · 2026-07-08 · conditional · none · ref 44 · internal anchor

    A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat adapted baselines on contact and pose metrics.