Pith. sign in

REVIEW 21 cited by

HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06553 v3 pith:W6TYX47P submitted 2023-12-11 cs.CV

classification cs.CV
keywords interactionscontactingdiffusionhumanmotionsobjectapdmapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop a dual-branch diffusion model (HOI-DM) to generate both human and object motions conditioned on the input text, and encourage coherent motions by a cross-attention communication module between the human and object motion generation branches. We also develop an affordance prediction diffusion model (APDM) to predict the contacting area between the human and object during the interactions driven by the textual prompt. The APDM is independent of the results by the HOI-DM and thus can correct potential errors by the latter. Moreover, it stochastically generates the contacting points to diversify the generated motions. Finally, we incorporate the estimated contacting points into the classifier-guidance to achieve accurate and close contact between humans and objects. To train and evaluate our approach, we annotate BEHAVE dataset with text descriptions. Experimental results on BEHAVE and OMOMO demonstrate that our approach produces realistic HOIs with various interactions and different types of objects.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Distilling frozen Motion-JEPA features into a compact 32-D latent whose geometry is coupled to the decoder lets a standard non-autoregressive flow-matching DiT reach state-of-the-art text-to-motion quality on HumanML3...

  2. HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

    cs.CV 2026-07 conditional novelty 7.0 of 10

    HarmoHOI jointly generates synchronized multi-view hand-object interaction videos and globally aligned 3D point tracks from a single reference image and target camera poses.

  3. Absolute Coordinates Make Motion Generation Easy

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Using absolute 3D joint coordinates with a plain Transformer and velocity-prediction diffusion outperforms the standard local-relative motion representation, improving fidelity and enabling direct control.

  4. Diffgrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model

    cs.CV 2024-12 conditional novelty 7.0 of 10

    DiffGrasp synthesizes full-body grasping motion sequences with realistic hand-object contact from object shape and motion via a single conditional diffusion model.

  5. BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects

    cs.CV 2024-12 conditional novelty 7.0 of 10

    BimArt synthesizes diverse, physically plausible bimanual hand motions for articulated objects from object trajectories alone, using a part-aware basis-point representation and predicted distance-based contact maps.

  6. AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    AgentHOI generates human-object interaction videos from text plus one human image and one object image, using multi-agent action planning and implicit text-to-motion feature alignment inside a video diffusion model.

  7. ContactMimic: Humanoid Object Interaction via Contact Control

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A humanoid tracking policy is trained with contact-following rewards and trajectory augmentation to decouple physical contact from keypoint geometry, enabling runtime contact control.

  8. GIRAF: Towards Generalizable Human Interactions with Articulated Objects

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat ad...

  9. InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    InterAct is a unified 21.81-hour 3D human-object interaction benchmark with text annotations, quality-corrected data, and a multi-task model that achieves state-of-the-art results across six generation tasks.

  10. Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Claimed first large-scale egocentric and multi-view dataset of human-object-human assistance (11.4 hours, 1.2M frames) with three benchmarks; only the abstract was assessable because the submitted body text is a diffe...

  11. GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GenHOI generates 4D human-object interaction sequences for unseen objects by predicting sparse 3D keyframes and interpolating them with a contact-aware diffusion model.

  12. DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.

  13. HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    HunyuanVideo-HOMA generates human-object interaction videos from weak, sparse inputs: one arm pose, an object center dot, a human photo, and an object photo.

  14. InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba

    cs.CV 2025-06 conditional novelty 6.0 of 10

    InterMamba introduces an adaptive spatial-temporal Mamba with self and cross interaction blocks for text-driven human-human interaction generation, reporting better text-motion alignment and much lower compute than InterGen.

  15. SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SViMo jointly generates HOI videos and explicit 3D hand-object motion via synchronized diffusion with a closed-loop 3D interaction diffusion model.

  16. InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing

    cs.CV 2025-05 conditional novelty 6.0 of 10

    InteractAnything combines LLM-based relation reasoning, inpainting-based object affordance parsing, and force-closure-inspired optimization to synthesize zero-shot 3D human-object interactions from text and an object mesh.

  17. CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects

    cs.GR 2025-05 conditional novelty 6.0 of 10

    CoDA generates coordinated whole-body articulated-object manipulation by optimizing the noise of three decoupled diffusion models, guided by BPS-based end-effector and object trajectories.

  18. SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SyncDiff synthesizes multi-body human-object interaction motions with one diffusion model plus explicit synchronization and frequency decomposition, improving contact and action-quality metrics over prior methods on f...

  19. Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking

    cs.RO 2024-12 conditional novelty 6.0 of 10

    Mimicking-Bench provides six humanoid-scene interaction tasks with 23K human motion references and a retarget-track-imitate pipeline that beats data-free RL on average success.

  20. SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion model generates human motion that simultaneously follows text instructions and adapts to complex 3D terrain, using goal-centric canonicalization and an ego-centric distance field.

  21. MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MulSMo uses a bidirectional control flow between style and content networks and contrastive learning over motion, text, and images to generate stylized human motions from multimodal style inputs.

Pith tools