REVIEW 21 cited by
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop a dual-branch diffusion model (HOI-DM) to generate both human and object motions conditioned on the input text, and encourage coherent motions by a cross-attention communication module between the human and object motion generation branches. We also develop an affordance prediction diffusion model (APDM) to predict the contacting area between the human and object during the interactions driven by the textual prompt. The APDM is independent of the results by the HOI-DM and thus can correct potential errors by the latter. Moreover, it stochastically generates the contacting points to diversify the generated motions. Finally, we incorporate the estimated contacting points into the classifier-guidance to achieve accurate and close contact between humans and objects. To train and evaluate our approach, we annotate BEHAVE dataset with text descriptions. Experimental results on BEHAVE and OMOMO demonstrate that our approach produces realistic HOIs with various interactions and different types of objects.
Forward citations
Cited by 21 Pith papers
-
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
Distilling frozen Motion-JEPA features into a compact 32-D latent whose geometry is coupled to the decoder lets a standard non-autoregressive flow-matching DiT reach state-of-the-art text-to-motion quality on HumanML3...
-
HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis
HarmoHOI jointly generates synchronized multi-view hand-object interaction videos and globally aligned 3D point tracks from a single reference image and target camera poses.
-
Absolute Coordinates Make Motion Generation Easy
Using absolute 3D joint coordinates with a plain Transformer and velocity-prediction diffusion outperforms the standard local-relative motion representation, improving fidelity and enabling direct control.
-
Diffgrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model
DiffGrasp synthesizes full-body grasping motion sequences with realistic hand-object contact from object shape and motion via a single conditional diffusion model.
-
BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects
BimArt synthesizes diverse, physically plausible bimanual hand motions for articulated objects from object trajectories alone, using a part-aware basis-point representation and predicted distance-based contact maps.
-
AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment
AgentHOI generates human-object interaction videos from text plus one human image and one object image, using multi-agent action planning and implicit text-to-motion feature alignment inside a video diffusion model.
-
ContactMimic: Humanoid Object Interaction via Contact Control
A humanoid tracking policy is trained with contact-following rewards and trajectory augmentation to decouple physical contact from keypoint geometry, enabling runtime contact control.
-
GIRAF: Towards Generalizable Human Interactions with Articulated Objects
A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat ad...
-
InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
InterAct is a unified 21.81-hour 3D human-object interaction benchmark with text annotations, quality-corrected data, and a multi-task model that achieves state-of-the-art results across six generation tasks.
-
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Claimed first large-scale egocentric and multi-view dataset of human-object-human assistance (11.4 hours, 1.2M frames) with three benchmarks; only the abstract was assessable because the submitted body text is a diffe...
-
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
GenHOI generates 4D human-object interaction sequences for unseen objects by predicting sparse 3D keyframes and interpolating them with a contact-aware diffusion model.
-
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.
-
HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
HunyuanVideo-HOMA generates human-object interaction videos from weak, sparse inputs: one arm pose, an object center dot, a human photo, and an object photo.
-
InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba
InterMamba introduces an adaptive spatial-temporal Mamba with self and cross interaction blocks for text-driven human-human interaction generation, reporting better text-motion alignment and much lower compute than InterGen.
-
SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
SViMo jointly generates HOI videos and explicit 3D hand-object motion via synchronized diffusion with a closed-loop 3D interaction diffusion model.
-
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
InteractAnything combines LLM-based relation reasoning, inpainting-based object affordance parsing, and force-closure-inspired optimization to synthesize zero-shot 3D human-object interactions from text and an object mesh.
-
CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects
CoDA generates coordinated whole-body articulated-object manipulation by optimizing the noise of three decoupled diffusion models, guided by BPS-based end-effector and object trajectories.
-
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
SyncDiff synthesizes multi-body human-object interaction motions with one diffusion model plus explicit synchronization and frequency decomposition, improving contact and action-quality metrics over prior methods on f...
-
Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking
Mimicking-Bench provides six humanoid-scene interaction tasks with 23K human motion references and a retarget-track-imitate pipeline that beats data-free RL on average success.
-
SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control
A diffusion model generates human motion that simultaneously follows text instructions and adapts to complex 3D terrain, using goal-centric canonicalization and an ego-centric distance field.
-
MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow
MulSMo uses a bidirectional control flow between style and content networks and contrastive learning over motion, text, and images to generate stylized human motions from multimodal style inputs.
Discussion (0). Continue with ORCID to comment.