REVIEW 14 cited by
THOR: Text to Human-Object Interaction Diffusion via Relation Intervention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper addresses new methodologies to deal with the challenging task of generating dynamic Human-Object Interactions from textual descriptions (Text2HOI). While most existing works assume interactions with limited body parts or static objects, our task involves addressing the variation in human motion, the diversity of object shapes, and the semantic vagueness of object motion simultaneously. To tackle this, we propose a novel Text-guided Human-Object Interaction diffusion model with Relation Intervention (THOR). THOR is a cohesive diffusion model equipped with a relation intervention mechanism. In each diffusion step, we initiate text-guided human and object motion and then leverage human-object relations to intervene in object motion. This intervention enhances the spatial-temporal relations between humans and objects, with human-centric interaction representation providing additional guidance for synthesizing consistent motion from text. To achieve more reasonable and realistic results, interaction losses is introduced at different levels of motion granularity. Moreover, we construct Text-BEHAVE, a Text2HOI dataset that seamlessly integrates textual descriptions with the currently largest publicly available 3D HOI dataset. Both quantitative and qualitative experiments demonstrate the effectiveness of our proposed model.
Forward citations
Cited by 14 Pith papers
-
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
Distilling frozen Motion-JEPA features into a compact 32-D latent whose geometry is coupled to the decoder lets a standard non-autoregressive flow-matching DiT reach state-of-the-art text-to-motion quality on HumanML3...
-
Absolute Coordinates Make Motion Generation Easy
Using absolute 3D joint coordinates with a plain Transformer and velocity-prediction diffusion outperforms the standard local-relative motion representation, improving fidelity and enabling direct control.
-
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
A masked-autoregressive diffusion model trained on a compact essential-feature latent space claims state-of-the-art text-to-motion generation under a new essential-dimension evaluation protocol.
-
GIRAF: Towards Generalizable Human Interactions with Articulated Objects
A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat ad...
-
InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
InterAct is a unified 21.81-hour 3D human-object interaction benchmark with text annotations, quality-corrected data, and a multi-task model that achieves state-of-the-art results across six generation tasks.
-
SimGenHOI: Physically Realistic Whole-Body Humanoid-Object Interaction via Generative Modeling and Reinforcement Learning
SimGenHOI generates physically plausible humanoid-object interaction sequences by predicting sparse key actions with a diffusion transformer and tracking them with a contact-aware reinforcement learning policy in simulation.
-
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
GenHOI generates 4D human-object interaction sequences for unseen objects by predicting sparse 3D keyframes and interpolating them with a contact-aware diffusion model.
-
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
SyncDiff synthesizes multi-body human-object interaction motions with one diffusion model plus explicit synchronization and frequency decomposition, improving contact and action-quality metrics over prior methods on f...
-
SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control
A diffusion model generates human motion that simultaneously follows text instructions and adapts to complex 3D terrain, using goal-centric canonicalization and an ego-centric distance field.
-
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models
InterDyn fine-tunes Stable Video Diffusion with a ControlNet-style branch so that a hand-mask control signal drives plausible, temporally consistent videos of object interactions.
-
MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow
MulSMo uses a bidirectional control flow between style and content networks and contrastive learning over motion, text, and images to generate stylized human motions from multimodal style inputs.
-
Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
Human-X jointly predicts actions and reactions in real time to produce physically plausible human-machine interaction motion.
-
Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
A temporal feature warping and attention fusion module for amodal completion improves occlusion handling and temporal stability in monocular HOI videos, and the completed frames support 3D Gaussian Splatting reconstruction.
-
UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes
UniHM merges continuous 6DoF motion with discrete local tokens in one autoregressive plus diffusion model for scene-aware text-to-motion and text-to-HOI, but its claimed first-ness and ablations are not fully consistent.
Discussion (0). Continue with ORCID to comment.