Pith. sign in

REVIEW 14 cited by

THOR: Text to Human-Object Interaction Diffusion via Relation Intervention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11208 v1 pith:LM456ARL submitted 2024-03-17 cs.CV

classification cs.CV
keywords motiondiffusionhuman-objectinteractioninterventionobjectmodelrelation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper addresses new methodologies to deal with the challenging task of generating dynamic Human-Object Interactions from textual descriptions (Text2HOI). While most existing works assume interactions with limited body parts or static objects, our task involves addressing the variation in human motion, the diversity of object shapes, and the semantic vagueness of object motion simultaneously. To tackle this, we propose a novel Text-guided Human-Object Interaction diffusion model with Relation Intervention (THOR). THOR is a cohesive diffusion model equipped with a relation intervention mechanism. In each diffusion step, we initiate text-guided human and object motion and then leverage human-object relations to intervene in object motion. This intervention enhances the spatial-temporal relations between humans and objects, with human-centric interaction representation providing additional guidance for synthesizing consistent motion from text. To achieve more reasonable and realistic results, interaction losses is introduced at different levels of motion granularity. Moreover, we construct Text-BEHAVE, a Text2HOI dataset that seamlessly integrates textual descriptions with the currently largest publicly available 3D HOI dataset. Both quantitative and qualitative experiments demonstrate the effectiveness of our proposed model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Distilling frozen Motion-JEPA features into a compact 32-D latent whose geometry is coupled to the decoder lets a standard non-autoregressive flow-matching DiT reach state-of-the-art text-to-motion quality on HumanML3...

  2. Absolute Coordinates Make Motion Generation Easy

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Using absolute 3D joint coordinates with a plain Transformer and velocity-prediction diffusion outperforms the standard local-relative motion representation, improving fidelity and enabling direct control.

  3. Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression

    cs.CV 2024-11 conditional novelty 7.0 of 10

    A masked-autoregressive diffusion model trained on a compact essential-feature latent space claims state-of-the-art text-to-motion generation under a new essential-dimension evaluation protocol.

  4. GIRAF: Towards Generalizable Human Interactions with Articulated Objects

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat ad...

  5. InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    InterAct is a unified 21.81-hour 3D human-object interaction benchmark with text annotations, quality-corrected data, and a multi-task model that achieves state-of-the-art results across six generation tasks.

  6. SimGenHOI: Physically Realistic Whole-Body Humanoid-Object Interaction via Generative Modeling and Reinforcement Learning

    cs.RO 2025-08 conditional novelty 6.0 of 10

    SimGenHOI generates physically plausible humanoid-object interaction sequences by predicting sparse key actions with a diffusion transformer and tracking them with a contact-aware reinforcement learning policy in simulation.

  7. GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GenHOI generates 4D human-object interaction sequences for unseen objects by predicting sparse 3D keyframes and interpolating them with a contact-aware diffusion model.

  8. SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SyncDiff synthesizes multi-body human-object interaction motions with one diffusion model plus explicit synchronization and frequency decomposition, improving contact and action-quality metrics over prior methods on f...

  9. SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion model generates human motion that simultaneously follows text instructions and adapts to complex 3D terrain, using goal-centric canonicalization and an ego-centric distance field.

  10. InterDyn: Controllable Interactive Dynamics with Video Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    InterDyn fine-tunes Stable Video Diffusion with a ControlNet-style branch so that a hand-mask control signal drives plausible, temporally consistent videos of object interactions.

  11. MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MulSMo uses a bidirectional control flow between style and content networks and contrastive learning over motion, text, and images to generate stylized human motions from multimodal style inputs.

  12. Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Human-X jointly predicts actions and reactions in real time to produce physically plausible human-machine interaction motion.

  13. Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A temporal feature warping and attention fusion module for amodal completion improves occlusion handling and temporal stability in monocular HOI videos, and the completed frames support 3D Gaussian Splatting reconstruction.

  14. UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes

    cs.GR 2025-05 conditional novelty 4.0 of 10

    UniHM merges continuous 6DoF motion with discrete local tokens in one autoregressive plus diffusion model for scene-aware text-to-motion and text-to-HOI, but its claimed first-ness and ablations are not fully consistent.

Pith tools