Pith. sign in

REVIEW 9 cited by

Contact-aware Human Motion Generation from Textual Descriptions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.15709 v2 pith:75QNTOGR submitted 2024-03-23 cs.CV cs.AI

classification cs.CVcs.AI
keywords motionhumantextualbodycontactcontactsrich-catsequences
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper addresses the problem of generating 3D interactive human motion from text. Given a textual description depicting the actions of different body parts in contact with static objects, we synthesize sequences of 3D body poses that are visually natural and physically plausible. Yet, this task poses a significant challenge due to the inadequate consideration of interactions by physical contacts in both motion and textual descriptions, leading to unnatural and implausible sequences. To tackle this challenge, we create a novel dataset named RICH-CAT, representing "Contact-Aware Texts" constructed from the RICH dataset. RICH-CAT comprises high-quality motion, accurate human-object contact labels, and detailed textual descriptions, encompassing over 8,500 motion-text pairs across 26 indoor/outdoor actions. Leveraging RICH-CAT, we propose a novel approach named CATMO for text-driven interactive human motion synthesis that explicitly integrates human body contacts as evidence. We employ two VQ-VAE models to encode motion and body contact sequences into distinct yet complementary latent spaces and an intertwined GPT for generating human motions and contacts in a mutually conditioned manner. Additionally, we introduce a pre-trained text encoder to learn textual embeddings that better discriminate among various contact types, allowing for more precise control over synthesized motions and contacts. Our experiments demonstrate the superior performance of our approach compared to existing text-to-motion methods, producing stable, contact-aware motion sequences. Code and data will be available for research purposes at https://xymsh.github.io/RICH-CAT/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Distilling frozen Motion-JEPA features into a compact 32-D latent whose geometry is coupled to the decoder lets a standard non-autoregressive flow-matching DiT reach state-of-the-art text-to-motion quality on HumanML3...

  2. Absolute Coordinates Make Motion Generation Easy

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Using absolute 3D joint coordinates with a plain Transformer and velocity-prediction diffusion outperforms the standard local-relative motion representation, improving fidelity and enabling direct control.

  3. Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression

    cs.CV 2024-11 conditional novelty 7.0 of 10

    A masked-autoregressive diffusion model trained on a compact essential-feature latent space claims state-of-the-art text-to-motion generation under a new essential-dimension evaluation protocol.

  4. Interactive Generative Motion Editing via Scheduled Inpainting

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Scheduled inpainting blends a base motion clip into a diffusion model's denoising process via a user-controlled schedule and spatiotemporal mask, enabling interactive editing of existing animations without retraining.

  5. GIRAF: Towards Generalizable Human Interactions with Articulated Objects

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat ad...

  6. PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning

    cs.CV 2025-04 conditional novelty 6.0 of 10

    ProMoGen generates human motion conditioned on both a trajectory and sparse anchor postures via a diffusion transformer trained with a dense-to-sparse curriculum.

  7. SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion model generates human motion that simultaneously follows text instructions and adapts to complex 3D terrain, using goal-centric canonicalization and an ego-centric distance field.

  8. Towards Consistent Long-Term Pose Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A one-stage Transformer with placeholder tokens generates continuous 2D pose sequences from a single image and text, avoiding autoregressive error accumulation.

  9. Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.

Pith tools