Pith. sign in

REVIEW 21 cited by

HumanTOMATO: Text-aligned Whole-body Motion Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12978 v1 pith:L7JA2TCT submitted 2023-10-19 cs.CV

classification cs.CV
keywords motiongenerationwhole-bodyhandalignmentbodydescriptionfine-grained
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

This work targets a novel text-driven whole-body motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and body motions simultaneously. Previous works on text-driven motion generation tasks mainly have two limitations: they ignore the key role of fine-grained hand and face controlling in vivid whole-body motion generation, and lack a good alignment between text and motion. To address such limitations, we propose a Text-aligned whOle-body Motion generATiOn framework, named HumanTOMATO, which is the first attempt to our knowledge towards applicable holistic motion generation in this research area. To tackle this challenging task, our solution includes two key designs: (1) a Holistic Hierarchical VQ-VAE (aka H$^2$VQ) and a Hierarchical-GPT for fine-grained body and hand motion reconstruction and generation with two structured codebooks; and (2) a pre-trained text-motion-alignment model to help generated motion align with the input textual description explicitly. Comprehensive experiments verify that our model has significant advantages in both the quality of generated motions and their alignment with text.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffgrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model

    cs.CV 2024-12 conditional novelty 7.0 of 10

    DiffGrasp synthesizes full-body grasping motion sequences with realistic hand-object contact from object shape and motion via a single conditional diffusion model.

  2. Motion4Motion: Motion Transfer Across Subjects at Inference

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training-free motion transfer across species works by extracting source motion flows, matching semantic points, and injecting them into DiT self-attention via TransPE positional padding.

  3. GIRAF: Towards Generalizable Human Interactions with Articulated Objects

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat ad...

  4. InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    InterAct is a unified 21.81-hour 3D human-object interaction benchmark with text annotations, quality-corrected data, and a multi-task model that achieves state-of-the-art results across six generation tasks.

  5. Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Being-M0.5 combines part-aware residual quantization with a 5M-sequence web-video dataset to reach real-time, part-controllable 3D motion generation, though its state-of-the-art claim does not hold on every standard b...

  6. The loss tolerance of cat breeding for fault-tolerant grid state generation

    quant-ph 2025-08 reject novelty 6.0 of 10

    Claims a 4% optical-loss ceiling for fault-tolerant GKP state generation via cat breeding, but the provided full text is an unrelated manuscript with no such analysis.

  7. Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dexterous VLA pretrained on a 2.5M-instance human hand motion dataset transfers skills to a real robot hand, outperforming baselines in manipulation tasks.

  8. Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling

    cs.MM 2025-07 conditional novelty 6.0 of 10

    A new captured music-dance dataset with facial expressions and a hierarchical residual VQ plus masked-transformer model that generates expressive 3D dance from music.

  9. PhysiInter: Integrating Physical Mapping for High-Fidelity Human Interaction Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A text-to-motion pipeline that projects motions through physics-based imitation for training and post-processing, plus new consistency and marker-interaction losses.

  10. Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues

    cs.AI 2025-06 conditional novelty 6.0 of 10

    VENUS is a large podcast-derived dataset aligning text with 3D facial and body cues, and MARS is an LLM fine-tuned on it to generate both words and nonverbal tokens.

  11. PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning

    cs.CV 2025-04 conditional novelty 6.0 of 10

    ProMoGen generates human motion conditioned on both a trajectory and sparse anchor postures via a diffusion transformer trained with a dense-to-sparse curriculum.

  12. MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm

    cs.CV 2025-02 conditional novelty 6.0 of 10

    MotionLab unifies text-based and trajectory-based motion generation with text-based editing, trajectory-based editing, motion in-betweening, and style transfer in one flow-based transformer.

  13. CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence

    cs.HC 2025-01 conditional novelty 6.0 of 10

    CARING-AI combines ChatGPT text generation, environment scanning, and smoothed text-to-motion diffusion to let authors create spatially grounded AR avatar instructions without coding or motion capture.

  14. A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A mask-guided motion correction module plus test-time adapted physics imitation restores physically plausible human motion for high-difficulty in-the-wild videos.

  15. SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SemTalk generates co-speech body motion by separating rhythm-based base gestures from semantically important sparse gestures and blending them with a learned frame-level semantic score.

  16. SCENIC: Scene-aware Semantic Navigation with Instruction-guided Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion model generates human motion that simultaneously follows text instructions and adapts to complex 3D terrain, using goal-centric canonicalization and an ego-centric distance field.

  17. Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer

    cs.CV 2024-12 conditional novelty 5.0 of 10

    This paper introduces a latent diffusion transformer that represents two-person interactive motions as one unified latent token sequence, improving text-to-motion generation quality and speed on InterHuman.

  18. KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    KptLLM++ unifies keypoint semantic understanding, visual-prompt detection, and text-prompt detection in a single multimodal LLM, reporting SOTA accuracy on COCO, AP-10K, Human-Art, and other benchmarks.

  19. From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control

    cs.RO 2025-05 reject novelty 4.0 of 10

    A new 124K-clip dataset with hierarchical text annotations, plus a pipeline that couples an LLM planner, a text-to-pose VAE, diffusion in-betweening, and physics control to generate long-horizon human behaviors.

  20. EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space

    cs.CV 2024-12 conditional novelty 4.0 of 10

    EnergyMoGen combines a latent-aware energy-based composition and a semantic-aware cross-attention energy model, fused by Synergistic Energy Fusion, to generate motions satisfying multiple textual concepts.

  21. Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.

Pith tools