Pith. sign in

REVIEW 13 cited by

GENMO: A GENeralist Model for Human MOtion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.01425 v1 pith:PMBYILRC submitted 2025-05-02 cs.GR cs.AIcs.CVcs.LGcs.RO

GENMO: A GENeralist Model for Human MOtion

classification cs.GR cs.AIcs.CVcs.LGcs.RO
keywords motiongenerationestimationgenmohumanmodelsdiversegeneralist
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes, while motion estimation models aim to reconstruct accurate motion trajectories from observations like videos. Despite sharing underlying representations of temporal dynamics and kinematics, this separation limits knowledge transfer between tasks and requires maintaining separate models. We present GENMO, a unified Generalist Model for Human Motion that bridges motion estimation and generation in a single framework. Our key insight is to reformulate motion estimation as constrained motion generation, where the output motion must precisely satisfy observed conditioning signals. Leveraging the synergy between regression and diffusion, GENMO achieves accurate global motion estimation while enabling diverse motion generation. We also introduce an estimation-guided training objective that exploits in-the-wild videos with 2D annotations and text descriptions to enhance generative diversity. Furthermore, our novel architecture handles variable-length motions and mixed multimodal conditions (text, audio, video) at different time intervals, offering flexible control. This unified approach creates synergistic benefits: generative priors improve estimated motions under challenging conditions like occlusions, while diverse video data enhances generation capabilities. Extensive experiments demonstrate GENMO's effectiveness as a generalist framework that successfully handles multiple human motion tasks within a single model.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation

    cs.AI 2026-06 unverdicted novelty 7.0

    STREAM decouples text and music conditioning in a diffusion transformer via AdaLN for structure and BEAM for beats, plus new Motorica++ dataset and editability metrics, claiming SOTA music alignment with preserved semantics.

  2. LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World

    cs.CV 2026-05 unverdicted novelty 7.0

    LAMP tracks 3D human motion from moving multi-camera headsets by converting 2D detections to a unified metric 3D world frame via device localization and fitting with an end-to-end spatio-temporal transformer.

  3. TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos

    cs.CV 2026-05 unverdicted novelty 7.0

    TT4D delivers a large-scale dataset of high-fidelity 3D table tennis gameplay reconstructed from monocular videos using a novel lift-first pipeline that infers ball trajectories and spin while handling occlusions.

  4. ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos

    cs.CV 2026-04 unverdicted novelty 7.0

    ExpertEdit edits novice motions to expert skill levels by learning a motion prior from unpaired videos and infilling masked skill-critical spans.

  5. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

    cs.CV 2025-12 unverdicted novelty 7.0

    CARI4D is the first category-agnostic pipeline that produces metric-scale, spatially and temporally consistent 4D reconstructions of human-object interactions from monocular RGB videos via foundation-model hypothesis ...

  6. Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation

    cs.AI 2026-06 conditional novelty 6.5

    STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.

  7. Video Generation Models are General-Purpose Vision Learners

    cs.CV 2026-07 conditional novelty 6.0

    A video-diffusion backbone fine-tuned as a single-step multi-task perceiver matches or beats specialists on depth, normals, pose and segmentation, with high data efficiency and sim-to-real transfer.

  8. TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes

    cs.RO 2026-06 unverdicted novelty 6.0

    TaskNPoint lets humanoid robots learn dynamic skills such as tennis backhands from single short human video demonstrations plus under one hour of single-GPU simulation training, achieving zero-shot generalization to n...

  9. MOCHI: Motion Enhancement of Collaborative Human-object Interactions

    cs.CV 2026-06 unverdicted novelty 6.0

    MOCHI enhances noisy collaborative human-object interaction captures via grasp optimization followed by diffusion-based full-body refinement that incorporates interaction information into single-person motion priors.

  10. VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

    cs.LG 2026-06 unverdicted novelty 6.0

    VideoMDM learns coherent 3D motion manifolds from 2D supervision alone by using a pretrained lifter as noisy teacher, depth-weighted 2D reprojection loss, and adapted regularizers, nearly matching fully 3D-supervised ...

  11. ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control

    cs.RO 2026-04 unverdicted novelty 6.0

    ExoActor uses exocentric video generation to implicitly model robot-environment-object interactions and converts the resulting videos into task-conditioned humanoid control sequences.

  12. Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs

    cs.CV 2026-04 unverdicted novelty 6.0

    IMU-to-4D uses wearable IMU data and repurposed LLMs to predict coherent 4D human motion plus coarse scene structure, outperforming cascaded state-of-the-art pipelines in temporal stability.

  13. FactorizedHMR: A Hybrid Framework for Video Human Mesh Recovery

    cs.CV 2026-05 unverdicted novelty 5.0

    FactorizedHMR recovers 3D human meshes from video by deterministically anchoring the torso-root then probabilistically completing distal articulations via flow-matching with geometry-aware supervision and a synthetic ...