Pith. sign in

REVIEW 3 cited by

Motion Mamba: Efficient and Long Sequence Motion Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07487 v4 pith:V2R6MQQX submitted 2024-03-12 cs.CV

classification cs.CV
keywords motiongenerationmambadesignefficientsequencelongmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human motion generation stands as a significant pursuit in generative computer vision, while achieving long-sequence and efficient motion generation remains challenging. Recent advancements in state space models (SSMs), notably Mamba, have showcased considerable promise in long sequence modeling with an efficient hardware-aware design, which appears to be a promising direction to build motion generation model upon it. Nevertheless, adapting SSMs to motion generation faces hurdles since the lack of a specialized design architecture to model motion sequence. To address these challenges, we propose Motion Mamba, a simple and efficient approach that presents the pioneering motion generation model utilized SSMs. Specifically, we design a Hierarchical Temporal Mamba (HTM) block to process temporal data by ensemble varying numbers of isolated SSM modules across a symmetric U-Net architecture aimed at preserving motion consistency between frames. We also design a Bidirectional Spatial Mamba (BSM) block to bidirectionally process latent poses, to enhance accurate motion generation within a temporal frame. Our proposed method achieves up to 50% FID improvement and up to 4 times faster on the HumanML3D and KIT-ML datasets compared to the previous best diffusion-based method, which demonstrates strong capabilities of high-quality long sequence motion modeling and real-time human motion generation. See project website https://steve-zeyu-zhang.github.io/MotionMamba/

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A semantically aligned latent space plus multi-token cross-attention conditioning sets a new state of the art in text-to-human-motion generation on HumanML3D.

  2. PhysiInter: Integrating Physical Mapping for High-Fidelity Human Interaction Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A text-to-motion pipeline that projects motions through physics-based imitation for training and post-processing, plus new consistency and marker-interaction losses.

  3. InterMamba: Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba

    cs.CV 2025-06 conditional novelty 6.0 of 10

    InterMamba introduces an adaptive spatial-temporal Mamba with self and cross interaction blocks for text-driven human-human interaction generation, reporting better text-motion alignment and much lower compute than InterGen.

Pith tools