Pith. sign in

REVIEW 8 cited by

MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20340 v1 pith:JAJZYIRN submitted 2024-05-30 cs.CV

classification cs.CV
keywords humanunderstandingmotionmotionllmbehaviorvideosdatallms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study delves into the realm of multi-modality (i.e., video and motion modalities) human behavior understanding by leveraging the powerful capabilities of Large Language Models (LLMs). Diverging from recent LLMs designed for video-only or motion-only understanding, we argue that understanding human behavior necessitates joint modeling from both videos and motion sequences (e.g., SMPL sequences) to capture nuanced body part dynamics and semantics effectively. In light of this, we present MotionLLM, a straightforward yet effective framework for human motion understanding, captioning, and reasoning. Specifically, MotionLLM adopts a unified video-motion training strategy that leverages the complementary advantages of existing coarse video-text data and fine-grained motion-text data to glean rich spatial-temporal insights. Furthermore, we collect a substantial dataset, MoVid, comprising diverse videos, motions, captions, and instructions. Additionally, we propose the MoVid-Bench, with carefully manual annotations, for better evaluation of human behavior understanding on video and motion. Extensive experiments show the superiority of MotionLLM in the caption, spatial-temporal comprehension, and reasoning ability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation

    cs.CV 2026-03 conditional novelty 7.0 of 10

    UniMotion unifies continuous human-motion, text, and RGB understanding/generation/editing in one LLM backbone via CMA-VAE, Dual-Posterior KL Alignment, and Latent Reconstruction Alignment, reporting SOTA on seven tri-...

  2. Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A single MLLM trained with a vision-guided hybrid VQ-VAE tokenizer reports state-of-the-art or competitive results for 3D pose estimation, motion prediction, and motion in-betweening on Human3.6M and 3DPW.

  3. HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical benchmark for multimodal models on human-centric visual understanding finds frontier models average under 60% and miss question-uncued visual evidence, with test-time scaling helping only marginally.

  4. Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Being-M0.5 combines part-aware residual quantization with a 5M-sequence web-video dataset to reach real-time, part-controllable 3D motion generation, though its state-of-the-art claim does not hold on every standard b...

  5. Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dexterous VLA pretrained on a 2.5M-instance human hand motion dataset transfers skills to a real robot hand, outperforming baselines in manipulation tasks.

  6. AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 2B-parameter video-language model using an RWKV linear-RNN backbone and sorted token merging achieves competitive long-video QA accuracy with far lower memory cost than transformer-based models.

  7. Hierarchical Motion Captioning Utilizing External Text Data Source

    cs.LG 2025-09 conditional novelty 5.0 of 10

    This paper introduces a hierarchical motion captioning system that generates low-level descriptions with an LLM and retrieves high-level captions from a database, reporting large gains over prior methods on three datasets.

  8. KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    KptLLM++ unifies keypoint semantic understanding, visual-prompt detection, and text-prompt detection in a single multimodal LLM, reporting SOTA accuracy on COCO, AP-10K, Human-Art, and other benchmarks.

Pith tools