Pith. sign in

REVIEW 4 major objections 5 minor 11 cited by

Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read An automatic annotation pipeline turns 80.8K internet videos into 19.5M 3D whole-body pose annotations with text and audio, and training on the data improves motion generation and recovery.

desk verdict A genuinely large and useful dataset expansion, but the core accuracy claim rests on an unvalidated scale-recovery step and the excerpt lacks independent verification. read the letter →

arxiv 2501.05098 v1 pith:GAJ57S6R submitted 2025-01-09 cs.CV

classification cs.CV
keywords 3Dwhole-bodymotionhumandatasetSMPL-Xtext-drivengenerationaudio-drivenmeshrecoveryposeestimationmultimodalannotationpipeline
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that an automatic annotation pipeline can convert unconstrained internet RGB videos into expressive 3D whole-body human motion with paired text and audio, at a scale not possible with manual labeling. The result is Motion-X++, with 19.5M 3D whole-body pose annotations across 120.5K sequences, 80.8K RGB videos, 45.3K audios, 19.5M frame-level pose descriptions, and 120.5K sequence-level semantic labels. The authors argue that this scale and multimodality improve text-driven and audio-driven motion generation, whole-body mesh recovery, and 2D keypoint estimation. If the pipeline's annotations are accurate enough to serve as ground truth, the dataset removes a key bottleneck in human-motion learning.

What carries the argument

The central object is the annotation pipeline, whose load-bearing stages are: whole-body keypoint estimation; camera tracking via dense bundle adjustment with the human masked out of feature extraction, correspondence fields, and flow; global trajectory optimization with reprojection, smoothness, ground-contact, and foot-skating losses; and text synthesis that turns estimated body-part spatial relations, hand shapes, and classifier-based emotions into frame-level descriptions, with a vision-language model providing sequence-level captions. All motion is represented in the SMPL-X whole-body parametric model, so body, hands, and face share one optimization. The pipeline's role is to convert raw RGB video into self-supervised training labels at scale.

What would settle it

Take a random sample of roughly 500 Motion-X++ clips, have annotators manually fit the whole-body model to the frames (or capture a subset with multi-view or motion-capture systems), and compare per-joint errors on hands, face, and global trajectory; if the median errors approach the typical inter-annotator variation or exceed the tolerances used in the paper's evaluation, the ground-truth claim is falsified.

Watch

Extended reading notes

Core claim

Motion-X++ asserts that every part of a human motion, including body pose, hand gestures, facial expressions, and the global trajectory, can be recovered from ordinary single-view or multi-view video by a staged pipeline, and that the recovered parameters are accurate enough to train and evaluate downstream models. The pipeline estimates whole-body 2D keypoints, tracks the camera with a masked dense-bundle-adjustment SLAM that excludes the moving person, optimizes the global trajectory with reprojection, smoothness, ground-contact, and foot-skating losses, fits SMPL-X parameters to the observations, and then generates frame-level text descriptions from body-part relations, hand gestures, and an emotion classifier, while a vision-language model supplies sequence-level captions. The paper presents the resulting dataset as ground truth and reports that training on it improves the four downstream task families.

Load-bearing premise

The load-bearing premise is that the automatic annotation pipeline, including SMPL-X fitting, masked dense-bundle-adjustment camera tracking, trajectory optimization, and vision-language captioning, produces 3D poses and text descriptions accurate enough to serve as ground truth for unconstrained internet videos.

Editorial extensions

If this is right

  • Text-driven whole-body motion generation can be trained on 120.5K sequence-level captions paired with expressive 3D pose, including hands and face.
  • Audio-driven motion generation gains 45.3K audio–motion pairs, enabling models that synthesize dance or performance from music.
  • Whole-body mesh recovery and 2D keypoint estimation can be trained on diverse internet scenes rather than lab captures.
  • Because the annotation pipeline runs on any RGB video collection, the dataset can keep growing without manual text labeling.
  • Frame-level whole-body pose descriptions give models dense language supervision instead of only sequence-level captions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the accuracy claim holds, Motion-X++ could also serve as pretraining data for video-to-motion and motion-to-video generation, since raw video and audio are stored alongside the motion.
  • The vision-language captions are likely noisier than the geometric pose labels, so a fair text-to-motion benchmark should include human judgments of caption–pose alignment before treating the captions as ground truth.
  • The strategy of masking the moving person out of dense-bundle-adjustment camera tracking could transfer to other dynamic-scene reconstruction problems, such as tracking cameras in crowds or sports footage.
  • Because audio, video, pose, and text are aligned for the same sequences, the dataset opens cross-modal alignment research that motion-only datasets cannot support.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Motion-X++, a large-scale multimodal dataset of 3D whole-body human motion, providing 19.5M pose annotations over 120.5K sequences, along with RGB video, audio, frame-level pose descriptions, and sequence-level semantic labels. The core contribution is a scalable automatic annotation pipeline that combines masked DROID-SLAM camera tracking, SMPL-X whole-body fitting, global trajectory optimization with foot-contact constraints, and GPT-4V-generated captions. The paper claims that training on Motion-X++ improves text-driven motion generation, audio-driven motion generation, 3D whole-body mesh recovery, and 2D whole-body keypoint estimation.

Significance. If substantiated, Motion-X++ would be a valuable community resource: it is substantially larger and more modality-rich than existing 3D whole-body motion datasets, and the pipeline explicitly targets known failure modes such as hand gestures and facial expressions. The manuscript deserves credit for providing concrete mathematical formulations of the optimization stages (Eqs. 6-14) and for acknowledging limitations of its predecessor, Motion-X. However, the load-bearing claim that the automatic pipeline yields metric, expressive 3D poses suitable as training and evaluation ground truth for unconstrained internet videos is not yet supported by visible evidence in the reviewed excerpt.

major comments (4)
  1. [§Human Trajectory Refinement, Eqs. (9)-(12)] The monocular camera/body scale ambiguity is not resolved or validated. Because the dataset is built from single-view internet videos, absolute scene scale is unobservable from the images alone. The optimization in Eq. (12) fixes camera scale α and SMPL-X shape β from the previous stage and optimizes only Φ_t and Γ_t, so the resulting root translations Γ_t inherit whatever scale is encoded in α. The paper does not show an experiment comparing recovered global trajectories against metric ground truth, nor does it quantify the effect of the scale choice on downstream pose accuracy. This is load-bearing because the 'accurate 3D whole-body pose annotation' claim includes global motion, not just relative pose.
  2. [Abstract and §1 (pipeline overview)] The central claim that 'Comprehensive experiments validate the accuracy of our annotation pipeline' is not supported by any quantitative result in the reviewed excerpt. There are no error bars, no comparisons against independent 3D ground truth, no human perceptual evaluation, and no failure-mode analysis. Because the paper's value proposition is annotation precision, the reader cannot assess whether the large scale is achieved at the cost of systematic errors in hands, faces, and global trajectories. Please add per-component accuracy evaluations and benchmark comparisons, or clearly state which experiments are deferred.
  3. [Downstream task evaluation and GPT-4V captions] The excerpt does not specify the evaluation protocol for the downstream tasks, so it is unclear whether test annotations come from the proposed pipeline itself or from independent human/external ground truth. If the same pipeline outputs are used as both training signal and evaluation target, the reported gains could reflect consistency with a biased annotation process rather than accuracy. Similarly, GPT-4V captions appear to be used as both the annotation and the evaluation target for caption quality. Please clarify the evaluation protocols and include at least one cross-dataset or human-validated evaluation for both 3D pose accuracy and text annotation quality.
  4. [Fig. 7 and single-view inference] Fig. 7 presents a multi-view annotation pipeline and the text states that the pipeline supports any number of viewpoints, but the dataset is constructed from single-view internet videos. The mechanism by which absolute scale is recovered in the single-view setting is therefore critical and under-specified. Concretely, the text should explain how α in Eq. (10) is computed and why it is metrically reliable for single-view inputs, or provide a validation experiment on sequences with known metric scale.
minor comments (5)
  1. [Eq. (9)] The parentheses do not balance: 'Jt = M (Φt, Θt, β) + Γt),' should be 'Jt = M(Φt, Θt, β) + Γt'.
  2. [Eq. (11)] There is an unmatched closing parenthesis in '||J I t − J I t+1)||2'; please correct the notation.
  3. [Fig. 7] The figure is numbered 'Fig. 7' but its caption reads 'Figure 1. Muti-view Annotation Pipeline'; the numbering and the typo 'Muti' should be fixed.
  4. [Eq. (10)] The symbol α is used in the reprojection loss before it is defined. Please introduce α explicitly, state its units, and explain how it is derived from the previous stage.
  5. [Masked DROID-SLAM description] The list of two key differences in the masked DROID-SLAM strategy is written as one paragraph with '1)We' missing a space; please reformat for readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: dataset construction pipeline is sequential and no prediction reduces to its inputs.

full rationale

The paper is a dataset-construction paper; it does not claim a first-principles derivation whose conclusion is equivalent to its premises. The core pipeline (masked DROID-SLAM camera tracking, SMPL-X fitting, trajectory optimization with Eqs. 6-14, and GPT-4V captioning) takes raw RGB video and detected 2D keypoints as inputs and produces 3D pose, trajectory, and text labels; no equation defines an output in terms of the same output. The monocular scale-depth ambiguity flagged in the skeptic attack is a real validation risk, but it is an identifiability/correctness issue rather than a circular reduction, because camera scale alpha is carried forward from an earlier stage, not derived from the trajectory it is used to refine. The manuscript's own limitation statement about Motion-X's 'inaccurate hand gestures, collapsed facial...' is an acknowledged failure mode of the prior pipeline, not a self-referential justification. The references to prior work such as SLAHMR and DROID-SLAM are standard external methods, and the reference to Motion-X is dataset lineage rather than load-bearing evidence. No quoted passage in the provided text shows downstream evaluation using the same pipeline's annotations as both training signal and evaluation target; without that specific reduction, no circularity can be scored.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

These are the main unverified premises the dataset's value rests on: the fidelity of SMPL-X, the reliability of masked DROID-SLAM in dynamic scenes, the accuracy of GPT-4V captions and rule-based pose descriptions, and the consistency of merged action datasets. The paper describes each component but does not, in the reviewed excerpt, supply independent error analysis for any of them.

free parameters (1)
  • Trajectory optimization hyperparameters = not reported
    Equations 12-14 introduce loss weights (lambda_data_G, lambda_smooth_G, lambda_skate, lambda_con) and a ground-contact threshold delta; these hand-chosen values control how the estimated whole-body motion is smoothed and grounded, directly affecting annotation quality.
assumptions (4)
  • domain assumption SMPL-X is a faithful whole-body body model, including hands and face.
    The entire annotation pipeline fits SMPL-X parameters; if SMPL-X's hand and face model is insufficient, the annotated gestures and expressions are biased.
  • domain assumption Masked DROID-SLAM yields reliable camera pose and depth in dynamic scenes.
    Equations 6-8 assume that dense bundle adjustment after masking moving humans gives correct camera motion, which is hard to guarantee when humans occupy large image areas.
  • domain assumption GPT-4V captions and rule-based pose descriptions are semantically aligned with the motion.
    The abstract and Fig. 1 state that text labels are produced by GPT-4V and rule-based scripts; no human verification rate is reported in the excerpt.
  • domain assumption Merging eight existing action datasets preserves label consistency.
    The pipeline incorporates data from eight existing action datasets; label semantics, capture conditions, and annotation conventions may differ across sources.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset." pith.science (2026). https://pith.science/paper/GAJ57S6R

@misc{pith2026250105098,
  author       = {Pith},
  title        = {Pith review of: Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GAJ57S6R}},
  note         = {Machine review of arXiv:2501.05098}
}
read the original abstract

In this paper, we introduce Motion-X++, a large-scale multimodal 3D expressive whole-body human motion dataset. Existing motion datasets predominantly capture body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions, and are typically limited to lab settings with manually labeled text descriptions, thereby restricting their scalability. To address this issue, we develop a scalable annotation pipeline that can automatically capture 3D whole-body human motion and comprehensive textural labels from RGB videos and build the Motion-X dataset comprising 81.1K text-motion pairs. Furthermore, we extend Motion-X into Motion-X++ by improving the annotation pipeline, introducing more data modalities, and scaling up the data quantities. Motion-X++ provides 19.5M 3D whole-body pose annotations covering 120.5K motion sequences from massive scenes, 80.8K RGB videos, 45.3K audios, 19.5M frame-level whole-body pose descriptions, and 120.5K sequence-level semantic labels. Comprehensive experiments validate the accuracy of our annotation pipeline and highlight Motion-X++'s significant benefits for generating expressive, precise, and natural motion with paired multimodal labels supporting several downstream tasks, including text-driven whole-body motion generation,audio-driven motion generation, 3D whole-body human mesh recovery, and 2D whole-body keypoints estimation, etc.

Figures

Figures reproduced from arXiv: 2501.05098 by the authors.

Figure 1
Figure 1. Compared to Motion-X, our enhanced dataset Motion-X++ offers (a) more precise human motion, including robust facial expressions and refined hand gestures. Facial expressions and hand gestures are highlighted. Additionally, Motion-X++ provides a broader range of modalities, such as audio and video, and improved quality in text (annotated by GPT-4V) and motion (refined annotation pipeline). The expanded modalities ena… view at source ↗
Figure 2
Figure 2. Diversity statistics of the face, hand, and body motions in Motion-X++. Multimodality. Compared to existing text-only or audio￾only motion datasets, Motion-X++ offers multiple modalities, including 120.5K videos with annotated motion sequences, 45.3K audio samples, 19.5M frame-level whole-body pose descriptions, and 120.5K sequence-level semantic labels. This diverse compilation, sourced from multiple data platforms… view at source ↗
Figure 5
Figure 5. Examples (a), (b), and (c) illustrate multi-shot scenarios [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗
Figures from the paper (10 more)
Figure 3
Figure 3. Figure 3: Illustration of the overall data collection and annotation pipeline. Play Clarinet Play Gaohu Play Jinghu Play Harp Play Guitar Aerial work Afraid of height Arm wrestling Balance beam Apply cream Playing drum handstand Tai Chi Play Guqin CPR Bmx riding skate forward br…
Figure 4
Figure 4. Figure 4: Overview of Motion-X++. It includes (a) diverse facial expressions extracted from BAUM [38], (b) indoor motion with expressive face and hand motions, (c) outdoor motion with diverse and challenging poses, and (d) several motion sequences. Purple SMPL-X is the observed …
Figure 6
Figure 6. Figure 6: The automatic pipeline for the whole-body motion capture from massive multi-shot videos. It comprises shot detection, 2D and 3D whole-body keypoints estimation, local pose estimation, and global trajectory optimization stages. This pipeline is designed to support both …
Figure 7
Figure 7. Figure 7: Our annotation pipeline can support input from any number of viewpoints. By simultaneously optimizing a fixed set [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 9
Figure 9. Figure 9: Frame distribution across the entire Game Motion Dataset using various clip method combinations. [30_80] means the frame length ranges from 30 to 80. The combined shot detection method (SceneDetect+Tracking+optical Flow) closely matches the ground truth (GT) distributi…
Figure 8
Figure 8. Figure 8: Qualitative comparisons of (a) 2D keypoints annotation with widely used methods [92, 93] and (b) the 3D mesh annotation with the popular fitting method [32] with ours. hand ↑ face ↑ whole-body ↑ Method AP AR AP AR AP AR OpenPose [92] 38.6 43.3 76.5 84.0 44.2 52.3 HRNet…
Figure 11
Figure 11. Figure 11: Global Camera Trajectory Comparison on EMDB: Our method yields more accurate global camera trajectories compared to DROID-SLAM, WHAM, and TRAM [PITH_FULL_IMAGE:figures/full_fig_p010_11.png]
Figure 10
Figure 10. Figure 10: Camera trajectory comparison between original Droid￾SLAM and our method on videos with static camera pose. Unlike the original DROID-SLAM, whose output dots (blue) indicate it as a moving camera, the masked DROID-SLAM (orange) is represented by a single point and is p…
Figure 13
Figure 13. Figure 13: Visual comparisons of motions generated by MLD [2] trained on HumanML3D (in purple) or Motion-X++ (in blue). Please zoom in for a detailed comparison. The model trained with Motion-X++ can generate more accurate and semantic￾corresponded motions. Comparison with Human…
Figure 14
Figure 14. Figure 14: Visualized results of 2D whole-body pose estimation on the COCO-Wholebody [69] dataset. Please zoom in for details. incorporating an additional 10% of the single-view data sampled from Motion-X++ while keeping the other setting the same. As shown in Tab. 9, the model …

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

    cs.RO 2026-08 conditional novelty 6.0 of 10

    HumanTracker introduces a 153-hour categorized humanoid tracking benchmark and a preference-trained metric, HumanScore, that agrees with human judgments better than kinematic error metrics.

  2. MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

    cs.CV 2026-08 conditional novelty 6.0 of 10

    MRBench is a multi-source, balanced, multi-granular human motion-text retrieval benchmark, and the proposed granularity-aware adapters improve mixed-granularity retrieval without degrading standard-caption retrieval.

  3. $\omega$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A single whole-body model with latent future prediction outperforms prior robot policies on 11 real-world humanoid household loco-manipulation tasks.

  4. UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A visual proxy that renders human motion under the driving camera and overlays camera trajectory markers lets a video diffusion model control both body motion and camera movement from a single visual conditioning space.

  5. EgoHTR: Egocentric 4D Demonstrations of Human Terrain Traversal

    cs.RO 2026-07 conditional novelty 6.0 of 10

    EgoHTR is a 55-sequence, 150k-frame egocentric 4D human-terrain dataset with a reconstruction pipeline, MoCap-validated benchmark, and perceptive locomotion policies deployed on a Unitree G1.

  6. GUSH3R: Everyone Everywhere All at Once as Gaussians

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A feed-forward framework predicts unified 3D Gaussian representations of dynamic humans and static scenes from monocular video in a single forward pass.

  7. Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation

    cs.CV 2026-01 conditional novelty 6.0 of 10

    RAM couples motion reconstruction with text-to-motion diffusion and adds reconstruction-anchored error guidance, reporting FID 0.032 on HumanML3D with 20 inference steps.

  8. Distinguishing Imitation Error from Intrinsic Motion Learning Difficulty

    cs.GR 2025-12 conditional novelty 6.0 of 10

    A physics-based score (MDS) predicts how hard a motion is for a humanoid to imitate by measuring how much joint torques must change under small pose perturbations.

  9. PHUMA: Physically Reliable Humanoid Locomotion Dataset

    cs.RO 2025-10 conditional novelty 6.0 of 10

    PHUMA is a curated 73-hour humanoid locomotion corpus whose physical-reliability metrics are partly defined by the same losses used to optimize it, and whose imitation success claims are confounded by in-distribution ...

  10. The loss tolerance of cat breeding for fault-tolerant grid state generation

    quant-ph 2025-08 reject novelty 6.0 of 10

    Claims a 4% optical-loss ceiling for fault-tolerant GKP state generation via cat breeding, but the provided full text is an unrelated manuscript with no such analysis.

  11. VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A unified multimodal motion LLM that handles nine generation and comprehension tasks across text, audio, and single/multi-agent motion, backed by a new dataset and tokenizer.

Reference graph

Works this paper leans on

103 extracted references · 66 canonical work pages · cited by 11 Pith papers

  1. [1]

    'E4nSt C777dYs3 Pi^w.x j իW?k׮ ].W Isssuuu^WQ ED ^н]@ /^

    11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEcompsocconfonly \@IEEEauthordefaulttextstyle \@IEEEcompsocnotconfonly \@IEEEauthor...

  2. [2]

    Ahuja and L.-P

    C. Ahuja and L.-P. Morency, ``Language2pose: Natural language grounded pose forecasting,'' in 3DV, 2019

  3. [3]

    X. Chen, B. Jiang, W. Liu, Z. Huang, B. Fu, T. Chen, J. Yu, and G. Yu, ``Executing your commands via motion diffusion in latent space,'' in CVPR, 2023

  4. [4]

    Delmas, P

    G. Delmas, P. Weinzaepfel, T. Lucas, F. Moreno-Noguer, and G. Rogez, ``Posescript: 3d human poses from natural language,'' in ECCV, 2022

  5. [5]

    C. Guo, S. Zou, X. Zuo, S. Wang, W. Ji, X. Li, and L. Cheng, ``Generating diverse and natural 3d human motions from text,'' in CVPR, 2022

  6. [6]

    Petrovich, M

    M. Petrovich, M. J. Black, and G. Varol, ``Temos: Generating diverse human motions from textual descriptions,'' in ECCV, 2022

  7. [7]

    Plappert, C

    M. Plappert, C. Mandery, and T. Asfour, ``The kit motion-language dataset,'' Big data, 2016

  8. [8]

    Plappert, C

    M. Plappert, C. Mandery, and T.Asfour, ``Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks,'' Robotics and Autonomous Systems, 2018

Show all 103 references
  1. [9]

    A. R. Punnakkal, A. Chandrasekaran, N. Athanasiou, A. Quiros-Ramirez, and M. J. Black, ``Babel: bodies, action and behavior with english labels,'' in CVPR, 2021

  2. [10]

    Zhang, Z

    M. Zhang, Z. Cai, L. Pan, F. Hong, X. Guo, L. Yang, and Z. Liu, ``Motiondiffuse: Text-driven human motion generation with diffusion model,'' arXiv preprint arXiv:2208.15001, 2022

  3. [11]

    M. Zhao, M. Liu, B. Ren, S. Dai, and N. Sebe, ``Modiff: Action-conditioned 3d motion generation with denoising diffusion probabilistic models,'' arXiv preprint arXiv:2301.03949, 2023

  4. [12]

    L.-H. Chen, S. Lu, A. Zeng, H. Zhang, B. Wang, R. Zhang, and L. Zhang, ``Motionllm: Understanding human behaviors from human motions and videos,'' arXiv preprint arXiv:2405.20340, 2024

  5. [13]

    F. Hong, L. Pan, Z. Cai, and Z. Liu, ``Versatile multi-modal pre-training for human-centric perception,'' arXiv preprint arXiv:2203.13815, 2022

  6. [14]

    Jiang, X

    B. Jiang, X. Chen, W. Liu, J. Yu, G. Yu, and T. Chen, ``Motiongpt: Human motion as a foreign language,'' Advances in Neural Information Processing Systems, vol. 36, 2024

  7. [15]

    Z. Zhou, Y. Wan, and B. Wang, ``Avatargpt: All-in-one framework for motion understanding, planning, generation and beyond,'' 2023. [Online]. Available: https://arxiv.org/abs/2311.16468

  8. [16]

    J. Wang, Y. Rong, J. Liu, S. Yan, D. Lin, and B. Dai, ``Towards diverse and natural scene-aware 3d human motion synthesis,'' 2022. [Online]. Available: https://arxiv.org/abs/2205.13001

  9. [17]

    J. Wang, Y. Yuan, Z. Luo, K. Xie, D. Lin, U. Iqbal, S. Fidler, and S. Khamis, ``Learning human dynamics in autonomous driving scenarios,'' in 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 1em plus 0.5em minus 0.4em Los Alamitos, CA, USA: IEEE Computer Socie...

  10. [18]

    Z. Xiao, T. Wang, J. Wang, J. Cao, W. Zhang, B. Dai, D. Lin, and J. Pang, ``Unified human-scene interaction via prompted chain-of-contacts,'' in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=1vCnDyQkjg

  11. [19]

    Mahmood, N

    N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, ``Amass: Archive of motion capture as surface shapes,'' in ICCV, 2019

  12. [20]

    J. Ho, A. Jain, and P. Abbeel, ``Denoising diffusion probabilistic models,'' NeurIPS, 2020

  13. [21]

    J. Song, C. Meng, and S. Ermon, ``Denoising diffusion implicit models,'' in ICLR, 2020

  14. [22]

    Tevet, S

    G. Tevet, S. Raab, B. Gordon, Y. Shafir, D. Cohen-Or, and A. H. Bermano, ``Human motion diffusion model,'' in ICLR, 2023

  15. [23]

    J. Lin, A. Zeng, H. Wang, L. Zhang, and Y. Li, ``One-stage 3d whole-body mesh recovery with component aware transformer,'' in CVPR, 2023

  16. [24]

    R. Li, S. Yang, D. A. Ross, and A. Kanazawa, ``Ai choreographer: Music conditioned 3d dance generation with aist++,'' in ICCV, 2021

  17. [25]

    G. Moon, H. Choi, and K. M. Lee, ``Accurate 3d hand pose estimation for whole-body 3d human mesh estimation,'' in CVPRW, 2020

  18. [26]

    Y. Xu, J. Zhang, Q. Zhang, and D. Tao, ``Vitpose: Simple vision transformer baselines for human pose estimation,'' in NeurIPS, 2022

  19. [27]

    Y. Yuan, U. Iqbal, P. Molchanov, K. Kitani, and J. Kautz, ``Glamr: Global occlusion-aware human mesh recovery with dynamic cameras,'' in CVPR, 2022

  20. [28]

    J. Yang, A. Zeng, S. Liu, F. Li, R. Zhang, and L. Zhang, ``Explicit box detection unifies end-to-end multi-person pose estimation,'' in ICLR, 2023

  21. [29]

    H. E. Pang, Z. Cai, L. Yang, T. Zhang, and Z. Liu, ``Benchmarking and analyzing 3d human pose and shape estimation beyond algorithms,'' in NeurIPS Datasets and Benchmarks Track, 2022

  22. [30]

    G. Moon, H. Choi, and K. M. Lee, ``Neuralannot: Neural annotator for 3d human mesh training sets,'' in CVPR, 2022

  23. [31]

    G. Moon, H. Choi, S. Chun, J. Lee, and S. Yun, ``Three recipes for better 3d pseudo-gts of 3d human mesh estimation in the wild,'' in CVPR, 2023

  24. [32]

    H. Yi, H. Liang, Y. Liu, Q. Cao, Y. Wen, T. Bolkart, D. Tao, and M. J. Black, ``Generating holistic 3d human motion from speech,'' in CVPR, 2023

  25. [33]

    Pavlakos, V

    G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, ``Expressive body capture: 3d hands, face, and body from a single image,'' in CVPR, 2019

  26. [34]

    Z. Cai, D. Ren, A. Zeng, Z. Lin, T. Yu, W. Wang, X. Fan, Y. Gao, Y. Yu, L. Pan, F. Hong, M. Zhang, C. C. Loy, L. Yang, and Z. Liu, ``Humman: Multi-modal 4d human dataset for versatile sensing and modeling,'' in ECCV, 2022

  27. [35]

    Chung, C.-h

    J. Chung, C.-h. Wuu, H.-r. Yang, Y.-W. Tai, and C.-K. Tang, ``Haa500: Human-centric atomic action dataset with curated videos,'' in ICCV, 2021

  28. [36]

    J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, ``Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,'' in TPAMI, 2019

  29. [37]

    Taheri, N

    O. Taheri, N. Ghorbani, M. J. Black, and D. Tzionas, ``Grab: A dataset of whole-body human grasping of objects,'' in ECCV, 2020

  30. [38]

    Tsuchida, S

    S. Tsuchida, S. Fukayama, M. Hamasaki, and M. Goto, ``Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing.'' in ISMIR, 2019

  31. [39]

    Zhalehpour, O

    S. Zhalehpour, O. Onder, Z. Akhtar, and C. E. Erdem, ``Baum-1: A spontaneous audio-visual face database of affective and mental states,'' IEEE Transactions on Affective Computing, 2016

  32. [40]

    Zhang, Q

    S. Zhang, Q. Ma, Y. Zhang, Z. Qian, T. Kwon, M. Pollefeys, F. Bogo, and S. Tang, ``Egobody: Human body shape and motion of interacting people from head-mounted devices,'' in ECCV, 2022

  33. [41]

    J. Lin, A. Zeng, S. Lu, Y. Cai, R. Zhang, H. Wang, and L. Zhang, ``Motion-x: A large-scale 3d expressive whole-body human motion dataset,'' Advances in Neural Information Processing Systems, 2023

  34. [42]

    Dan e c ek, M

    R. Dan e c ek, M. J. Black, and T. Bolkart, ``Emoca: Emotion driven monocular face capture and animation,'' in CVPR, 2022

  35. [43]

    Pavlakos, D

    G. Pavlakos, D. Shan, I. Radosavovic, A. Kanazawa, D. Fouhey, and J. Malik, ``Reconstructing hands in 3d with transformers,'' in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 9826--9836

  36. [44]

    Z. Cai, W. Yin, A. Zeng, C. Wei, Q. Sun, W. Yanjun, H. E. Pang, H. Mei, M. Zhang, L. Zhang et al., ``Smpler-x: Scaling up expressive human pose and shape estimation,'' Advances in Neural Information Processing Systems, vol. 36, 2024

  37. [45]

    V. Ye, G. Pavlakos, J. Malik, and A. Kanazawa, ``Decoupling human and camera motion from videos in the wild,'' in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2023

  38. [46]

    Carreira, E

    J. Carreira, E. Noland, C. Hillier, and A. Zisserman, ``A short note on the kinetics-700 human action dataset,'' arXiv preprint arXiv:1907.06987, 2019

  39. [47]

    C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar et al., ``Ava: A video dataset of spatio-temporally localized atomic visual actions,'' in CVPR, 2018

  40. [48]

    Shahroudy, J

    A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, ``Ntu rgb+ d: A large scale dataset for 3d human activity analysis,'' in CVPR, 2016

  41. [49]

    Trivedi, A

    N. Trivedi, A. Thatipelli, and R. K. Sarvadevabhatla, ``Ntu-x: an enhanced large-scale dataset for improving pose-based recognition of subtle human actions,'' in ICVGIP, 2021

  42. [50]

    Hassan, D

    M. Hassan, D. Ceylan, R. Villegas, J. Saito, J. Yang, Y. Zhou, and M. J. Black, ``Stochastic scene-aware motion prediction,'' in ICCV, 2021

  43. [51]

    Hassan, V

    M. Hassan, V. Choutas, D. Tzionas, and M. J. Black, ``Resolving 3d human pose ambiguities with 3d scene constraints,'' in ICCV, 2019

  44. [52]

    Y.-L. Li, X. Liu, X. Wu, Y. Li, Z. Qiu, L. Xu, Y. Xu, H.-S. Fang, and C. Lu, ``Hake: a knowledge engine foundation for human activity understanding,'' in TPAMI, 2022

  45. [53]

    Zheng, Y

    Y. Zheng, Y. Yang, K. Mo, J. Li, T. Yu, Y. Liu, C. K. Liu, and L. J. Guibas, ``Gimo: Gaze-informed human motion prediction in context,'' in ECCV, 2022

  46. [54]

    C. Guo, X. Zuo, S. Wang, S. Zou, Q. Sun, A. Deng, M. Gong, and L. Cheng, ``Action2motion: Conditioned generation of 3d human motions,'' in ACM MM, 2020

  47. [55]

    Gross and J

    R. Gross and J. Shi, ``The cmu motion of body (mobo) database,'' 2001

  48. [56]

    Ionescu, D

    C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, `` Human3.6M : Large scale datasets and predictive methods for 3d human sensing in natural environments,'' in TPAMI, 2014

  49. [57]

    Sigal, A

    L. Sigal, A. O. Balan, and M. J. Black, ``Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion,'' IJCV, 2010

  50. [58]

    Trumble, A

    M. Trumble, A. Gilbert, C. Malleson, A. Hilton, and J. P. Collomosse, ``Total capture: 3d human pose estimation fusing video and inertial sensors,'' in BMVC, 2017

  51. [59]

    Loper, N

    M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, ``Smpl: A skinned multi-person linear model,'' ACM TOG, 2015

  52. [60]

    Y. Yuan, J. Song, U. Iqbal, A. Vahdat, and J. Kautz, ``Physdiff: Physics-guided human motion diffusion model,'' in ICCV, 2023

  53. [61]

    Zhang, Y

    J. Zhang, Y. Zhang, X. Cun, S. Huang, Y. Zhang, H. Zhao, H. Lu, and X. Shen, ``T2m-gpt: Generating human motion from textual descriptions with discrete representations,'' in CVPR, 2023

  54. [62]

    Zhuang, C

    W. Zhuang, C. Wang, J. Chai, Y. Wang, M. Shao, and S. Xia, ``Music2dance: Dancenet for music-driven dance generation,'' ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 18, no. 2, pp. 1--21, 2022

  55. [63]

    R. Li, J. Zhao, Y. Zhang, M. Su, Z. Ren, H. Zhang, Y. Tang, and X. Li, ``Finedance: A fine-grained choreography dataset for 3d full body dance generation,'' in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 10\,234--10\,243

  56. [64]

    W. Jiao, W. Wang, J.-t. Huang, X. Wang, and Z. Tu, ``Is chatgpt a good translator? a preliminary study,'' arXiv preprint arXiv:2301.08745, 2023

  57. [65]

    Zablotskaia, A

    P. Zablotskaia, A. Siarohin, B. Zhao, and L. Sigal, ``Dwnet: Dense warp-based network for pose-guided human video generation,'' arXiv preprint arXiv:1910.09139, 2019

  58. [66]

    Huang, Y

    Q. Huang, Y. Xiong, A. Rao, J. Wang, and D. Lin, ``Movienet: A holistic dataset for movie understanding,'' in The European Conference on Computer Vision (ECCV), 2020

  59. [67]

    Contributors, `` MMTracking: OpenMMLab video perception toolbox and benchmark,'' https://github.com/open-mmlab/mmtracking, 2020

    M. Contributors, `` MMTracking: OpenMMLab video perception toolbox and benchmark,'' https://github.com/open-mmlab/mmtracking, 2020

  60. [68]

    R. E. K \'a lm \'a n and R. S. Bucy, ``New results in linear filtering and prediction theory,'' Journal of Basic Engineering, vol. 83, pp. 95--108, 1961. [Online]. Available: https://api.semanticscholar.org/CorpusID:8141345

  61. [69]

    Teed and J

    Z. Teed and J. Deng, ``Raft: Recurrent all-pairs field transforms for optical flow,'' 2020. [Online]. Available: https://arxiv.org/abs/2003.12039

  62. [70]

    S. Jin, L. Xu, J. Xu, C. Wang, W. Liu, C. Qian, W. Ouyang, and P. Luo, ``Whole-body human pose estimation in the wild,'' in ECCV, 2020

  63. [71]

    L. Xu, S. Jin, W. Liu, C. Qian, W. Ouyang, P. Luo, and X. Wang, ``Zoomnas: searching for whole-body human pose estimation in the wild,'' TPAMI, 2022

  64. [72]

    Narasimhaswamy, T

    S. Narasimhaswamy, T. Nguyen, M. Huang, and M. Hoai, ``Whose hands are these? hand detection and hand-body association in the wild,'' in CVPR, 2022

  65. [73]

    Savitzky and M

    A. Savitzky and M. J. Golay, ``Smoothing and differentiation of data by simplified least squares procedures.'' Analytical chemistry, 1964

  66. [74]

    A. Zeng, X. Ju, L. Yang, R. Gao, X. Zhu, B. Dai, and Q. Xu, ``Deciwatch: A simple baseline for 10x efficient 2d and 3d pose estimation,'' in ECCV, 2022

  67. [75]

    A. Zeng, L. Yang, X. Ju, J. Li, J. Wang, and Q. Xu, ``Smoothnet: A plug-and-play network for refining human poses in videos,'' in ECCV, 2022

  68. [76]

    S \'a r \'a ndi, A

    I. S \'a r \'a ndi, A. Hermans, and B. Leibe, ``Learning 3d human pose estimation from dozens of datasets using a geometry-aware autoencoder to bridge between skeleton formats,'' in WACV, 2023

  69. [77]

    Ionescu, D

    C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, ``Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,'' IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, no. 7, pp. 1325--1339, jul 2014

  70. [78]

    Mehta, O

    D. Mehta, O. Sotnychenko, F. Mueller, W. Xu, S. Sridhar, G. Pons-Moll, and C. Theobalt, ``Single-shot multi-person 3d pose estimation from monocular rgb,'' in 3D Vision (3DV), 2018 Sixth International Conference on. 1em plus 0.5em minus 0.4em IEEE, sep 2018. [Online]. Availabl...

  71. [79]

    H. Joo, T. Simon, X. Li, H. Liu, L. Tan, L. Gui, S. Banerjee, T. S. Godisart, B. Nabbe, I. Matthews, T. Kanade, S. Nobuhara, and Y. Sheikh, ``Panoptic studio: A massively multiview system for social interaction capture,'' IEEE Transactions on Pattern Analysis and Machine Intel...

  72. [80]

    Hu, H.-S

    Y.-T. Hu, H.-S. Chen, K. Hui, J.-B. Huang, and A. G. Schwing, `` SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation -- A Synthetic Dataset and Baselines ,'' in Proc. CVPR, 2019

  73. [81]

    Tsuchida, S

    S. Tsuchida, S. Fukayama, M. Hamasaki, and M. Goto, ``Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing,'' in Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR 2019 , D...

  74. [82]

    F. Bogo, A. Kanazawa, C. Lassner, P. Gehler, J. Romero, and M. J. Black, ``Keep it smpl: Automatic estimation of 3d human pose and shape from a single image,'' in ECCV, 2016

  75. [83]

    Shimada, V

    S. Shimada, V. Golyanik, W. Xu, and C. Theobalt, ``Physcap: Physically plausible monocular 3d motion capture in real time,'' ACM ToG, 2020

  76. [84]

    Teed and J

    Z. Teed and J. Deng, ``Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras,'' in Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. 1em plus 0.5em minus 0.4em Curran Associate...

  77. [85]

    Y. Wang, Z. Wang, L. Liu, and K. Daniilidis, ``Tram: Global trajectory and motion of 3d humans from in-the-wild videos,'' arXiv preprint arXiv:2403.17346, 2024

  78. [86]

    Kirillov, E

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Doll \'a r, and R. Girshick, ``Segment anything,'' arXiv:2304.02643, 2023

  79. [87]

    Geman and D

    S. Geman and D. E. McClure, ``Statistical methods for tomographic image reconstruction,'' 1987. [Online]. Available: https://api.semanticscholar.org/CorpusID:118824639

  80. [88]

    Y. Zou, J. Yang, D. Ceylan, J. Zhang, F. Perazzi, and J.-B. Huang, ``Reducing footskate in human motion reconstruction with ground contact constraints,'' in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2020, pp. 459--468

  81. [89]

    [Online]

    OpenAI, ``Gpt-4 technical report,'' 2024. [Online]. Available: https://arxiv.org/abs/2303.08774

  82. [90]

    Dutta and A

    A. Dutta and A. Zisserman, ``The VIA annotation software for images, audio and video,'' in ACM MM, 2019

  83. [91]

    Mollahosseini, B

    A. Mollahosseini, B. Hasani, and M. H. Mahoor, ``Affectnet: A database for facial expression, valence, and arousal computing in the wild,'' IEEE Transactions on Affective Computing, 2017

  84. [92]

    Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh, ``Realtime multi-person 2d pose estimation using part affinity fields,'' in CVPR, 2017

  85. [93]

    Zhang, V

    F. Zhang, V. Bazarevsky, A. Vakunov, A. Tkachenka, G. Sung, C.-L. Chang, and M. Grundmann, ``Mediapipe hands: On-device real-time hand tracking,'' arXiv preprint arXiv:2006.10214, 2020

  86. [94]

    K. Sun, B. Xiao, D. Liu, and J. Wang, ``Deep high-resolution representation learning for human pose estimation,'' in CVPR, 2019

  87. [95]

    Zhang, D

    J. Zhang, D. Zhang, X. Xu, F. Jia, Y. Liu, X. Liu, J. Ren, and Y. Zhang, ``Mobipose: Real-time multi-person pose estimation on mobile devices,'' in SenSys, 2020

  88. [96]

    Zhang, Y

    H. Zhang, Y. Tian, Y. Zhang, M. Li, L. An, Z. Sun, and Y. Liu, ``Pymaf-x: Towards well-aligned full-body model regression from monocular images,'' in TPAMI, 2023

  89. [97]

    Kaufmann, J

    M. Kaufmann, J. Song, C. Guo, K. Shen, T. Jiang, C. Tang, J. J. Z \'a rate, and O. Hilliges, `` EMDB : The E lectromagnetic D atabase of G lobal 3 D H uman P ose and S hape in the W ild,'' in International Conference on Computer Vision (ICCV), 2023

  90. [98]

    S. Shin, J. Kim, E. Halilaj, and M. J. Black, `` WHAM : Reconstructing world-grounded humans with accurate 3D motion,'' in IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Jun. 2024

  91. [99]

    Z. Teed, L. Lipson, and J. Deng, ``Deep patch visual odometry,'' Advances in Neural Information Processing Systems, 2023

  92. [100]

    C. Guo, X. Zuo, S. Wang, X. Liu, S. Zou, M. Gong, and L. Cheng, ``Action2video: Generating videos of human 3d actions,'' IJCV, 2022

  93. [101]

    Tseng, R

    J. Tseng, R. Castellon, and K. Liu, ``Edge: Editable dance generation from music,'' in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 448--458

  94. [102]

    Patel, C.-H

    P. Patel, C.-H. P. Huang, J. Tesch, D. T. Hoffmann, S. Tripathi, and M. J. Black, `` AGORA : Avatars in geography optimized for regression analysis,'' in CVPR, 2021

  95. [103]

    Andriluka, L

    M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele, ``2d human pose estimation: New benchmark and state of the art analysis,'' in CVPR, 2014

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.