REVIEW 4 major objections 5 minor 11 cited by
Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read An automatic annotation pipeline turns 80.8K internet videos into 19.5M 3D whole-body pose annotations with text and audio, and training on the data improves motion generation and recovery.
desk verdict A genuinely large and useful dataset expansion, but the core accuracy claim rests on an unvalidated scale-recovery step and the excerpt lacks independent verification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the annotation pipeline, whose load-bearing stages are: whole-body keypoint estimation; camera tracking via dense bundle adjustment with the human masked out of feature extraction, correspondence fields, and flow; global trajectory optimization with reprojection, smoothness, ground-contact, and foot-skating losses; and text synthesis that turns estimated body-part spatial relations, hand shapes, and classifier-based emotions into frame-level descriptions, with a vision-language model providing sequence-level captions. All motion is represented in the SMPL-X whole-body parametric model, so body, hands, and face share one optimization. The pipeline's role is to convert raw RGB video into self-supervised training labels at scale.
What would settle it
Take a random sample of roughly 500 Motion-X++ clips, have annotators manually fit the whole-body model to the frames (or capture a subset with multi-view or motion-capture systems), and compare per-joint errors on hands, face, and global trajectory; if the median errors approach the typical inter-annotator variation or exceed the tolerances used in the paper's evaluation, the ground-truth claim is falsified.
Extended reading notes
Core claim
Motion-X++ asserts that every part of a human motion, including body pose, hand gestures, facial expressions, and the global trajectory, can be recovered from ordinary single-view or multi-view video by a staged pipeline, and that the recovered parameters are accurate enough to train and evaluate downstream models. The pipeline estimates whole-body 2D keypoints, tracks the camera with a masked dense-bundle-adjustment SLAM that excludes the moving person, optimizes the global trajectory with reprojection, smoothness, ground-contact, and foot-skating losses, fits SMPL-X parameters to the observations, and then generates frame-level text descriptions from body-part relations, hand gestures, and an emotion classifier, while a vision-language model supplies sequence-level captions. The paper presents the resulting dataset as ground truth and reports that training on it improves the four downstream task families.
Load-bearing premise
The load-bearing premise is that the automatic annotation pipeline, including SMPL-X fitting, masked dense-bundle-adjustment camera tracking, trajectory optimization, and vision-language captioning, produces 3D poses and text descriptions accurate enough to serve as ground truth for unconstrained internet videos.
Editorial extensions
If this is right
- Text-driven whole-body motion generation can be trained on 120.5K sequence-level captions paired with expressive 3D pose, including hands and face.
- Audio-driven motion generation gains 45.3K audio–motion pairs, enabling models that synthesize dance or performance from music.
- Whole-body mesh recovery and 2D keypoint estimation can be trained on diverse internet scenes rather than lab captures.
- Because the annotation pipeline runs on any RGB video collection, the dataset can keep growing without manual text labeling.
- Frame-level whole-body pose descriptions give models dense language supervision instead of only sequence-level captions.
Reading between the lines
- If the accuracy claim holds, Motion-X++ could also serve as pretraining data for video-to-motion and motion-to-video generation, since raw video and audio are stored alongside the motion.
- The vision-language captions are likely noisier than the geometric pose labels, so a fair text-to-motion benchmark should include human judgments of caption–pose alignment before treating the captions as ground truth.
- The strategy of masking the moving person out of dense-bundle-adjustment camera tracking could transfer to other dynamic-scene reconstruction problems, such as tracking cameras in crowds or sports footage.
- Because audio, video, pose, and text are aligned for the same sequences, the dataset opens cross-modal alignment research that motion-only datasets cannot support.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Motion-X++, a large-scale multimodal dataset of 3D whole-body human motion, providing 19.5M pose annotations over 120.5K sequences, along with RGB video, audio, frame-level pose descriptions, and sequence-level semantic labels. The core contribution is a scalable automatic annotation pipeline that combines masked DROID-SLAM camera tracking, SMPL-X whole-body fitting, global trajectory optimization with foot-contact constraints, and GPT-4V-generated captions. The paper claims that training on Motion-X++ improves text-driven motion generation, audio-driven motion generation, 3D whole-body mesh recovery, and 2D whole-body keypoint estimation.
Significance. If substantiated, Motion-X++ would be a valuable community resource: it is substantially larger and more modality-rich than existing 3D whole-body motion datasets, and the pipeline explicitly targets known failure modes such as hand gestures and facial expressions. The manuscript deserves credit for providing concrete mathematical formulations of the optimization stages (Eqs. 6-14) and for acknowledging limitations of its predecessor, Motion-X. However, the load-bearing claim that the automatic pipeline yields metric, expressive 3D poses suitable as training and evaluation ground truth for unconstrained internet videos is not yet supported by visible evidence in the reviewed excerpt.
major comments (4)
- [§Human Trajectory Refinement, Eqs. (9)-(12)] The monocular camera/body scale ambiguity is not resolved or validated. Because the dataset is built from single-view internet videos, absolute scene scale is unobservable from the images alone. The optimization in Eq. (12) fixes camera scale α and SMPL-X shape β from the previous stage and optimizes only Φ_t and Γ_t, so the resulting root translations Γ_t inherit whatever scale is encoded in α. The paper does not show an experiment comparing recovered global trajectories against metric ground truth, nor does it quantify the effect of the scale choice on downstream pose accuracy. This is load-bearing because the 'accurate 3D whole-body pose annotation' claim includes global motion, not just relative pose.
- [Abstract and §1 (pipeline overview)] The central claim that 'Comprehensive experiments validate the accuracy of our annotation pipeline' is not supported by any quantitative result in the reviewed excerpt. There are no error bars, no comparisons against independent 3D ground truth, no human perceptual evaluation, and no failure-mode analysis. Because the paper's value proposition is annotation precision, the reader cannot assess whether the large scale is achieved at the cost of systematic errors in hands, faces, and global trajectories. Please add per-component accuracy evaluations and benchmark comparisons, or clearly state which experiments are deferred.
- [Downstream task evaluation and GPT-4V captions] The excerpt does not specify the evaluation protocol for the downstream tasks, so it is unclear whether test annotations come from the proposed pipeline itself or from independent human/external ground truth. If the same pipeline outputs are used as both training signal and evaluation target, the reported gains could reflect consistency with a biased annotation process rather than accuracy. Similarly, GPT-4V captions appear to be used as both the annotation and the evaluation target for caption quality. Please clarify the evaluation protocols and include at least one cross-dataset or human-validated evaluation for both 3D pose accuracy and text annotation quality.
- [Fig. 7 and single-view inference] Fig. 7 presents a multi-view annotation pipeline and the text states that the pipeline supports any number of viewpoints, but the dataset is constructed from single-view internet videos. The mechanism by which absolute scale is recovered in the single-view setting is therefore critical and under-specified. Concretely, the text should explain how α in Eq. (10) is computed and why it is metrically reliable for single-view inputs, or provide a validation experiment on sequences with known metric scale.
minor comments (5)
- [Eq. (9)] The parentheses do not balance: 'Jt = M (Φt, Θt, β) + Γt),' should be 'Jt = M(Φt, Θt, β) + Γt'.
- [Eq. (11)] There is an unmatched closing parenthesis in '||J I t − J I t+1)||2'; please correct the notation.
- [Fig. 7] The figure is numbered 'Fig. 7' but its caption reads 'Figure 1. Muti-view Annotation Pipeline'; the numbering and the typo 'Muti' should be fixed.
- [Eq. (10)] The symbol α is used in the reprojection loss before it is defined. Please introduce α explicitly, state its units, and explain how it is derived from the previous stage.
- [Masked DROID-SLAM description] The list of two key differences in the masked DROID-SLAM strategy is written as one paragraph with '1)We' missing a space; please reformat for readability.
Circularity Check
No significant circularity: dataset construction pipeline is sequential and no prediction reduces to its inputs.
full rationale
The paper is a dataset-construction paper; it does not claim a first-principles derivation whose conclusion is equivalent to its premises. The core pipeline (masked DROID-SLAM camera tracking, SMPL-X fitting, trajectory optimization with Eqs. 6-14, and GPT-4V captioning) takes raw RGB video and detected 2D keypoints as inputs and produces 3D pose, trajectory, and text labels; no equation defines an output in terms of the same output. The monocular scale-depth ambiguity flagged in the skeptic attack is a real validation risk, but it is an identifiability/correctness issue rather than a circular reduction, because camera scale alpha is carried forward from an earlier stage, not derived from the trajectory it is used to refine. The manuscript's own limitation statement about Motion-X's 'inaccurate hand gestures, collapsed facial...' is an acknowledged failure mode of the prior pipeline, not a self-referential justification. The references to prior work such as SLAHMR and DROID-SLAM are standard external methods, and the reference to Motion-X is dataset lineage rather than load-bearing evidence. No quoted passage in the provided text shows downstream evaluation using the same pipeline's annotations as both training signal and evaluation target; without that specific reduction, no circularity can be scored.
Assumptions & free parameters
free parameters (1)
- Trajectory optimization hyperparameters =
not reported
assumptions (4)
- domain assumption SMPL-X is a faithful whole-body body model, including hands and face.
- domain assumption Masked DROID-SLAM yields reliable camera pose and depth in dynamic scenes.
- domain assumption GPT-4V captions and rule-based pose descriptions are semantically aligned with the motion.
- domain assumption Merging eight existing action datasets preserves label consistency.
Cite this review
Pith. "Pith review of Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset." pith.science (2026). https://pith.science/paper/GAJ57S6R
@misc{pith2026250105098,
author = {Pith},
title = {Pith review of: Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset},
year = {2026},
howpublished = {\url{https://pith.science/paper/GAJ57S6R}},
note = {Machine review of arXiv:2501.05098}
}
read the original abstract
In this paper, we introduce Motion-X++, a large-scale multimodal 3D expressive whole-body human motion dataset. Existing motion datasets predominantly capture body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions, and are typically limited to lab settings with manually labeled text descriptions, thereby restricting their scalability. To address this issue, we develop a scalable annotation pipeline that can automatically capture 3D whole-body human motion and comprehensive textural labels from RGB videos and build the Motion-X dataset comprising 81.1K text-motion pairs. Furthermore, we extend Motion-X into Motion-X++ by improving the annotation pipeline, introducing more data modalities, and scaling up the data quantities. Motion-X++ provides 19.5M 3D whole-body pose annotations covering 120.5K motion sequences from massive scenes, 80.8K RGB videos, 45.3K audios, 19.5M frame-level whole-body pose descriptions, and 120.5K sequence-level semantic labels. Comprehensive experiments validate the accuracy of our annotation pipeline and highlight Motion-X++'s significant benefits for generating expressive, precise, and natural motion with paired multimodal labels supporting several downstream tasks, including text-driven whole-body motion generation,audio-driven motion generation, 3D whole-body human mesh recovery, and 2D whole-body keypoints estimation, etc.
Figures
Figures from the paper (10 more)
Forward citations
Cited by 11 Pith papers
-
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
HumanTracker introduces a 153-hour categorized humanoid tracking benchmark and a preference-trained metric, HumanScore, that agrees with human judgments better than kinematic error metrics.
-
MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval
MRBench is a multi-source, balanced, multi-granular human motion-text retrieval benchmark, and the proposed granularity-aware adapters improve mixed-granularity retrieval without degrading standard-caption retrieval.
-
$\omega$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
A single whole-body model with latent future prediction outperforms prior robot policies on 11 real-world humanoid household loco-manipulation tasks.
-
UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation
A visual proxy that renders human motion under the driving camera and overlays camera trajectory markers lets a video diffusion model control both body motion and camera movement from a single visual conditioning space.
-
EgoHTR: Egocentric 4D Demonstrations of Human Terrain Traversal
EgoHTR is a 55-sequence, 150k-frame egocentric 4D human-terrain dataset with a reconstruction pipeline, MoCap-validated benchmark, and perceptive locomotion policies deployed on a Unitree G1.
-
GUSH3R: Everyone Everywhere All at Once as Gaussians
A feed-forward framework predicts unified 3D Gaussian representations of dynamic humans and static scenes from monocular video in a single forward pass.
-
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
RAM couples motion reconstruction with text-to-motion diffusion and adds reconstruction-anchored error guidance, reporting FID 0.032 on HumanML3D with 20 inference steps.
-
Distinguishing Imitation Error from Intrinsic Motion Learning Difficulty
A physics-based score (MDS) predicts how hard a motion is for a humanoid to imitate by measuring how much joint torques must change under small pose perturbations.
-
PHUMA: Physically Reliable Humanoid Locomotion Dataset
PHUMA is a curated 73-hour humanoid locomotion corpus whose physical-reliability metrics are partly defined by the same losses used to optimize it, and whose imitation success claims are confounded by in-distribution ...
-
The loss tolerance of cat breeding for fault-tolerant grid state generation
Claims a 4% optical-loss ceiling for fault-tolerant GKP state generation via cat breeding, but the provided full text is an unrelated manuscript with no such analysis.
-
VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
A unified multimodal motion LLM that handles nine generation and comprehension tasks across text, audio, and single/multi-agent motion, backed by a new dataset and tokenizer.
Reference graph
Works this paper leans on
-
[1]
'E4nSt C777dYs3 Pi^w.x j իW?k ].W Isssuuu^WQ ED ^н]@ /^
11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 \@IEEEnotcompsoconly \@IEEEcompsoconly #1 * [1] 0pt [0pt][0pt] #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEcompsocconfonly \@IEEEauthordefaulttextstyle \@IEEEcompsocnotconfonly \@IEEEauthor...
arXiv 1939
-
[2]
Ahuja and L.-P
C. Ahuja and L.-P. Morency, ``Language2pose: Natural language grounded pose forecasting,'' in 3DV, 2019
2019
-
[3]
X. Chen, B. Jiang, W. Liu, Z. Huang, B. Fu, T. Chen, J. Yu, and G. Yu, ``Executing your commands via motion diffusion in latent space,'' in CVPR, 2023
2023
-
[4]
Delmas, P
G. Delmas, P. Weinzaepfel, T. Lucas, F. Moreno-Noguer, and G. Rogez, ``Posescript: 3d human poses from natural language,'' in ECCV, 2022
2022
-
[5]
C. Guo, S. Zou, X. Zuo, S. Wang, W. Ji, X. Li, and L. Cheng, ``Generating diverse and natural 3d human motions from text,'' in CVPR, 2022
2022
-
[6]
Petrovich, M
M. Petrovich, M. J. Black, and G. Varol, ``Temos: Generating diverse human motions from textual descriptions,'' in ECCV, 2022
2022
-
[7]
Plappert, C
M. Plappert, C. Mandery, and T. Asfour, ``The kit motion-language dataset,'' Big data, 2016
2016
-
[8]
Plappert, C
M. Plappert, C. Mandery, and T.Asfour, ``Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks,'' Robotics and Autonomous Systems, 2018
2018
Show all 103 references
-
[9]
A. R. Punnakkal, A. Chandrasekaran, N. Athanasiou, A. Quiros-Ramirez, and M. J. Black, ``Babel: bodies, action and behavior with english labels,'' in CVPR, 2021
2021
-
[10]
Zhang, Z
M. Zhang, Z. Cai, L. Pan, F. Hong, X. Guo, L. Yang, and Z. Liu, ``Motiondiffuse: Text-driven human motion generation with diffusion model,'' arXiv preprint arXiv:2208.15001, 2022
2022 arXiv
-
[11]
M. Zhao, M. Liu, B. Ren, S. Dai, and N. Sebe, ``Modiff: Action-conditioned 3d motion generation with denoising diffusion probabilistic models,'' arXiv preprint arXiv:2301.03949, 2023
2023 arXiv
-
[12]
L.-H. Chen, S. Lu, A. Zeng, H. Zhang, B. Wang, R. Zhang, and L. Zhang, ``Motionllm: Understanding human behaviors from human motions and videos,'' arXiv preprint arXiv:2405.20340, 2024
2024 arXiv
-
[13]
F. Hong, L. Pan, Z. Cai, and Z. Liu, ``Versatile multi-modal pre-training for human-centric perception,'' arXiv preprint arXiv:2203.13815, 2022
2022 arXiv
-
[14]
Jiang, X
B. Jiang, X. Chen, W. Liu, J. Yu, G. Yu, and T. Chen, ``Motiongpt: Human motion as a foreign language,'' Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[15]
Z. Zhou, Y. Wan, and B. Wang, ``Avatargpt: All-in-one framework for motion understanding, planning, generation and beyond,'' 2023. [Online]. Available: https://arxiv.org/abs/2311.16468
2023 arXiv
-
[16]
J. Wang, Y. Rong, J. Liu, S. Yan, D. Lin, and B. Dai, ``Towards diverse and natural scene-aware 3d human motion synthesis,'' 2022. [Online]. Available: https://arxiv.org/abs/2205.13001
2022 arXiv
-
[17]
J. Wang, Y. Yuan, Z. Luo, K. Xie, D. Lin, U. Iqbal, S. Fidler, and S. Khamis, ``Learning human dynamics in autonomous driving scenarios,'' in 2023 IEEE/CVF International Conference on Computer Vision (ICCV). 1em plus 0.5em minus 0.4em Los Alamitos, CA, USA: IEEE Computer Socie...
2023
-
[18]
Z. Xiao, T. Wang, J. Wang, J. Cao, W. Zhang, B. Dai, D. Lin, and J. Pang, ``Unified human-scene interaction via prompted chain-of-contacts,'' in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=1vCnDyQkjg
2024
-
[19]
Mahmood, N
N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, and M. J. Black, ``Amass: Archive of motion capture as surface shapes,'' in ICCV, 2019
2019
-
[20]
J. Ho, A. Jain, and P. Abbeel, ``Denoising diffusion probabilistic models,'' NeurIPS, 2020
2020
-
[21]
J. Song, C. Meng, and S. Ermon, ``Denoising diffusion implicit models,'' in ICLR, 2020
2020
-
[22]
Tevet, S
G. Tevet, S. Raab, B. Gordon, Y. Shafir, D. Cohen-Or, and A. H. Bermano, ``Human motion diffusion model,'' in ICLR, 2023
2023
-
[23]
J. Lin, A. Zeng, H. Wang, L. Zhang, and Y. Li, ``One-stage 3d whole-body mesh recovery with component aware transformer,'' in CVPR, 2023
2023
-
[24]
R. Li, S. Yang, D. A. Ross, and A. Kanazawa, ``Ai choreographer: Music conditioned 3d dance generation with aist++,'' in ICCV, 2021
2021
-
[25]
G. Moon, H. Choi, and K. M. Lee, ``Accurate 3d hand pose estimation for whole-body 3d human mesh estimation,'' in CVPRW, 2020
2020
-
[26]
Y. Xu, J. Zhang, Q. Zhang, and D. Tao, ``Vitpose: Simple vision transformer baselines for human pose estimation,'' in NeurIPS, 2022
2022
-
[27]
Y. Yuan, U. Iqbal, P. Molchanov, K. Kitani, and J. Kautz, ``Glamr: Global occlusion-aware human mesh recovery with dynamic cameras,'' in CVPR, 2022
2022
-
[28]
J. Yang, A. Zeng, S. Liu, F. Li, R. Zhang, and L. Zhang, ``Explicit box detection unifies end-to-end multi-person pose estimation,'' in ICLR, 2023
2023
-
[29]
H. E. Pang, Z. Cai, L. Yang, T. Zhang, and Z. Liu, ``Benchmarking and analyzing 3d human pose and shape estimation beyond algorithms,'' in NeurIPS Datasets and Benchmarks Track, 2022
2022
-
[30]
G. Moon, H. Choi, and K. M. Lee, ``Neuralannot: Neural annotator for 3d human mesh training sets,'' in CVPR, 2022
2022
-
[31]
G. Moon, H. Choi, S. Chun, J. Lee, and S. Yun, ``Three recipes for better 3d pseudo-gts of 3d human mesh estimation in the wild,'' in CVPR, 2023
2023
-
[32]
H. Yi, H. Liang, Y. Liu, Q. Cao, Y. Wen, T. Bolkart, D. Tao, and M. J. Black, ``Generating holistic 3d human motion from speech,'' in CVPR, 2023
2023
-
[33]
Pavlakos, V
G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. Osman, D. Tzionas, and M. J. Black, ``Expressive body capture: 3d hands, face, and body from a single image,'' in CVPR, 2019
2019
-
[34]
Z. Cai, D. Ren, A. Zeng, Z. Lin, T. Yu, W. Wang, X. Fan, Y. Gao, Y. Yu, L. Pan, F. Hong, M. Zhang, C. C. Loy, L. Yang, and Z. Liu, ``Humman: Multi-modal 4d human dataset for versatile sensing and modeling,'' in ECCV, 2022
2022
-
[35]
Chung, C.-h
J. Chung, C.-h. Wuu, H.-r. Yang, Y.-W. Tai, and C.-K. Tang, ``Haa500: Human-centric atomic action dataset with curated videos,'' in ICCV, 2021
2021
-
[36]
J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, ``Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,'' in TPAMI, 2019
2019
-
[37]
Taheri, N
O. Taheri, N. Ghorbani, M. J. Black, and D. Tzionas, ``Grab: A dataset of whole-body human grasping of objects,'' in ECCV, 2020
2020
-
[38]
Tsuchida, S
S. Tsuchida, S. Fukayama, M. Hamasaki, and M. Goto, ``Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing.'' in ISMIR, 2019
2019
-
[39]
Zhalehpour, O
S. Zhalehpour, O. Onder, Z. Akhtar, and C. E. Erdem, ``Baum-1: A spontaneous audio-visual face database of affective and mental states,'' IEEE Transactions on Affective Computing, 2016
2016
-
[40]
Zhang, Q
S. Zhang, Q. Ma, Y. Zhang, Z. Qian, T. Kwon, M. Pollefeys, F. Bogo, and S. Tang, ``Egobody: Human body shape and motion of interacting people from head-mounted devices,'' in ECCV, 2022
2022
-
[41]
J. Lin, A. Zeng, S. Lu, Y. Cai, R. Zhang, H. Wang, and L. Zhang, ``Motion-x: A large-scale 3d expressive whole-body human motion dataset,'' Advances in Neural Information Processing Systems, 2023
2023
-
[42]
Dan e c ek, M
R. Dan e c ek, M. J. Black, and T. Bolkart, ``Emoca: Emotion driven monocular face capture and animation,'' in CVPR, 2022
2022
-
[43]
Pavlakos, D
G. Pavlakos, D. Shan, I. Radosavovic, A. Kanazawa, D. Fouhey, and J. Malik, ``Reconstructing hands in 3d with transformers,'' in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 9826--9836
2024
-
[44]
Z. Cai, W. Yin, A. Zeng, C. Wei, Q. Sun, W. Yanjun, H. E. Pang, H. Mei, M. Zhang, L. Zhang et al., ``Smpler-x: Scaling up expressive human pose and shape estimation,'' Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[45]
V. Ye, G. Pavlakos, J. Malik, and A. Kanazawa, ``Decoupling human and camera motion from videos in the wild,'' in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2023
2023
-
[46]
Carreira, E
J. Carreira, E. Noland, C. Hillier, and A. Zisserman, ``A short note on the kinetics-700 human action dataset,'' arXiv preprint arXiv:1907.06987, 2019
1907 arXiv
-
[47]
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar et al., ``Ava: A video dataset of spatio-temporally localized atomic visual actions,'' in CVPR, 2018
2018
-
[48]
Shahroudy, J
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, ``Ntu rgb+ d: A large scale dataset for 3d human activity analysis,'' in CVPR, 2016
2016
-
[49]
Trivedi, A
N. Trivedi, A. Thatipelli, and R. K. Sarvadevabhatla, ``Ntu-x: an enhanced large-scale dataset for improving pose-based recognition of subtle human actions,'' in ICVGIP, 2021
2021
-
[50]
Hassan, D
M. Hassan, D. Ceylan, R. Villegas, J. Saito, J. Yang, Y. Zhou, and M. J. Black, ``Stochastic scene-aware motion prediction,'' in ICCV, 2021
2021
-
[51]
Hassan, V
M. Hassan, V. Choutas, D. Tzionas, and M. J. Black, ``Resolving 3d human pose ambiguities with 3d scene constraints,'' in ICCV, 2019
2019
-
[52]
Y.-L. Li, X. Liu, X. Wu, Y. Li, Z. Qiu, L. Xu, Y. Xu, H.-S. Fang, and C. Lu, ``Hake: a knowledge engine foundation for human activity understanding,'' in TPAMI, 2022
2022
-
[53]
Zheng, Y
Y. Zheng, Y. Yang, K. Mo, J. Li, T. Yu, Y. Liu, C. K. Liu, and L. J. Guibas, ``Gimo: Gaze-informed human motion prediction in context,'' in ECCV, 2022
2022
-
[54]
C. Guo, X. Zuo, S. Wang, S. Zou, Q. Sun, A. Deng, M. Gong, and L. Cheng, ``Action2motion: Conditioned generation of 3d human motions,'' in ACM MM, 2020
2020
-
[55]
Gross and J
R. Gross and J. Shi, ``The cmu motion of body (mobo) database,'' 2001
2001
-
[56]
Ionescu, D
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, `` Human3.6M : Large scale datasets and predictive methods for 3d human sensing in natural environments,'' in TPAMI, 2014
2014
-
[57]
Sigal, A
L. Sigal, A. O. Balan, and M. J. Black, ``Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion,'' IJCV, 2010
2010
-
[58]
Trumble, A
M. Trumble, A. Gilbert, C. Malleson, A. Hilton, and J. P. Collomosse, ``Total capture: 3d human pose estimation fusing video and inertial sensors,'' in BMVC, 2017
2017
-
[59]
Loper, N
M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black, ``Smpl: A skinned multi-person linear model,'' ACM TOG, 2015
2015
-
[60]
Y. Yuan, J. Song, U. Iqbal, A. Vahdat, and J. Kautz, ``Physdiff: Physics-guided human motion diffusion model,'' in ICCV, 2023
2023
-
[61]
Zhang, Y
J. Zhang, Y. Zhang, X. Cun, S. Huang, Y. Zhang, H. Zhao, H. Lu, and X. Shen, ``T2m-gpt: Generating human motion from textual descriptions with discrete representations,'' in CVPR, 2023
2023
-
[62]
Zhuang, C
W. Zhuang, C. Wang, J. Chai, Y. Wang, M. Shao, and S. Xia, ``Music2dance: Dancenet for music-driven dance generation,'' ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), vol. 18, no. 2, pp. 1--21, 2022
2022
-
[63]
R. Li, J. Zhao, Y. Zhang, M. Su, Z. Ren, H. Zhang, Y. Tang, and X. Li, ``Finedance: A fine-grained choreography dataset for 3d full body dance generation,'' in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 10\,234--10\,243
2023
-
[64]
W. Jiao, W. Wang, J.-t. Huang, X. Wang, and Z. Tu, ``Is chatgpt a good translator? a preliminary study,'' arXiv preprint arXiv:2301.08745, 2023
2023 arXiv
-
[65]
Zablotskaia, A
P. Zablotskaia, A. Siarohin, B. Zhao, and L. Sigal, ``Dwnet: Dense warp-based network for pose-guided human video generation,'' arXiv preprint arXiv:1910.09139, 2019
1910 arXiv
-
[66]
Huang, Y
Q. Huang, Y. Xiong, A. Rao, J. Wang, and D. Lin, ``Movienet: A holistic dataset for movie understanding,'' in The European Conference on Computer Vision (ECCV), 2020
2020
-
[67]
Contributors, `` MMTracking: OpenMMLab video perception toolbox and benchmark,'' https://github.com/open-mmlab/mmtracking, 2020
M. Contributors, `` MMTracking: OpenMMLab video perception toolbox and benchmark,'' https://github.com/open-mmlab/mmtracking, 2020
2020
-
[68]
R. E. K \'a lm \'a n and R. S. Bucy, ``New results in linear filtering and prediction theory,'' Journal of Basic Engineering, vol. 83, pp. 95--108, 1961. [Online]. Available: https://api.semanticscholar.org/CorpusID:8141345
1961
-
[69]
Teed and J
Z. Teed and J. Deng, ``Raft: Recurrent all-pairs field transforms for optical flow,'' 2020. [Online]. Available: https://arxiv.org/abs/2003.12039
2020 arXiv
-
[70]
S. Jin, L. Xu, J. Xu, C. Wang, W. Liu, C. Qian, W. Ouyang, and P. Luo, ``Whole-body human pose estimation in the wild,'' in ECCV, 2020
2020
-
[71]
L. Xu, S. Jin, W. Liu, C. Qian, W. Ouyang, P. Luo, and X. Wang, ``Zoomnas: searching for whole-body human pose estimation in the wild,'' TPAMI, 2022
2022
-
[72]
Narasimhaswamy, T
S. Narasimhaswamy, T. Nguyen, M. Huang, and M. Hoai, ``Whose hands are these? hand detection and hand-body association in the wild,'' in CVPR, 2022
2022
-
[73]
Savitzky and M
A. Savitzky and M. J. Golay, ``Smoothing and differentiation of data by simplified least squares procedures.'' Analytical chemistry, 1964
1964
-
[74]
A. Zeng, X. Ju, L. Yang, R. Gao, X. Zhu, B. Dai, and Q. Xu, ``Deciwatch: A simple baseline for 10x efficient 2d and 3d pose estimation,'' in ECCV, 2022
2022
-
[75]
A. Zeng, L. Yang, X. Ju, J. Li, J. Wang, and Q. Xu, ``Smoothnet: A plug-and-play network for refining human poses in videos,'' in ECCV, 2022
2022
-
[76]
S \'a r \'a ndi, A
I. S \'a r \'a ndi, A. Hermans, and B. Leibe, ``Learning 3d human pose estimation from dozens of datasets using a geometry-aware autoencoder to bridge between skeleton formats,'' in WACV, 2023
2023
-
[77]
Ionescu, D
C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu, ``Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,'' IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, no. 7, pp. 1325--1339, jul 2014
2014
-
[78]
Mehta, O
D. Mehta, O. Sotnychenko, F. Mueller, W. Xu, S. Sridhar, G. Pons-Moll, and C. Theobalt, ``Single-shot multi-person 3d pose estimation from monocular rgb,'' in 3D Vision (3DV), 2018 Sixth International Conference on. 1em plus 0.5em minus 0.4em IEEE, sep 2018. [Online]. Availabl...
2018
-
[79]
H. Joo, T. Simon, X. Li, H. Liu, L. Tan, L. Gui, S. Banerjee, T. S. Godisart, B. Nabbe, I. Matthews, T. Kanade, S. Nobuhara, and Y. Sheikh, ``Panoptic studio: A massively multiview system for social interaction capture,'' IEEE Transactions on Pattern Analysis and Machine Intel...
2017
-
[80]
Hu, H.-S
Y.-T. Hu, H.-S. Chen, K. Hui, J.-B. Huang, and A. G. Schwing, `` SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation -- A Synthetic Dataset and Baselines ,'' in Proc. CVPR, 2019
2019
-
[81]
Tsuchida, S
S. Tsuchida, S. Fukayama, M. Hamasaki, and M. Goto, ``Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing,'' in Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR 2019 , D...
2019
-
[82]
F. Bogo, A. Kanazawa, C. Lassner, P. Gehler, J. Romero, and M. J. Black, ``Keep it smpl: Automatic estimation of 3d human pose and shape from a single image,'' in ECCV, 2016
2016
-
[83]
Shimada, V
S. Shimada, V. Golyanik, W. Xu, and C. Theobalt, ``Physcap: Physically plausible monocular 3d motion capture in real time,'' ACM ToG, 2020
2020
-
[84]
Teed and J
Z. Teed and J. Deng, ``Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras,'' in Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. 1em plus 0.5em minus 0.4em Curran Associate...
2021
-
[85]
Y. Wang, Z. Wang, L. Liu, and K. Daniilidis, ``Tram: Global trajectory and motion of 3d humans from in-the-wild videos,'' arXiv preprint arXiv:2403.17346, 2024
2024 arXiv
-
[86]
Kirillov, E
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Doll \'a r, and R. Girshick, ``Segment anything,'' arXiv:2304.02643, 2023
2023 arXiv
-
[87]
Geman and D
S. Geman and D. E. McClure, ``Statistical methods for tomographic image reconstruction,'' 1987. [Online]. Available: https://api.semanticscholar.org/CorpusID:118824639
1987
-
[88]
Y. Zou, J. Yang, D. Ceylan, J. Zhang, F. Perazzi, and J.-B. Huang, ``Reducing footskate in human motion reconstruction with ground contact constraints,'' in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2020, pp. 459--468
2020
-
[89]
[Online]
OpenAI, ``Gpt-4 technical report,'' 2024. [Online]. Available: https://arxiv.org/abs/2303.08774
2024 arXiv
-
[90]
Dutta and A
A. Dutta and A. Zisserman, ``The VIA annotation software for images, audio and video,'' in ACM MM, 2019
2019
-
[91]
Mollahosseini, B
A. Mollahosseini, B. Hasani, and M. H. Mahoor, ``Affectnet: A database for facial expression, valence, and arousal computing in the wild,'' IEEE Transactions on Affective Computing, 2017
2017
-
[92]
Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh, ``Realtime multi-person 2d pose estimation using part affinity fields,'' in CVPR, 2017
2017
-
[93]
Zhang, V
F. Zhang, V. Bazarevsky, A. Vakunov, A. Tkachenka, G. Sung, C.-L. Chang, and M. Grundmann, ``Mediapipe hands: On-device real-time hand tracking,'' arXiv preprint arXiv:2006.10214, 2020
2006 arXiv
-
[94]
K. Sun, B. Xiao, D. Liu, and J. Wang, ``Deep high-resolution representation learning for human pose estimation,'' in CVPR, 2019
2019
-
[95]
Zhang, D
J. Zhang, D. Zhang, X. Xu, F. Jia, Y. Liu, X. Liu, J. Ren, and Y. Zhang, ``Mobipose: Real-time multi-person pose estimation on mobile devices,'' in SenSys, 2020
2020
-
[96]
Zhang, Y
H. Zhang, Y. Tian, Y. Zhang, M. Li, L. An, Z. Sun, and Y. Liu, ``Pymaf-x: Towards well-aligned full-body model regression from monocular images,'' in TPAMI, 2023
2023
-
[97]
Kaufmann, J
M. Kaufmann, J. Song, C. Guo, K. Shen, T. Jiang, C. Tang, J. J. Z \'a rate, and O. Hilliges, `` EMDB : The E lectromagnetic D atabase of G lobal 3 D H uman P ose and S hape in the W ild,'' in International Conference on Computer Vision (ICCV), 2023
2023
-
[98]
S. Shin, J. Kim, E. Halilaj, and M. J. Black, `` WHAM : Reconstructing world-grounded humans with accurate 3D motion,'' in IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), Jun. 2024
2024
-
[99]
Z. Teed, L. Lipson, and J. Deng, ``Deep patch visual odometry,'' Advances in Neural Information Processing Systems, 2023
2023
-
[100]
C. Guo, X. Zuo, S. Wang, X. Liu, S. Zou, M. Gong, and L. Cheng, ``Action2video: Generating videos of human 3d actions,'' IJCV, 2022
2022
-
[101]
Tseng, R
J. Tseng, R. Castellon, and K. Liu, ``Edge: Editable dance generation from music,'' in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 448--458
2023
-
[102]
Patel, C.-H
P. Patel, C.-H. P. Huang, J. Tesch, D. T. Hoffmann, S. Tripathi, and M. J. Black, `` AGORA : Avatars in geography optimized for regression analysis,'' in CVPR, 2021
2021
-
[103]
Andriluka, L
M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele, ``2d human pose estimation: New benchmark and state of the art analysis,'' in CVPR, 2014
2014
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.