{"id":"11f18243-6523-432a-a793-149297886c8f","arxiv_id":"2501.05098","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Motion-X++ provides 19.5M 3D whole-body pose annotations across 120.5K sequences with text, audio, video, and motion modalities.","lead":"This paper introduces Motion-X++, a large-scale multimodal 3D whole-body human motion dataset with 120.5K sequences, 19.5M pose annotations, and paired text, audio, and video. It offers a new automated annotation pipeline and reports benefits for expressive motion generation and human pose estimation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Monocular camera/body scale ambiguity is not validated; claimed 3D accuracy rests on an untested scale transfer from SMPL-X shape to DROID-SLAM depth.","rationale":"The reader's weakest assumption correctly identifies the automatic annotation pipeline's accuracy as load-bearing. My concern is more specific: the monocular scale ambiguity between camera translation α and body shape β (Eqs. 9–12) is the concrete mechanism that could corrupt global trajectories and absolute pose scale, and it is not addressed by any validation shown in the excerpt. This is not an accusation of wrongdoing; it is an ordinary scientific gap. The reader's CONDITIONAL verdict already captures the need for validation and data release, so my stress test does not move the verdict. I partially agree with the reader because the reader also mentioned hand/face and caption quality, which are related but distinct failure modes; the scale ambiguity is the deepest technical risk. The proposed test—reconstructing a small set of held-out sequences with known ground truth—would settle whether the scale-transfer step works, which is the weakest link in the pipeline.","tokens_in":3580,"tokens_out":3265,"duration_ms":36167,"concrete_test":"Sample 200 sequences across the dataset's scene and motion categories (e.g., game, animation, Tai Chi, outdoor sports). Re-run the full annotation pipeline, then compare the resulting SMPL-X joint positions and root trajectories against a high-precision reference: synchronously capture these scenes with a calibrated multi-view camera rig and a commercial MoCap system, or use an existing in-the-wild benchmark with ground-truth 3D pose if coverage is sufficient. Report mean per-joint position error (MPJPE) with and without Procrustes alignment, and absolute trajectory error (ATE) in meters. If MPJPE exceeds 50 mm or ATE exceeds 0.2 m on this subset, the claimed metric accuracy is not substantiated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Motion-X++ provides accurate 3D whole-body pose annotations for 120.5K sequences from diverse internet videos. The described pipeline estimates camera poses and depths monocularly via masked DROID-SLAM (Eqs. 6–8), then optimizes global trajectory with a fixed camera scale α derived from a previous stage (Eqs. 9–12). In monocular reconstruction, absolute scene scale is unconstrained; if α comes from SMPL-X body shape, then α and body shape β are coupled through the well-known scale-depth ambiguity. The excerpt provides no ablation, error bars, or comparison against ground-truth 3D data showing that root translations Γ_t are metrically accurate, nor that hand and face parameters avoid the 'collapsed facial' and 'inaccurate hand gestures' problems acknowledged in Motion-X. Fig. 7 illustrates a multi-view pipeline, but the dataset is built from single-view internet videos, so the actual scale-recovery mechanism is critical yet unvalidated. Without this validation, the downstream task gains—while plausible from data scale alone—cannot be attributed to annotation precision, which is the paper's core value proposition.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Motion-X++, a large-scale multimodal dataset of 3D whole-body human motion, providing 19.5M pose annotations over 120.5K sequences, along with RGB video, audio, frame-level pose descriptions, and sequence-level semantic labels. The core contribution is a scalable automatic annotation pipeline that combines masked DROID-SLAM camera tracking, SMPL-X whole-body fitting, global trajectory optimization with foot-contact constraints, and GPT-4V-generated captions. The paper claims that training on Motion-X++ improves text-driven motion generation, audio-driven motion generation, 3D whole-body mesh recovery, and 2D whole-body keypoint estimation.","tokens_in":3780,"tokens_out":4967,"duration_ms":48061,"significance":"If substantiated, Motion-X++ would be a valuable community resource: it is substantially larger and more modality-rich than existing 3D whole-body motion datasets, and the pipeline explicitly targets known failure modes such as hand gestures and facial expressions. The manuscript deserves credit for providing concrete mathematical formulations of the optimization stages (Eqs. 6-14) and for acknowledging limitations of its predecessor, Motion-X. However, the load-bearing claim that the automatic pipeline yields metric, expressive 3D poses suitable as training and evaluation ground truth for unconstrained internet videos is not yet supported by visible evidence in the reviewed excerpt.","major_comments":[{"comment":"The monocular camera/body scale ambiguity is not resolved or validated. Because the dataset is built from single-view internet videos, absolute scene scale is unobservable from the images alone. The optimization in Eq. (12) fixes camera scale α and SMPL-X shape β from the previous stage and optimizes only Φ_t and Γ_t, so the resulting root translations Γ_t inherit whatever scale is encoded in α. The paper does not show an experiment comparing recovered global trajectories against metric ground truth, nor does it quantify the effect of the scale choice on downstream pose accuracy. This is load-bearing because the 'accurate 3D whole-body pose annotation' claim includes global motion, not just relative pose.","section":"§Human Trajectory Refinement, Eqs. (9)-(12)"},{"comment":"The central claim that 'Comprehensive experiments validate the accuracy of our annotation pipeline' is not supported by any quantitative result in the reviewed excerpt. There are no error bars, no comparisons against independent 3D ground truth, no human perceptual evaluation, and no failure-mode analysis. Because the paper's value proposition is annotation precision, the reader cannot assess whether the large scale is achieved at the cost of systematic errors in hands, faces, and global trajectories. Please add per-component accuracy evaluations and benchmark comparisons, or clearly state which experiments are deferred.","section":"Abstract and §1 (pipeline overview)"},{"comment":"The excerpt does not specify the evaluation protocol for the downstream tasks, so it is unclear whether test annotations come from the proposed pipeline itself or from independent human/external ground truth. If the same pipeline outputs are used as both training signal and evaluation target, the reported gains could reflect consistency with a biased annotation process rather than accuracy. Similarly, GPT-4V captions appear to be used as both the annotation and the evaluation target for caption quality. Please clarify the evaluation protocols and include at least one cross-dataset or human-validated evaluation for both 3D pose accuracy and text annotation quality.","section":"Downstream task evaluation and GPT-4V captions"},{"comment":"Fig. 7 presents a multi-view annotation pipeline and the text states that the pipeline supports any number of viewpoints, but the dataset is constructed from single-view internet videos. The mechanism by which absolute scale is recovered in the single-view setting is therefore critical and under-specified. Concretely, the text should explain how α in Eq. (10) is computed and why it is metrically reliable for single-view inputs, or provide a validation experiment on sequences with known metric scale.","section":"Fig. 7 and single-view inference"}],"minor_comments":[{"comment":"The parentheses do not balance: 'Jt = M (Φt, Θt, β) + Γt),' should be 'Jt = M(Φt, Θt, β) + Γt'.","section":"Eq. (9)"},{"comment":"There is an unmatched closing parenthesis in '||J I t − J I t+1)||2'; please correct the notation.","section":"Eq. (11)"},{"comment":"The figure is numbered 'Fig. 7' but its caption reads 'Figure 1. Muti-view Annotation Pipeline'; the numbering and the typo 'Muti' should be fixed.","section":"Fig. 7"},{"comment":"The symbol α is used in the reprojection loss before it is defined. Please introduce α explicitly, state its units, and explain how it is derived from the previous stage.","section":"Eq. (10)"},{"comment":"The list of two key differences in the masked DROID-SLAM strategy is written as one paragraph with '1)We' missing a space; please reformat for readability.","section":"Masked DROID-SLAM description"}],"recommendation":"major_revision","confidential_remarks":"The reviewed excerpt appears to be an incomplete version of the manuscript, with no experimental section visible. The scale ambiguity is the main technical risk, and a focused validation on a small set of sequences with known metric ground truth would substantially address it. I would recommend requesting the complete manuscript before a final accept/reject decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a serious dataset paper, but the accuracy claim is not yet proven. Motion-X++ is a genuine expansion of the authors' earlier Motion-X—more sequences, new audio/video modalities, frame-level pose descriptions, and a detailed annotation pipeline. The scale is impressive: 19.5M 3D pose annotations across 120.5K sequences, 80.8K videos, 45.3K audios. If the annotations are trustworthy, this is a major resource for whole-body motion generation, mesh recovery, and pose estimation.\n\nWhat the paper does well: the pipeline is described at a level of detail that makes it auditable. Masked DROID-SLAM with explicit equations for camera pose and depth, trajectory optimization with foot-contact constraints, and GPT-4V captioning are all laid out. The authors are also transparent about the predecessor's problems—'inaccurate hand gestures' and 'collapsed facial' expressions—and claim the new pipeline fixes them.\n\nThe soft spot is the one the stress-test note identifies: the monocular scale ambiguity. The camera scale α is derived from SMPL-X shape, and that couples α to body shape β. The excerpt shows no ablation or ground-truth comparison demonstrating that root translations Γ_t are metrically accurate. Without that, the claim of 'accurate' 3D annotation is not fully supported. Relatedly, the downstream evaluations appear to use the same pipeline for both training and evaluation, and GPT-4V captions are both the annotation and the evaluation target. That makes it hard to attribute gains to annotation precision rather than data scale alone. There is also no visible human validation or error bars, and no release artifacts—no code, data, or prompts.\n\nThese are real weaknesses, but they are not fatal. The scale statistics are internally consistent, and the pipeline is detailed enough to reproduce and test. This is precisely the kind of work that should go to peer review: a potentially significant dataset with a method that can be evaluated for a specific weakness. A serious referee should ask for scale validation, independent hand/face metrics, and a release plan before acceptance.\n\nRecommendation: send it to review, conditional on the missing validation. I would engage with it, and I'd cite it if I work in this area.","headline":"A genuinely large and useful dataset expansion, but the core accuracy claim rests on an unvalidated scale-recovery step and the excerpt lacks independent verification.","tokens_in":4315,"tokens_out":2910,"would_cite":true,"duration_ms":27716,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An automatic annotation pipeline turns 80.8K internet videos into 19.5M 3D whole-body pose annotations with text and audio, and training on the data improves motion generation and recovery.","keywords":["3D whole-body motion","human motion dataset","SMPL-X","text-driven motion generation","audio-driven motion generation","mesh recovery","whole-body pose estimation","multimodal annotation pipeline"],"falsifier":"Take a random sample of roughly 500 Motion-X++ clips, have annotators manually fit the whole-body model to the frames (or capture a subset with multi-view or motion-capture systems), and compare per-joint errors on hands, face, and global trajectory; if the median errors approach the typical inter-annotator variation or exceed the tolerances used in the paper's evaluation, the ground-truth claim is falsified.","tokens_in":3395,"feed_emoji":"🕺","tokens_out":7799,"duration_ms":66318,"temperature":0.7,"pith_summary":"The paper claims that an automatic annotation pipeline can convert unconstrained internet RGB videos into expressive 3D whole-body human motion with paired text and audio, at a scale not possible with manual labeling. The result is Motion-X++, with 19.5M 3D whole-body pose annotations across 120.5K sequences, 80.8K RGB videos, 45.3K audios, 19.5M frame-level pose descriptions, and 120.5K sequence-level semantic labels. The authors argue that this scale and multimodality improve text-driven and audio-driven motion generation, whole-body mesh recovery, and 2D keypoint estimation. If the pipeline's annotations are accurate enough to serve as ground truth, the dataset removes a key bottleneck in human-motion learning.","feed_headline":"80.8K videos yield 19.5M whole-body motion labels","feed_subtitle":"New dataset pairs expressive 3D poses with text and audio to improve motion generation and mesh recovery.","key_machinery":"The central object is the annotation pipeline, whose load-bearing stages are: whole-body keypoint estimation; camera tracking via dense bundle adjustment with the human masked out of feature extraction, correspondence fields, and flow; global trajectory optimization with reprojection, smoothness, ground-contact, and foot-skating losses; and text synthesis that turns estimated body-part spatial relations, hand shapes, and classifier-based emotions into frame-level descriptions, with a vision-language model providing sequence-level captions. All motion is represented in the SMPL-X whole-body parametric model, so body, hands, and face share one optimization. The pipeline's role is to convert raw RGB video into self-supervised training labels at scale.","core_discovery":"Motion-X++ asserts that every part of a human motion, including body pose, hand gestures, facial expressions, and the global trajectory, can be recovered from ordinary single-view or multi-view video by a staged pipeline, and that the recovered parameters are accurate enough to train and evaluate downstream models. The pipeline estimates whole-body 2D keypoints, tracks the camera with a masked dense-bundle-adjustment SLAM that excludes the moving person, optimizes the global trajectory with reprojection, smoothness, ground-contact, and foot-skating losses, fits SMPL-X parameters to the observations, and then generates frame-level text descriptions from body-part relations, hand gestures, and an emotion classifier, while a vision-language model supplies sequence-level captions. The paper presents the resulting dataset as ground truth and reports that training on it improves the four downstream task families.","pith_inferences":["If the accuracy claim holds, Motion-X++ could also serve as pretraining data for video-to-motion and motion-to-video generation, since raw video and audio are stored alongside the motion.","The vision-language captions are likely noisier than the geometric pose labels, so a fair text-to-motion benchmark should include human judgments of caption–pose alignment before treating the captions as ground truth.","The strategy of masking the moving person out of dense-bundle-adjustment camera tracking could transfer to other dynamic-scene reconstruction problems, such as tracking cameras in crowds or sports footage.","Because audio, video, pose, and text are aligned for the same sequences, the dataset opens cross-modal alignment research that motion-only datasets cannot support."],"forward_implications":["Text-driven whole-body motion generation can be trained on 120.5K sequence-level captions paired with expressive 3D pose, including hands and face.","Audio-driven motion generation gains 45.3K audio–motion pairs, enabling models that synthesize dance or performance from music.","Whole-body mesh recovery and 2D keypoint estimation can be trained on diverse internet scenes rather than lab captures.","Because the annotation pipeline runs on any RGB video collection, the dataset can keep growing without manual text labeling.","Frame-level whole-body pose descriptions give models dense language supervision instead of only sequence-level captions."],"supporting_citations":[{"why":"defines the SMPL-X whole-body parametric model that all motion annotations are expressed in.","marker":"[32]"},{"why":"supplies the dual-mask dynamic-object camera tracking approach that the pipeline adapts.","marker":"[85]"},{"why":"provides segmentation masks used to exclude the moving human from camera tracking.","marker":"[86]"},{"why":"gives the global trajectory optimization formulation the pipeline follows.","marker":"[44]"},{"why":"supplies the robust Geman-McClure loss used in trajectory reprojection.","marker":"[87]"},{"why":"provides ground-contact detection used to impose foot-skating constraints.","marker":"[88]"}],"fun_headline_variants":["19.5M whole-body poses from 80.8K videos","Whole-body motion: faces, hands, and text at scale","Motion-X++: 120K sequences, 19.5M labels","Automated pipeline yields multimodal 3D motion","81K text-motion pairs, 19.5M pose annotations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automatic annotation pipeline, including SMPL-X fitting, masked dense-bundle-adjustment camera tracking, trajectory optimization, and vision-language captioning, produces 3D poses and text descriptions accurate enough to serve as ground truth for unconstrained internet videos.","fun_headline_variants_meta":{"raw":{"variants":["19.5M whole-body poses from 80.8K videos","Whole-body motion: faces, hands, and text at scale","Motion-X++: 120K sequences, 19.5M labels","Automated pipeline yields multimodal 3D motion","81K text-motion pairs, 19.5M pose annotations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1511,"prompt_tokens":965,"completion_tokens":546,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":456}},"tokens_in":581,"tokens_out":546,"duration_ms":5366,"temperature":1.0,"reasoning_tokens":456,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:18:39.413735+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of roughly 500 Motion-X++ clips, have annotators manually fit the whole-body model to the frames (or capture a subset with multi-view or motion-capture systems), and compare per-joint errors on hands, face, and global trajectory; if the median errors approach the typical inter-annotator variation or exceed the tolerances used in the paper's evaluation, the ground-truth claim is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the SMPL-X whole-body parametric model that all motion annotations are expressed in."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"gives the global trajectory optimization formulation the pipeline follows."},{"cited_title":"Geman and D","cited_arxiv_id":null,"evidence_quote":"supplies the robust Geman-McClure loss used in trajectory reprojection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides ground-contact detection used to impose foot-skating constraints."}],"review_version":1}