{"id":"0661a839-d963-4027-9c5f-c6ac6406b01d","arxiv_id":"2501.09782","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Training on 40 datasets and larger vision transformers steadily improves whole-body and hand pose estimation, achieving state-of-the-art results on multiple benchmarks.","lead":"This paper studies how scaling up data and model size improves expressive human pose and shape estimation, building generalist models that outperform prior methods on several benchmarks. It introduces a new synthetic dataset called SynHand and a custom generalization metric, MPE.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark-in-training-set issue is real but Table 13 partially addresses it; the more decisive gap is that the central scaling claim is never evaluated on a single benchmark that is completely held out from all 40 training datasets.","rationale":"The reader identified the same core weakness: training on the test benchmarks' training splits weakens the generalist claim. My reading of Table 13 confirms that the paper is aware of and partially controls for this, since R2 vs R3 shows a clear in-domain boost and the R3-R9 trajectory shows out-of-domain gains too. However, I do not think the concern is fully resolved by the existing ablations. The strongest quantitative claim (MPE from 110+ to below 60 mm, hand MPE to 31 mm) is computed over benchmarks that are largely seen at train time, so the headline is not a clean generalization result. The additional 'unseen' benchmarks cited (EHF, ARCTIC, DNA-Rendering) are limited: EHF is a 100-frame single-subject set; ARCTIC and DNA-Rendering-HiRes are progressively included in the 40-dataset pool at the larger scales (ARCTIC appears in the 40-dataset list with 1.5M instances; DNA-Rendering-HiRes appears with 998K instances), so they are not held out for the final H40/H40+ models. The paper's train-set ranking protocol in Appendix C.1 is a real strength: it avoids test-set leakage for dataset selection. But it does not remove the need for a truly external held-out evaluation. My main load-bearing concern is therefore not circularity or dishonesty, but a missing experimental arm: a fully held-out whole-body benchmark. This is a concrete, checkable gap rather than a refutation. Credit is due for the extensive benchmarking table, the proposed MPE metric, the SynHand dataset, and the honest Table 13 and diminishing-returns analysis. The correct disposition remains CONDITIONAL: accept the scaling trend and resource contributions, but require a genuinely held-out evaluation before endorsing the 'generalist foundation model' and the headline error reductions as evidence of cross-domain generalization.","tokens_in":47518,"tokens_out":2045,"duration_ms":19791,"concrete_test":"Retrain SMPLest-X-H40+ (or reuse the released checkpoints) and evaluate on a held-out whole-body benchmark with a training split not included in the 40 datasets, e.g., EMDB test, or the validation split of a recently released dataset such as HuMMan or GTA-Human-II, while keeping the inference protocol identical. Report PVE, PA-PVE, and hand PVE. If the held-out performance remains far better than prior SOTA (which was trained on confined datasets), the generalist claim is substantiated; if the advantage shrinks to a few mm, the headline MPE reduction is mostly attributable to in-domain training leakage, and the verdict should be CONDITIONAL on re-scoped claims.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that scaling to 40 datasets / 10M instances produces generalist foundation models that transfer to unseen environments. The results supporting this claim (Tables 3-5) are measured on AGORA-val, UBody, EgoBody, 3DPW, and EHF. Table 13 shows the training sets of AGORA, UBody, EgoBody, and 3DPW are all included among the 40 training datasets (rows R6-R9 list #Seen=5 or 6; the 'seen' benchmark train splits include AGORA, UBody, EgoBody, 3DPW, and by R8 likely ARCTIC or DNA-R too). EHF is test-only and thus genuinely unseen, which supports some cross-dataset transfer claim. However, the headline numbers in Table 5 (MPE 58.9, Hand-MPE 31.6) average over four in-domain-train benchmarks plus EHF, so the 'generalist' claim rests substantially on in-domain training. The paper itself acknowledges in Section 3.4 that using test-set rankings leaks information and constructs a train-set benchmark in Appendix C.1 for dataset selection, which is good practice. The unaddressed gap is that no evaluation is reported on a whole-body benchmark whose train split is absent from the 40-dataset pool; EHF (one subject, 100 frames) and DNA-Rendering-HiRes (seen at 40-dataset scale, see Table 12 text) are too narrow or not held out. Therefore the claim that large-scale data mixing yields a true generalist is not crisply tested, and the relative contribution of in-domain training versus genuine cross-domain generalization remains entangled in the MPE/Hand-MPE headline numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript investigates data- and model-scaling for expressive human pose and shape estimation (EHPS). The authors assemble a large pool of 40 datasets (up to 10M instances), introduce a new synthetic hand-focused benchmark called SynHand, define a mean-primary-error (MPE) metric for whole-body and hand evaluation, and train two deliberately simple architectures (SMPLer-X and SMPLest-X) with ViT backbones up to ViT-Huge. They report monotonic improvements with more data and larger models, specialist fine-tuning results on AGORA, UBody, EgoBody, and a claim of strong transfer to unseen environments. The central quantitative claims are in Tables 1-5, with an in-domain-training analysis in Table 13.","tokens_in":47825,"tokens_out":7805,"duration_ms":79464,"significance":"The systematic dataset benchmark, the train-set-based selection protocol in Appendix C.1, the release of code, and the inclusion of Table 13 are real strengths; if the scaling conclusions survive a clean held-out evaluation, this would be a valuable reference for data mixing and model scaling in EHPS. However, the headline MPE and hand-MPE numbers are computed largely on benchmarks whose training splits are present in the training pool, so the central 'generalist / transfer to unseen environments' claim is not yet crisply demonstrated. The paper is honest about the in-domain effect in Table 13, but that analysis also exposes the gap between the paper's claims and its evaluation protocol.","major_comments":[{"comment":"The headline MPE and Hand-MPE in Table 5 are averaged over AGORA-val, UBody, EgoBody, 3DPW, and EHF, with ARCTIC and SynHand added for the hand metrics. Table 13 (rows R6-R9) shows that AGORA, UBody, EgoBody, and 3DPW train splits are seen during training, and from R7 onward ARCTIC and DNA-Rendering-HiRes are also seen, leaving only EHF (one subject, 100 curated frames, Section 3.3) as a fully held-out benchmark at the final 40-dataset scale. The central claim of transferability to unseen environments is therefore not crisply tested, and part of the reported improvement over prior methods reflects in-domain training. Please add a genuinely held-out whole-body benchmark that is absent from all training configurations, or at minimum report seen and unseen benchmark averages separately and state explicitly which benchmark train splits are included in each training pool.","section":"Sec. 5.1.3, Tables 5 and 13"},{"comment":"The newly proposed SynHand dataset appears both as a benchmarked data source in Table 1 (462.8K instances) and as an evaluation benchmark in Table 4, and the hand-MPE in Table 5 averages over SynHand and ARCTIC. The manuscript does not state explicitly whether the SynHand training split is included in the final 40-dataset training pool, nor whether the ARCTIC train split is seen by the H40/H40+ models. If either is included, the hand-MPE is partly in-domain by construction. Please clarify these memberships and, where possible, recompute hand-MPE excluding any benchmark whose training split was used, or report a fully held-out hand benchmark separately.","section":"Sec. 3.5, Tables 1 and 4"},{"comment":"The proposed MPE and Hand-MPE are averages over heterogeneous primary metrics (MPJPE for 3DPW, PVE elsewhere) and are not validated against established aggregate protocols or leaderboard orderings. No uncertainty estimates or multiple seeds are reported, and several entries are non-monotonic, e.g., SMPLest-X-H40 vs H40+ in Table 3/Table 4 (MPE 58.9 vs 59.9; Hand-MPE 31.6 vs 32.7), and ARCTIC hand-PA-PVE in Table 4 worsens monotonically from S5 to S40 (16.7 to 19.2). These patterns make the scaling-law claims hard to separate from noise. Please report per-benchmark results for every scale and backbone, add multiple seeds or confidence intervals, and show that MPE ordering agrees with standard benchmark ordering before using it as the main summary statistic.","section":"Sec. 3.1, Tables 3-5"},{"comment":"The balanced sampling protocol resamples all selected datasets to the same length, so the effective per-dataset instance count and upsampling factor change as the dataset count grows (e.g., 0.75M/5 = 150K per dataset for the 5-dataset setting vs 10M/40 = 250K per dataset for the 40-dataset setting). The comparison across #Datasets therefore changes both dataset diversity and per-dataset repetition/epoch exposure simultaneously, and the reported 'scaling law' may partly reflect training intensity rather than dataset count alone. Please report the total epochs, per-dataset repetition factors, and effective number of unique instances per dataset for each row in Tables 3 and 4.","section":"Sec. 5.1.1, Tables 3 and 4"}],"minor_comments":[{"comment":"The caption contains the fragment 'Real/eot', which appears to be a typo for 'Real/Photo' or a similar label; please correct it.","section":"Fig. 2 caption"},{"comment":"The text refers to 'RenBody' in some places and 'DNA-Rendering' elsewhere, and Table 22 is titled 'DNA-Rendering-HiRes' while the appendix text uses 'RenBody-HiRes'; please standardize the dataset name throughout.","section":"Appendix B and Table 22"},{"comment":"The sentence 'Ee use the annotations generated by EFT' contains a typo and should read 'We use the annotations generated by EFT'.","section":"Appendix D, PoseTrack paragraph"},{"comment":"Table 1 lists 41 rows, while the text repeatedly says 40 datasets; please clarify whether SynHand is the 40th dataset or whether one row is counted differently, and state explicitly whether SynHand is part of the 40-dataset training pool.","section":"Table 1 vs Section 3.2"},{"comment":"The sentence 'We adopt the official train and test split provided for InterHand datasets' is ambiguous because SynHand combines hand poses from InterHand, AGORA, and GRAB with body poses from AMASS; please state explicitly how the official InterHand split is mapped onto SynHand images and whether the AGORA-derived hand poses overlap with AGORA train images.","section":"Sec. 3.5, SynHand description"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious large-scale empirical study and the authors have been transparent in some respects, especially with Table 13 and the train-set-based dataset selection in Appendix C.1. The main stumbling block is that the headline generalist claim is evaluated on benchmarks whose training splits are in the training pool, so the contribution of true cross-domain generalization versus in-domain training remains entangled. If the authors add a clean held-out benchmark, separate seen/unseen reporting, and uncertainty quantification, this paper could become a strong candidate for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a serious, large-scale empirical study that will likely be influential for the EHPS subfield, but the headline 'generalist foundation model' claim is not actually tested on a single fully held-out whole-body benchmark. The paper does a lot right: it assembles 40 datasets, trains SMPLer-X and SMPLest-X at up to ViT-Huge, and systematically shows that more data and larger models reduce mean primary error across five benchmarks, with diminishing returns at 10M instances. That scaling trend is visible in the tables and is the paper's real contribution. The new SynHand dataset and the MPE metric are useful resources, though MPE is just a mean of primary errors and is not validated as a holistic metric.\n\nThe soft spots are real and concentrated in the evaluation design. The training pool includes the training splits of AGORA, UBody, EgoBody, and 3DPW, and Table 13 shows that in-domain training gives a large boost. So the reported gains over prior methods on those benchmarks partly reflect having seen the target domain, not pure cross-domain generalization. The paper is aware of this and constructs a train-set-based selection benchmark to avoid test-set leakage, which is good practice. But the 'generalist transfers to unseen environments' claim rests on EHF (100 frames, one subject) and ARCTIC/DNA-Rendering, which are either tiny or not fully held out (DNA-Rendering appears in the 40-dataset pool at the 40-dataset scale). The hand-MPE metric includes SynHand, which is also a training dataset, creating a mild circularity for that specific claim. I'd also like error bars; some hand metrics are non-monotonic across data scale, which the paper doesn't address.\n\nNone of this looks like sloppiness or bad faith. The paper is transparent about its limitations and its diminishing-returns discussion is honest. The central empirical observation—that scaling data and model size helps within this protocol—probably holds. But the 'generalist foundation model' framing is oversold relative to the evidence, and a referee should ask for evaluation on a benchmark whose training split is absent from the 40-dataset pool.\n\nWho is this for? Anyone working on whole-body or hand pose estimation, and anyone interested in scaling laws for perception tasks. It deserves a serious referee and, with a held-out evaluation added, would be a solid publication.\n\nMy take: engage with it, but read the tables carefully and don't take the 'unseen' transfer claims at face value.","headline":"Big, useful data-scaling study for EHPS, but the 'generalist' claim is not tested on a truly held-out whole-body benchmark.","tokens_in":48410,"tokens_out":1908,"would_cite":true,"duration_ms":21302,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By training two deliberately simple vision-transformer architectures on 40 datasets and up to 10 million images, this paper shows that scaling data and model size alone produces state-of-the-art expressive human pose and shape estimation…","keywords":["expressive human pose and shape estimation","SMPL-X","data scaling","model scaling","vision transformer","foundation models","hand pose estimation","synthetic dataset"],"falsifier":"Retrain the identical SMPLest-X-H40 recipe on the same 40 datasets with the training splits of every evaluation benchmark that has one (AGORA, UBody, EgoBody, 3DPW, ARCTIC, and DNA-Rendering) removed, then compare mean primary error on those test sets against the paper's numbers; if the margin over prior single-domain methods shrinks to a few millimeters, the central generalization claim is largely an in-domain training effect, whereas a large surviving margin would confirm genuine scaling-driven transfer.","tokens_in":1996,"feed_emoji":"🧍","tokens_out":3337,"duration_ms":109251,"temperature":0.7,"pith_summary":"This paper argues that expressive human pose and shape estimation, which recovers body, hands, and face as a single parametric mesh, is limited less by network design than by the scale of training data and model. It assembles 40 datasets totaling up to 10 million instances, benchmarks each dataset on five test sets, and trains two minimalist architectures, SMPLer-X and the simpler SMPLest-X, using backbones from ViT-Small to ViT-Huge. The result is a family of generalist models that lowers mean whole-body error across five benchmarks from above 110 mm to below 60 mm and hand error from above 62 mm to 31 mm, while also transferring to benchmarks not seen in training and adapting through finetuning into specialists. The paper also introduces SynHand, a synthetic whole-body dataset biased toward complex hand poses, and a mean primary error (MPE) metric for gauging generalization, and it reports diminishing returns past roughly 7.5 to 10 million training instances within the same data domain.","feed_headline":"One recipe cuts whole-body pose error from 110 to 60 mm","feed_subtitle":"Training a plain vision transformer on 40 datasets and 10M images beats prior state-of-the-art on seven benchmarks.","key_machinery":"The carrying object is a minimal encoder-decoder pipeline: a Vision Transformer that turns a 512 by 384 crop into image tokens, plus, in SMPLest-X, 80 learnable task tokens (30 for hand poses, the rest for body pose, face expression, orientations, and translation) that are concatenated with the image tokens and read out through a six-layer transformer decoder. SMPLer-X instead keeps a component-guiding module that predicts hand and face bounding boxes and crops features before regression. The decisive comparison is that removing that module, making the network one-stage and letting task tokens attend to fingers on their own, cuts hand error by roughly 13 to 15 percent, which the paper attributes to avoiding the information loss and error accumulation of feature cropping. Scaling operates through two axes this design is built for: dataset count and instance count under balanced sampling, and backbone size from ViT-Small to ViT-Huge, with the newly proposed MPE and hand-PA-MPE metrics used to read the scaling curves.","core_discovery":"The paper's claim is that expressive human pose and shape estimation, recovering body, hands, and face as a single SMPL-X mesh from one image, is bottlenecked less by algorithmic design than by training scale. Using two deliberately minimal architectures, SMPLer-X (a ViT encoder with a hand-and-face guiding module plus regression heads) and SMPLest-X (the same encoder with a bare transformer decoder and task tokens), and training on 40 datasets totaling up to 10 million instances, the authors report mean whole-body primary error across AGORA, UBody, EgoBody, 3DPW, and EHF falling from over 110 mm for prior methods to 58.9 mm, and hand error across six hand benchmarks from over 62 mm to 31.6 mm. They further report that the one-stage SMPLest-X beats the two-stage SMPLer-X on hands by roughly 13 to 15 percent, that results transfer to unseen benchmarks such as ARCTIC and DNA-Rendering, that finetuning the generalist sets new records including the first 96.2 mm NMVE on the AGORA leaderboard, and that growing instances from 7.5M to 10M within the same 40-dataset pool yields diminishing returns, which they take as evidence that scaling has reached a saturation point and algorithmic work should resume on top of large diverse pretraining.","pith_inferences":["The paper's own Table 13 (rows R2 versus R3) implies that part of the reported advantage over prior methods comes from seeing each benchmark's training split; a fair head-to-head would hold out all evaluation-benchmark training data and measure how much of the gain is genuine cross-domain transfer rather than in-domain training.","The saturation observed at 10M instances is a statement about the current minimalist architectures; a richer decoder or explicit hand-focusing supervision could shift the saturation point, so the right reading may be that architecture now matters more, not that data no longer matters.","SynHand's deliberate bias toward non-relaxed hand poses could serve as a diagnostic: if a model trained without SynHand degrades selectively on complex-hand images, that would confirm that standard whole-body datasets underspecify hand articulation, a claim about data composition rather than scale.","Balanced sampling treats all 40 datasets equally despite large quality differences; quality-weighted sampling might reach the same MPE with fewer total instances, an experiment hinted at by the paper's weighted-strategy results but not pursued."],"forward_implications":["Whole-body pose and shape estimation follows a scaling law: increasing dataset count and model capacity monotonically reduces mean primary error until roughly 10M training instances, after which returns diminish.","A one-stage design without hand-specific cropping outperforms a two-stage design with explicit component localization on hand pose, suggesting that component guidance is not required for accurate articulations.","Generalists trained on many datasets beat models trained only on a benchmark's own training split, and finetuning a generalist into a specialist sets new state-of-the-art numbers on AGORA, UBody, and EgoBody.","The data-scaling conclusion holds for CNN architectures as well: scaling a Hand4Whole-style model from 0.65M to 4.5 to 5.6 million instances cuts its MPE from about 116.6 mm to between 96.9 and 98.3 mm.","Adding instances from 7.5M to 10M within the same 40-dataset pool raises per-epoch compute cost by about 43 percent with only marginal gains, indicating data scale has saturated for the current model family."],"supporting_citations":[{"why":"Defines the SMPL-X parametric output space (body, hands, face) that the entire task and all metrics are built on.","marker":"[1]"},{"why":"OSX is the one-stage ViT baseline whose architecture SMPLest-X simplifies and whose UBody dataset serves as a key benchmark.","marker":"[6]"},{"why":"Supplies the Vision Transformer backbone family whose scalability from Small to Huge carries the model-scaling study.","marker":"[52]"},{"why":"AGORA is the primary SMPL-X benchmark and leaderboard target where the specialists set new records.","marker":"[12]"},{"why":"HybrIK-X is a state-of-the-art baseline whose whole-body and hand errors the foundation models are compared against.","marker":"[15]"},{"why":"Hand4Whole is the CNN-based baseline reused to show that data scaling benefits non-transformer architectures as well.","marker":"[51]"},{"why":"Prior systematic study of data scaling for body-only SMPL estimation that this work extends to hands, face, and expression.","marker":"[58]"},{"why":"EgoBody provides the egocentric, heavily truncated scenario included in the mean primary error basket.","marker":"[5]"},{"why":"3DPW is the in-the-wild body benchmark whose MPJPE contributes to the MPE metric.","marker":"[13]"},{"why":"EHF, a test-only dataset with no training split, is used to demonstrate cross-dataset transfer and hand evaluation.","marker":"[14]"}],"fun_headline_variants":["Scale, not design, breaks pose error records","10M images make plain ViT best at body, hands, face","SMPLest-X: huge data makes minimal model a top scorer","From 110 to 59 mm: scaling wins for whole-body pose","One ViT, 40 datasets: state of the art on 7 benchmarks"],"cache_read_input_tokens":50432,"weakest_assumption_plain":"The headline gains are measured on benchmark test sets whose training splits are inside the model's 40-dataset training pool, and the paper's own Table 13 shows that seeing a benchmark's training split accounts for a large part of the improvement, so the reported margin over prior methods bundles in-domain training together with genuine cross-domain transfer.","fun_headline_variants_meta":{"raw":{"variants":["Scale, not design, breaks pose error records","10M images make plain ViT best at body, hands, face","SMPLest-X: huge data makes minimal model a top scorer","From 110 to 59 mm: scaling wins for whole-body pose","One ViT, 40 datasets: state of the art on 7 benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001007,"raw_usage":{"total_tokens":4368,"prompt_tokens":1164,"completion_tokens":3204,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":780,"completion_tokens_details":{"reasoning_tokens":3111}},"tokens_in":780,"tokens_out":3204,"duration_ms":22008,"temperature":1.0,"reasoning_tokens":3111,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:41:10.459570+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the identical SMPLest-X-H40 recipe on the same 40 datasets with the training splits of every evaluation benchmark that has one (AGORA, UBody, EgoBody, 3DPW, ARCTIC, and DNA-Rendering) removed, then compare mean primary error on those test sets against the paper's numbers; if the margin over prior single-domain methods shrinks to a few millimeters, the central generalization claim is largely an in-domain training effect, whereas a large surviving margin would confirm genuine scaling-driven transfer.","supporting_citations":[{"cited_title":"AGORA: Avatars in geography optimized for regression analysis,","cited_arxiv_id":null,"evidence_quote":"AGORA is the primary SMPL-X benchmark and leaderboard target where the specialists set new records."},{"cited_title":"Expressive body capture: 3d hands, face, and body from a single image,","cited_arxiv_id":null,"evidence_quote":"Defines the SMPL-X parametric output space (body, hands, face) that the entire task and all metrics are built on."},{"cited_title":"One-stage 3d whole-body mesh recovery with component aware transformer,","cited_arxiv_id":null,"evidence_quote":"OSX is the one-stage ViT baseline whose architecture SMPLest-X simplifies and whose UBody dataset serves as a key benchmark."},{"cited_title":"An image is worth 16x16 words: Transformers for image recognition at scale,","cited_arxiv_id":null,"evidence_quote":"Supplies the Vision Transformer backbone family whose scalability from Small to Huge carries the model-scaling study."},{"cited_title":"Accurate 3d hand pose estima- tion for whole-body 3d human mesh estimation,","cited_arxiv_id":null,"evidence_quote":"Hand4Whole is the CNN-based baseline reused to show that data scaling benefits non-transformer architectures as well."},{"cited_title":"Benchmarking and analyzing 3d human pose and shape estimation beyond algorithms,","cited_arxiv_id":null,"evidence_quote":"Prior systematic study of data scaling for body-only SMPL estimation that this work extends to hands, face, and expression."},{"cited_title":"Egobody: Human body shape and motion of interacting people from head-mounted devices,","cited_arxiv_id":null,"evidence_quote":"EgoBody provides the egocentric, heavily truncated scenario included in the mean primary error basket."},{"cited_title":"Recovering accurate 3d human pose in the wild using imus and a moving camera,","cited_arxiv_id":null,"evidence_quote":"3DPW is the in-the-wild body benchmark whose MPJPE contributes to the MPE metric."},{"cited_title":"Expressive body capture: 3d hands, face, and body from a single image,","cited_arxiv_id":null,"evidence_quote":"EHF, a test-only dataset with no training split, is used to demonstrate cross-dataset transfer and hand evaluation."}],"review_version":1}