Pith. sign in

REVIEW 3 cited by

SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09782 v1 pith:T5D7W6OQ submitted 2025-01-16 cs.CV cs.GRcs.HCcs.MMcs.RO

SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation

classification cs.CV cs.GRcs.HCcs.MMcs.RO
keywords ehpsscalingmodelmodelsdatadatasetsfoundationsmplest-x
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods focus on training innovative architectural designs on confined datasets. In this work, we investigate the impact of scaling up EHPS towards a family of generalist foundation models. 1) For data scaling, we perform a systematic investigation on 40 EHPS datasets, encompassing a wide range of scenarios that a model trained on any single dataset cannot handle. More importantly, capitalizing on insights obtained from the extensive benchmarking process, we optimize our training scheme and select datasets that lead to a significant leap in EHPS capabilities. Ultimately, we achieve diminishing returns at 10M training instances from diverse data sources. 2) For model scaling, we take advantage of vision transformers (up to ViT-Huge as the backbone) to study the scaling law of model sizes in EHPS. To exclude the influence of algorithmic design, we base our experiments on two minimalist architectures: SMPLer-X, which consists of an intermediate step for hand and face localization, and SMPLest-X, an even simpler version that reduces the network to its bare essentials and highlights significant advances in the capture of articulated hands. With big data and the large model, the foundation models exhibit strong performance across diverse test benchmarks and excellent transferability to even unseen environments. Moreover, our finetuning strategy turns the generalist into specialist models, allowing them to achieve further performance boosts. Notably, our foundation models consistently deliver state-of-the-art results on seven benchmarks such as AGORA, UBody, EgoBody, and our proposed SynHand dataset for comprehensive hand evaluation. (Code is available at: https://github.com/wqyin/SMPLest-X).

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video

    cs.CV 2026-04 unverdicted novelty 7.0

    UniCon3R reconstructs 3D human motion and scene geometry from monocular video by inferring human-scene contacts and using them as an active corrective prior to enforce physical plausibility.

  2. VRGaussianAvatar: Integrating 3D Gaussian Avatars into VR

    cs.CV 2026-02 conditional novelty 7.0

    VRGaussianAvatar enables real-time full-body 3D Gaussian Splatting avatars in VR from HMD tracking alone via inverse kinematics and binocular batching for efficient stereo rendering, outperforming mesh baselines in pe...

  3. UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video

    cs.CV 2026-04 unverdicted novelty 6.0

    UniCon3R infers 4D contact from pose and geometry to correct human mesh and scene alignment in monocular video, yielding more physically plausible joint reconstructions than prior feed-forward methods.