Pith. sign in

REVIEW 8 cited by

RTMW: Real-Time Multi-Person 2D and 3D Whole-body Pose Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08634 v1 pith:BP7HB2US submitted 2024-07-11 cs.CV

classification cs.CV
keywords poseestimationwhole-bodyrtmwbodymodelmodelsapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Whole-body pose estimation is a challenging task that requires simultaneous prediction of keypoints for the body, hands, face, and feet. Whole-body pose estimation aims to predict fine-grained pose information for the human body, including the face, torso, hands, and feet, which plays an important role in the study of human-centric perception and generation and in various applications. In this work, we present RTMW (Real-Time Multi-person Whole-body pose estimation models), a series of high-performance models for 2D/3D whole-body pose estimation. We incorporate RTMPose model architecture with FPN and HEM (Hierarchical Encoding Module) to better capture pose information from different body parts with various scales. The model is trained with a rich collection of open-source human keypoint datasets with manually aligned annotations and further enhanced via a two-stage distillation strategy. RTMW demonstrates strong performance on multiple whole-body pose estimation benchmarks while maintaining high inference efficiency and deployment friendliness. We release three sizes: m/l/x, with RTMW-l achieving a 70.2 mAP on the COCO-Wholebody benchmark, making it the first open-source model to exceed 70 mAP on this benchmark. Meanwhile, we explored the performance of RTMW in the task of 3D whole-body pose estimation, conducting image-based monocular 3D whole-body pose estimation in a coordinate classification manner. We hope this work can benefit both academic research and industrial applications. The code and models have been made publicly available at: https://github.com/open-mmlab/mmpose/tree/main/projects/rtmpose

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Wearing A Coat: Dual-Arm Robot-Assisted Dressing with Differentiable Clothing Simulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Explicit high-order-bias differentiable cloth simulation plus multi-stage dual-arm MPC with local compensation enables successful full-coat dressing under contact constraints on real humans and dummies.

  2. AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    AnyMo is a masked-modeling framework for any-modality human motion generation trained on the new OmniHuMo dataset of 5,000+ hours of multimodal motion sequences.

  3. DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    DanceHMR uses residual body-hand fusion and close-up-aware augmentation to achieve temporally stable whole-body SMPL-X mesh recovery from monocular videos with improved hand accuracy.

  4. MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Using a 3D foundation model to produce viewpoint-aware anchors plus multi-view reference textures enables realistic human-object-interaction reenactment with large out-of-plane rotations.

  5. SegSLR: Promptable Video Segmentation for Isolated Sign Language Recognition

    cs.CV 2025-09 conditional novelty 6.0 of 10

    SegSLR uses pose-guided SAM 2 video segmentations to focus RGB streams on the signer's body and hands, improving isolated sign language recognition on ChaLearn249 IsoGD.

  6. DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    DanceHMR proposes a temporal whole-body mesh recovery method that fuses body context with hand observations for improved hand detail and stability in monocular videos.

  7. DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    DanceHMR uses residual body-hand fusion and close-up-aware augmentation in a single temporal model to achieve stable whole-body SMPL-X recovery with improved hand detail from monocular videos.

  8. Vision-Based Human Awareness Estimation for Enhanced Safety and Efficiency of AMRs in Industrial Warehouses

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    A single-camera pipeline estimates human awareness of AMRs via 3D pose lifting and head orientation to allow adaptive robot motion in mixed human-robot warehouses.

Pith tools