Pith. sign in

REVIEW 28 cited by

RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07399 v2 pith:GRSVHUYV submitted 2023-03-13 cs.CV

RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

classification cs.CV
keywords estimationposertmposeachievesmmposereal-timecocomodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent studies on 2D pose estimation have achieved excellent performance on public benchmarks, yet its application in the industrial community still suffers from heavy model parameters and high latency. In order to bridge this gap, we empirically explore key factors in pose estimation including paradigm, model architecture, training strategy, and deployment, and present a high-performance real-time multi-person pose estimation framework, RTMPose, based on MMPose. Our RTMPose-m achieves 75.8% AP on COCO with 90+ FPS on an Intel i7-11700 CPU and 430+ FPS on an NVIDIA GTX 1660 Ti GPU, and RTMPose-l achieves 67.0% AP on COCO-WholeBody with 130+ FPS. To further evaluate RTMPose's capability in critical real-time applications, we also report the performance after deploying on the mobile device. Our RTMPose-s achieves 72.2% AP on COCO with 70+ FPS on a Snapdragon 865 chip, outperforming existing open-source libraries. Code and models are released at https://github.com/open-mmlab/mmpose/tree/1.x/projects/rtmpose.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. iMiGUE-3K: A Large-Scale Benchmark for Micro-Gesture Analysis with Self-Supervised Learning

    cs.CV 2026-05 unverdicted novelty 8.0

    iMiGUE-3K is the largest in-the-wild micro-gesture video dataset with 3.4K clips and 37M frames from real interviews, supporting self-supervised foundation models and benchmarks that show micro-gestures improve emotio...

  2. AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computing

    cs.CV 2026-05 unverdicted novelty 7.0

    AIGaitor is the first claimed end-to-end on-device monocular motion-capture and deep-learning gait analysis pipeline demonstrated on consumer smartphones.

  3. AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computing

    cs.CV 2026-05 unverdicted novelty 7.0

    The paper presents AIGaitor, a privacy-preserving on-device monocular motion analysis system that performs end-to-end pose estimation and deep learning gait analysis on consumer smartphones.

  4. PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition

    cs.CV 2026-05 unverdicted novelty 7.0

    PoseBridge recovers semantic information lost during skeletonization by extracting pose-anchored cues from human pose estimation and transferring them via skeleton-conditioned bridging and semantic prototype adaptatio...

  5. OmniRobotHome: A Multi-Camera Platform for Real-Time Multiadic Human-Robot Interaction

    cs.RO 2026-04 unverdicted novelty 7.0

    A 48-camera residential platform delivers real-time occlusion-robust 3D perception and coordinated actuation for multi-human multi-robot interaction in a shared home workspace.

  6. BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognition

    cs.CV 2026-04 unverdicted novelty 7.0

    BarbieGait is a new synthetic gait dataset with identity-consistent cloth changes paired with the GaitCLIF model that improves cross-clothing recognition on the new data and existing benchmarks.

  7. Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

    cs.CV 2026-04 unverdicted novelty 7.0

    The paper presents a multimodal framework, dataset, and reconstruction pipeline to create immersive volumetric videos supporting large 6-DoF audiovisual interaction from real multi-view captures.

  8. LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition

    cs.CV 2026-03 unverdicted novelty 7.0

    LA-Sign achieves state-of-the-art skeleton-based sign language recognition on WLASL and MSASL by using recurrent looped transformers with adaptive hyperbolic geometry alignment.

  9. Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study

    cs.CV 2026-07 conditional novelty 6.5

    Identity-specific free-throw motion signatures exist and are learnable by video models, but models prefer static appearance shortcuts unless silhouettes or skeletons suppress them.

  10. BackTranslation2.0 -- A Linguistically Motivated Metric to Assess Sign Language Production

    cs.CV 2026-06 unverdicted novelty 6.0

    BackTranslation2.0 is a linguistically motivated evaluation metric for sign language production that uses an agentic tool pipeline and LLM cross-referencing to score four dimensions and shows strong human correlation ...

  11. SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation

    cs.CV 2026-06 unverdicted novelty 6.0

    SIGNET uses attention-based aggregation of multiple pretrained sign language backbones and gated fusion to achieve state-of-the-art cross-language sign language translation on How2Sign, Phoenix14T, CSL-Daily, and MeineDGS.

  12. Emotion Recognition in Sign Language Conversation

    cs.CL 2026-05 unverdicted novelty 6.0

    Introduces the eJSL Dialog dataset (1,920 videos in 480 dialogues from STUDIES corpus) for conversational sign language emotion recognition and benchmarks models revealing a domain gap with generic multimodal approaches.

  13. DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 6.0

    DanceHMR uses residual body-hand fusion and close-up-aware augmentation to achieve temporally stable whole-body SMPL-X mesh recovery from monocular videos with improved hand accuracy.

  14. High-Fidelity Single-Image Head Modeling with Industry-Grade Topology

    cs.CV 2026-05 unverdicted novelty 6.0

    A single-image head reconstruction method uses coarse-to-fine optimization with normal consistency, landmarks, and geometry-aware constraints on curvature and conformality to produce meshes with industry-grade topolog...

  15. EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

    cs.CV 2025-09 unverdicted novelty 6.0

    EditVerse unifies image and video editing and generation in one transformer model via unified token sequences and in-context learning, trained jointly on curated video editing data plus image/video corpora and evaluat...

  16. Biomechanics-aware Multi-view Markerless Motion Capture of Dexterous Hand Movements

    cs.CV 2026-07 conditional novelty 5.0

    End-to-end biomechanics-aware optimization recovers plausible multi-joint hand kinematics from multi-view markerless video more robustly than two-stage triangulation-plus-IK, especially under object occlusion.

  17. Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

    cs.CV 2026-06 unverdicted novelty 5.0

    A multi-frame network with SAM 3-derived mask priors achieves 72.4% F1 tip and 58.0% F1 anchor localization in surgical videos without manual mask annotations for training.

  18. SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks

    cs.CV 2026-06 unverdicted novelty 5.0

    The paper releases SignNet-1M, a 1M-scale augmented dataset for ASL, CSL and DGS with 3DGS and diffusion-based variations, plus benchmarks showing improved cross-shift generalization.

  19. VDSB-GWSyn: Diffusion Schr\"{o}dinger Bridge for Controllable and Anatomically Feasible Guidewire Synthesis in Coronary Angiography

    cs.CV 2026-05 unverdicted novelty 5.0

    VDSB-GWSyn uses DSB conditioned on vessel masks and a shape prior to synthesize guidewires, yielding downstream localization gains when used for pre-training.

  20. Emotion Recognition in Sign Language Conversation

    cs.CL 2026-05 unverdicted novelty 5.0

    Isolated-sign emotion models fail in dialogue; eJSL Dialog benchmarks show a domain gap for generic multimodal ERC models on sign language.

  21. OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

    cs.CV 2026-05 accept novelty 5.0

    The OSS Challenge provides benchmarks showing spatiotemporal video models excel at open suturing skill classification and OSATS scoring but struggle with keypoint tracking under occlusion.

  22. DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 5.0

    DanceHMR proposes a temporal whole-body mesh recovery method that fuses body context with hand observations for improved hand detail and stability in monocular videos.

  23. Inferring World Belief States in Dynamic Real-World Environments

    cs.RO 2026-04 unverdicted novelty 5.0

    A robot infers human world belief states from observations in dynamic 3D household environments to enable fluent human-robot teamwork.

  24. Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation

    cs.CV 2025-09 unverdicted novelty 5.0

    Extends online 2D multi-camera tracking to 3D via depth-based point cloud reconstruction, clustering for 3D boxes, and local ID consistency for global data association, placing 3rd on 2025 AI City Challenge 3D MTMC dataset.

  25. SPARK: Low Latency Single-Camera 3D Pose Estimation for Autonomous Racing using Keypoints

    cs.RO 2026-06 unverdicted novelty 4.0

    SPARK applies keypoint detection with YOLO models to monocular images for low-latency 3D pose estimation of racing opponents, claiming better accuracy and speed than prior camera methods on real racing data.

  26. Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models

    cs.CV 2026-06 unverdicted novelty 4.0

    YOLO26 presents a unified real-time vision model family with dual-head end-to-end design, new training components, and task-specific heads that reports improved mAP-latency tradeoffs on COCO and LVIS benchmarks across...

  27. DanceHMR: Hand-Aware Whole-Body Human Mesh Recovery from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 4.0

    DanceHMR uses residual body-hand fusion and close-up-aware augmentation in a single temporal model to achieve stable whole-body SMPL-X recovery with improved hand detail from monocular videos.

  28. World Simulation with Video Foundation Models for Physical AI

    cs.CV 2025-10 unverdicted novelty 4.0

    Cosmos-Predict2.5 unifies text-to-world, image-to-world, and video-to-world generation in one model trained on 200M clips with RL post-training, delivering improved quality and control for physical AI.