Pith. sign in

REVIEW 20 cited by

OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.08008 v2 pith:4UFXRBF7 submitted 2018-12-18 cs.CV

classification cs.CV
keywords bodyrealtimepartposeaccuracyestimationfootimage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Realtime multi-person 2D pose estimation is a key component in enabling machines to have an understanding of people in images and videos. In this work, we present a realtime approach to detect the 2D pose of multiple people in an image. The proposed method uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. This bottom-up system achieves high accuracy and realtime performance, regardless of the number of people in the image. In previous work, PAFs and body part location estimation were refined simultaneously across training stages. We demonstrate that a PAF-only refinement rather than both PAF and body part location refinement results in a substantial increase in both runtime performance and accuracy. We also present the first combined body and foot keypoint detector, based on an internal annotated foot dataset that we have publicly released. We show that the combined detector not only reduces the inference time compared to running them sequentially, but also maintains the accuracy of each component individually. This work has culminated in the release of OpenPose, the first open-source realtime system for multi-person 2D pose detection, including body, foot, hand, and facial keypoints.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  2. VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    VideoMDM learns coherent 3D motion manifolds from 2D supervision alone by using a pretrained lifter as noisy teacher, depth-weighted 2D reprojection loss, and adapted regularizers, nearly matching fully 3D-supervised ...

  3. Deep Learning Pose Estimation for Multi-Label Recognition of Combined Hyperkinetic Movement Disorders

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A pose-based machine-learning pipeline detects and jointly labels eight hyperkinetic movement disorder signs from routine clinical videos, achieving patient-level macro-AUPRC 0.821 and 86% correct per-label decisions ...

  4. Learning Object-Action Relations from Bimanual Human Demonstration Using Graph Networks

    cs.RO 2019-08 conditional novelty 6.0 of 10

    A graph network classifier trained on scene graphs of symbolic spatial relations achieves per-hand action recognition (macro F1 0.86 top-3) on a new bimanual RGB-D dataset.

  5. Multi-timescale Trajectory Prediction for Abnormal Human Activity Detection

    cs.CV 2019-08 conditional novelty 6.0 of 10

    Multi-timescale pose prediction, trained with supervision at intermediate layers, detects both short-term and long-term human anomalies and outperforms prior pose-based detectors on HR-ShanghaiTech and HR-Avenue.

  6. Zero-Shot Sign Language Recognition: Can Textual Data Uncover Sign Languages?

    cs.CV 2019-07 unverdicted novelty 6.0 of 10

    Introduces ZSSLR problem and ASL-Text dataset; uses text embeddings with 3D-CNN and bi-LSTM video features for zero-shot sign recognition.

  7. Linking Art through Human Poses

    cs.CV 2019-07 unverdicted novelty 6.0 of 10

    Human pose similarity matching with spatial verification outperforms standard content-based image retrieval for discovering composition transfers in art on a manually annotated dataset.

  8. Weight and Height Estimation from a Single Human Image Captured in the Wild

    cs.CV 2026-07 reject novelty 5.0 of 10

    Full-body images and multi-task regression improve CNN-based BMI/weight/height estimation on a new self-reported in-the-wild dataset, though no code, data, or trivial baseline is provided.

  9. Classifying Simulated Gait Impairments using Privacy-preserving Explainable Artificial Intelligence and Mobile Phone Videos

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A mobile-phone pose-estimation pipeline with classic ML classifiers distinguishes seven simulated gait patterns at 86.5% accuracy on a new dataset, pending clinical validation.

  10. Learning to gesticulate by observation using a deep generative approach

    cs.RO 2019-09 conditional novelty 5.0 of 10

    A GAN trained on human Kinect motion capture generates talking gestures for a Pepper humanoid robot using direct kinematic mapping and glove-based hand state recognition.

  11. Safe physical HRI: Toward a unified treatment of speed and separation monitoring together with power and force limiting

    cs.RO 2019-08 conditional novelty 5.0 of 10

    A collaborative robot that switches between full speed, reduced speed, and stop based on body-part-specific distance thresholds completes a mock task faster than zone-monitoring alternatives, with the head-only stop t...

  12. Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review

    eess.IV 2025-02 conditional novelty 4.0 of 10

    A survey and benchmark that groups unpaired image-to-image translation tasks into fully, partially, and non-content preserving categories, and evaluates six models on a vehicle-focused Sim2Real benchmark.

  13. A Wearable Gait Monitoring System for 17 Gait Parameters Based on Computer Vision

    eess.SY 2024-11 conditional novelty 4.0 of 10

    A shoe-mounted stereo-camera and pressure-sensor system measures 17 gait parameters with reported accuracy above 93.61% and identifies individuals from gait sequences with 95.7% accuracy.

  14. Video Affective Effects Prediction with Multi-modal Fusion and Shot-Long Temporal Context

    cs.CV 2019-09 reject novelty 4.0 of 10

    A multi-modal two-time-scale framework with progressive training is claimed to beat prior work on LIRIS-ACCEDE, but the results rely on test-set parameter selection and do not improve arousal MSE.

  15. In defense of OSVOS

    cs.CV 2019-08 reject novelty 4.0 of 10

    Auxiliary video losses help an under-trained OSVOS on DAVIS-2016, but the gains are small and the comparison setting is non-standard.

  16. Fitting, Comparison, and Alignment of Trajectories on Positive Semi-Definite Matrices with Application to Action Recognition

    cs.CV 2019-08 conditional novelty 4.0 of 10

    A skeleton-only action recognition method on PSD matrices with a quotient metric, Bezier curve fitting, and global alignment kernel achieves 97.99% on UTKinect, 96.16% on KTH, and 92.44% on UAV-Gesture.

  17. Gesture Recognition in RGB Videos UsingHuman Body Keypoints and Dynamic Time Warping

    cs.CV 2019-06 unverdicted novelty 4.0 of 10

    A pipeline using OpenPose keypoints and DTW-1NN classifies gestures in RGB videos with flexibility to add new classes via few examples.

  18. A Machine Learning Framework for Real-Time Personalized Ergonomic Pose Analysis

    cs.CV 2026-06 unverdicted novelty 3.0 of 10

    A framework for real-time ergonomic pose prediction from 3D volumetric video that trains personalized classifiers on user-labeled poses captured by RGB-D cameras.

  19. Stereo Vision-Based Fall Prediction and Detection using Human Pose Estimation on the AMD Kria K26 SOM

    cs.CV 2026-06 unverdicted novelty 3.0 of 10

    An applied system combines quantized YOLOX detection, A2J pose estimation, and a CNN classifier on an AMD Kria K26 SOM with a RealSense D455 camera, reporting component accuracies of 74%, 84.13%, and 75.85% and throug...

  20. Sequence-to-Sequence Natural Language to Humanoid Robot Sign Language

    cs.RO 2019-07 unverdicted novelty 3.0 of 10

    Applies established seq2seq neural networks to convert text to Spanish sign language for humanoid robot TEO, proposing OpenPose for skeleton data collection to handle sequence length differences and non-manual markers.

Pith tools