Pith. sign in

REVIEW 12 cited by

OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.08008 v2 pith:4UFXRBF7 submitted 2018-12-18 cs.CV

classification cs.CV
keywords bodyrealtimepartposeaccuracyestimationfootimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Realtime multi-person 2D pose estimation is a key component in enabling machines to have an understanding of people in images and videos. In this work, we present a realtime approach to detect the 2D pose of multiple people in an image. The proposed method uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. This bottom-up system achieves high accuracy and realtime performance, regardless of the number of people in the image. In previous work, PAFs and body part location estimation were refined simultaneously across training stages. We demonstrate that a PAF-only refinement rather than both PAF and body part location refinement results in a substantial increase in both runtime performance and accuracy. We also present the first combined body and foot keypoint detector, based on an internal annotated foot dataset that we have publicly released. We show that the combined detector not only reduces the inference time compared to running them sequentially, but also maintains the accuracy of each component individually. This work has culminated in the release of OpenPose, the first open-source realtime system for multi-person 2D pose detection, including body, foot, hand, and facial keypoints.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  2. VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    VideoMDM learns coherent 3D motion manifolds from 2D supervision alone by using a pretrained lifter as noisy teacher, depth-weighted 2D reprojection loss, and adapted regularizers, nearly matching fully 3D-supervised ...

  3. Zero-Shot Sign Language Recognition: Can Textual Data Uncover Sign Languages?

    cs.CV 2019-07 unverdicted novelty 6.0 of 10

    Introduces ZSSLR problem and ASL-Text dataset; uses text embeddings with 3D-CNN and bi-LSTM video features for zero-shot sign recognition.

  4. Linking Art through Human Poses

    cs.CV 2019-07 unverdicted novelty 6.0 of 10

    Human pose similarity matching with spatial verification outperforms standard content-based image retrieval for discovering composition transfers in art on a manually annotated dataset.

  5. Weight and Height Estimation from a Single Human Image Captured in the Wild

    cs.CV 2026-07 reject novelty 5.0 of 10

    Full-body images and multi-task regression improve CNN-based BMI/weight/height estimation on a new self-reported in-the-wild dataset, though no code, data, or trivial baseline is provided.

  6. Classifying Simulated Gait Impairments using Privacy-preserving Explainable Artificial Intelligence and Mobile Phone Videos

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A mobile-phone pose-estimation pipeline with classic ML classifiers distinguishes seven simulated gait patterns at 86.5% accuracy on a new dataset, pending clinical validation.

  7. Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review

    eess.IV 2025-02 conditional novelty 4.0 of 10

    A survey and benchmark that groups unpaired image-to-image translation tasks into fully, partially, and non-content preserving categories, and evaluates six models on a vehicle-focused Sim2Real benchmark.

  8. A Wearable Gait Monitoring System for 17 Gait Parameters Based on Computer Vision

    eess.SY 2024-11 conditional novelty 4.0 of 10

    A shoe-mounted stereo-camera and pressure-sensor system measures 17 gait parameters with reported accuracy above 93.61% and identifies individuals from gait sequences with 95.7% accuracy.

  9. Gesture Recognition in RGB Videos UsingHuman Body Keypoints and Dynamic Time Warping

    cs.CV 2019-06 unverdicted novelty 4.0 of 10

    A pipeline using OpenPose keypoints and DTW-1NN classifies gestures in RGB videos with flexibility to add new classes via few examples.

  10. A Machine Learning Framework for Real-Time Personalized Ergonomic Pose Analysis

    cs.CV 2026-06 unverdicted novelty 3.0 of 10

    A framework for real-time ergonomic pose prediction from 3D volumetric video that trains personalized classifiers on user-labeled poses captured by RGB-D cameras.

  11. Stereo Vision-Based Fall Prediction and Detection using Human Pose Estimation on the AMD Kria K26 SOM

    cs.CV 2026-06 unverdicted novelty 3.0 of 10

    An applied system combines quantized YOLOX detection, A2J pose estimation, and a CNN classifier on an AMD Kria K26 SOM with a RealSense D455 camera, reporting component accuracies of 74%, 84.13%, and 75.85% and throug...

  12. Sequence-to-Sequence Natural Language to Humanoid Robot Sign Language

    cs.RO 2019-07 unverdicted novelty 3.0 of 10

    Applies established seq2seq neural networks to convert text to Spanish sign language for humanoid robot TEO, proposing OpenPose for skeleton data collection to handle sequence length differences and non-manual markers.

Pith tools