REVIEW 20 cited by
OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Realtime multi-person 2D pose estimation is a key component in enabling machines to have an understanding of people in images and videos. In this work, we present a realtime approach to detect the 2D pose of multiple people in an image. The proposed method uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. This bottom-up system achieves high accuracy and realtime performance, regardless of the number of people in the image. In previous work, PAFs and body part location estimation were refined simultaneously across training stages. We demonstrate that a PAF-only refinement rather than both PAF and body part location refinement results in a substantial increase in both runtime performance and accuracy. We also present the first combined body and foot keypoint detector, based on an internal annotated foot dataset that we have publicly released. We show that the combined detector not only reduces the inference time compared to running them sequentially, but also maintains the accuracy of each component individually. This work has culminated in the release of OpenPose, the first open-source realtime system for multi-person 2D pose detection, including body, foot, hand, and facial keypoints.
Forward citations
Cited by 20 Pith papers
-
MultiAnimate: A Unified Framework for Controllable Multi-Character Animation
A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.
-
VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
VideoMDM learns coherent 3D motion manifolds from 2D supervision alone by using a pretrained lifter as noisy teacher, depth-weighted 2D reprojection loss, and adapted regularizers, nearly matching fully 3D-supervised ...
-
Deep Learning Pose Estimation for Multi-Label Recognition of Combined Hyperkinetic Movement Disorders
A pose-based machine-learning pipeline detects and jointly labels eight hyperkinetic movement disorder signs from routine clinical videos, achieving patient-level macro-AUPRC 0.821 and 86% correct per-label decisions ...
-
Learning Object-Action Relations from Bimanual Human Demonstration Using Graph Networks
A graph network classifier trained on scene graphs of symbolic spatial relations achieves per-hand action recognition (macro F1 0.86 top-3) on a new bimanual RGB-D dataset.
-
Multi-timescale Trajectory Prediction for Abnormal Human Activity Detection
Multi-timescale pose prediction, trained with supervision at intermediate layers, detects both short-term and long-term human anomalies and outperforms prior pose-based detectors on HR-ShanghaiTech and HR-Avenue.
-
Zero-Shot Sign Language Recognition: Can Textual Data Uncover Sign Languages?
Introduces ZSSLR problem and ASL-Text dataset; uses text embeddings with 3D-CNN and bi-LSTM video features for zero-shot sign recognition.
-
Linking Art through Human Poses
Human pose similarity matching with spatial verification outperforms standard content-based image retrieval for discovering composition transfers in art on a manually annotated dataset.
-
Weight and Height Estimation from a Single Human Image Captured in the Wild
Full-body images and multi-task regression improve CNN-based BMI/weight/height estimation on a new self-reported in-the-wild dataset, though no code, data, or trivial baseline is provided.
-
Classifying Simulated Gait Impairments using Privacy-preserving Explainable Artificial Intelligence and Mobile Phone Videos
A mobile-phone pose-estimation pipeline with classic ML classifiers distinguishes seven simulated gait patterns at 86.5% accuracy on a new dataset, pending clinical validation.
-
Learning to gesticulate by observation using a deep generative approach
A GAN trained on human Kinect motion capture generates talking gestures for a Pepper humanoid robot using direct kinematic mapping and glove-based hand state recognition.
-
Safe physical HRI: Toward a unified treatment of speed and separation monitoring together with power and force limiting
A collaborative robot that switches between full speed, reduced speed, and stop based on body-part-specific distance thresholds completes a mock task faster than zone-monitoring alternatives, with the head-only stop t...
-
Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review
A survey and benchmark that groups unpaired image-to-image translation tasks into fully, partially, and non-content preserving categories, and evaluates six models on a vehicle-focused Sim2Real benchmark.
-
A Wearable Gait Monitoring System for 17 Gait Parameters Based on Computer Vision
A shoe-mounted stereo-camera and pressure-sensor system measures 17 gait parameters with reported accuracy above 93.61% and identifies individuals from gait sequences with 95.7% accuracy.
-
Video Affective Effects Prediction with Multi-modal Fusion and Shot-Long Temporal Context
A multi-modal two-time-scale framework with progressive training is claimed to beat prior work on LIRIS-ACCEDE, but the results rely on test-set parameter selection and do not improve arousal MSE.
-
In defense of OSVOS
Auxiliary video losses help an under-trained OSVOS on DAVIS-2016, but the gains are small and the comparison setting is non-standard.
-
Fitting, Comparison, and Alignment of Trajectories on Positive Semi-Definite Matrices with Application to Action Recognition
A skeleton-only action recognition method on PSD matrices with a quotient metric, Bezier curve fitting, and global alignment kernel achieves 97.99% on UTKinect, 96.16% on KTH, and 92.44% on UAV-Gesture.
-
Gesture Recognition in RGB Videos UsingHuman Body Keypoints and Dynamic Time Warping
A pipeline using OpenPose keypoints and DTW-1NN classifies gestures in RGB videos with flexibility to add new classes via few examples.
-
A Machine Learning Framework for Real-Time Personalized Ergonomic Pose Analysis
A framework for real-time ergonomic pose prediction from 3D volumetric video that trains personalized classifiers on user-labeled poses captured by RGB-D cameras.
-
Stereo Vision-Based Fall Prediction and Detection using Human Pose Estimation on the AMD Kria K26 SOM
An applied system combines quantized YOLOX detection, A2J pose estimation, and a CNN classifier on an AMD Kria K26 SOM with a RealSense D455 camera, reporting component accuracies of 74%, 84.13%, and 75.85% and throug...
-
Sequence-to-Sequence Natural Language to Humanoid Robot Sign Language
Applies established seq2seq neural networks to convert text to Spanish sign language for humanoid robot TEO, proposing OpenPose for skeleton data collection to handle sequence length differences and non-manual markers.
Discussion (0). Continue with ORCID to comment.