Pith. sign in

Multi-person 3D Pose Estimation in Crowded Scenes Based on Multi-View Geometry

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Epipolar constraints are at the core of feature matching and depth estimation in current multi-person multi-camera 3D human pose estimation methods. Despite the satisfactory performance of this formulation in sparser crowd scenes, its effectiveness is frequently challenged under denser crowd circumstances mainly due to two sources of ambiguity. The first is the mismatch of human joints resulting from the simple cues provided by the Euclidean distances between joints and epipolar lines. The second is the lack of robustness from the naive formulation of the problem as a least squares minimization. In this paper, we depart from the multi-person 3D pose estimation formulation, and instead reformulate it as crowd pose estimation. Our method consists of two key components: a graph model for fast cross-view matching, and a maximum a posteriori (MAP) estimator for the reconstruction of the 3D human poses. We demonstrate the effectiveness and superiority of our proposed method on four benchmark datasets.

citation-role summary

baseline 1

citation-polarity summary

fields

cs.CV 1

years

2024 1

verdicts

CONDITIONAL 1

roles

baseline 1

polarities

baseline 1

representative citing papers

Reconstructing People, Places, and Cameras

cs.CV · 2024-12-23 · conditional · novelty 6.0

HSfM jointly optimizes human meshes, dense scene pointmaps, and camera poses in a metric world frame, reducing world-frame human joint error from 3.5m to 1.0m on EgoHumans.

citing papers explorer

Showing 1 of 1 citing paper.

  • Reconstructing People, Places, and Cameras cs.CV · 2024-12-23 · conditional · none · ref 6 · internal anchor

    HSfM jointly optimizes human meshes, dense scene pointmaps, and camera poses in a metric world frame, reducing world-frame human joint error from 3.5m to 1.0m on EgoHumans.