Pith. sign in

REVIEW 1 cited by

VoxelKeypointFusion: Generalizable Multi-View Multi-Person Pose Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.18723 v3 pith:M6JBLMGL submitted 2024-10-24 cs.CV cs.HC

VoxelKeypointFusion: Generalizable Multi-View Multi-Person Pose Estimation

classification cs.CV cs.HC
keywords multi-personmulti-viewpresentsdatasetsposetaskunseenwell
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In the rapidly evolving field of computer vision, the task of accurately estimating the poses of multiple individuals from various viewpoints presents a formidable challenge, especially if the estimations should be reliable as well. This work presents an extensive evaluation of the generalization capabilities of multi-view multi-person pose estimators to unseen datasets and presents a new algorithm with strong performance in this task. It also studies the improvements by additionally using depth information. Since the new approach can not only generalize well to unseen datasets, but also to different keypoints, the first multi-view multi-person whole-body estimator is presented. To support further research on those topics, all of the work is publicly accessible.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Skarimva: Skeleton-based Action Recognition is a Multi-view Application

    cs.CV 2026-02 unverdicted novelty 4.0

    Multi-view camera setups that triangulate higher-quality 3D skeletons measurably improve state-of-the-art skeleton-based action recognition models, implying input data quality is the current bottleneck.