REVIEW 2 cited by
Look Ma, no markers: holistic performance capture without the hassle
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We tackle the problem of highly-accurate, holistic performance capture for the face, body and hands simultaneously. Motion-capture technologies used in film and game production typically focus only on face, body or hand capture independently, involve complex and expensive hardware and a high degree of manual intervention from skilled operators. While machine-learning-based approaches exist to overcome these problems, they usually only support a single camera, often operate on a single part of the body, do not produce precise world-space results, and rarely generalize outside specific contexts. In this work, we introduce the first technique for marker-free, high-quality reconstruction of the complete human body, including eyes and tongue, without requiring any calibration, manual intervention or custom hardware. Our approach produces stable world-space results from arbitrary camera rigs as well as supporting varied capture environments and clothing. We achieve this through a hybrid approach that leverages machine learning models trained exclusively on synthetic data and powerful parametric models of human shape and motion. We evaluate our method on a number of body, face and hand reconstruction benchmarks and demonstrate state-of-the-art results that generalize on diverse datasets.
Forward citations
Cited by 2 Pith papers
-
Generative Zoo
A pipeline uses a diffusion image generator to create realistic animal photos with exact 3D pose and shape labels, achieving state-of-the-art 3D animal pose and shape estimation when trained only on synthetic data.
-
SEREP: Semantic Facial Expression Representation for Robust In-the-Wild Capture and Retargeting
SEREP learns a semantic facial expression code from unpaired 3D scans and a semi-supervised image encoder, reporting lower 3D expression error than DECA, EMICA, and SMIRK on the new MultiREX benchmark.
Discussion (0). Continue with ORCID to comment.