Pith. sign in

Fast Registration of Photorealistic Avatars for VR Facial Animation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Virtual Reality (VR) bares promise of social interactions that can feel more immersive than other media. Key to this is the ability to accurately animate a personalized photorealistic avatar, and hence the acquisition of the labels for headset-mounted camera (HMC) images need to be efficient and accurate, while wearing a VR headset. This is challenging due to oblique camera views and differences in image modality. In this work, we first show that the domain gap between the avatar and HMC images is one of the primary sources of difficulty, where a transformer-based architecture achieves high accuracy on domain-consistent data, but degrades when the domain-gap is re-introduced. Building on this finding, we propose a system split into two parts: an iterative refinement module that takes in-domain inputs, and a generic avatar-guided image-to-image domain transfer module conditioned on current estimates. These two modules reinforce each other: domain transfer becomes easier when close-to-groundtruth examples are shown, and better domain-gap removal in turn improves the registration. Our system obviates the need for costly offline optimization, and produces online registration of higher quality than direct regression method. We validate the accuracy and efficiency of our approach through extensive experiments on a commodity headset, demonstrating significant improvements over these baselines. To stimulate further research in this direction, we make our large-scale dataset and code publicly available.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2024 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

clear filters

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper after filters.

  • CFSynthesis: Controllable and Free-view 3D Human Video Synthesis cs.CV · 2024-12-15 · conditional · none · ref 44 · internal anchor

    CFSynthesis generates free-view human videos from one reference image by conditioning a diffusion model on a textured SMPL body model and separately encoded foreground and background.