Pith. sign in

REVIEW 8 cited by

Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.11094 v1 pith:YX6ZT35A submitted 2022-07-22 cs.CV

classification cs.CV
keywords reconstructionvideosimagemethodmouthaudiodatadatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent state of the art on monocular 3D face reconstruction from image data has made some impressive advancements, thanks to the advent of Deep Learning. However, it has mostly focused on input coming from a single RGB image, overlooking the following important factors: a) Nowadays, the vast majority of facial image data of interest do not originate from single images but rather from videos, which contain rich dynamic information. b) Furthermore, these videos typically capture individuals in some form of verbal communication (public talks, teleconferences, audiovisual human-computer interactions, interviews, monologues/dialogues in movies, etc). When existing 3D face reconstruction methods are applied in such videos, the artifacts in the reconstruction of the shape and motion of the mouth area are often severe, since they do not match well with the speech audio. To overcome the aforementioned limitations, we present the first method for visual speech-aware perceptual reconstruction of 3D mouth expressions. We do this by proposing a "lipread" loss, which guides the fitting process so that the elicited perception from the 3D reconstructed talking head resembles that of the original video footage. We demonstrate that, interestingly, the lipread loss is better suited for 3D reconstruction of mouth movements compared to traditional landmark losses, and even direct 3D supervision. Furthermore, the devised method does not rely on any text transcriptions or corresponding audio, rendering it ideal for training in unlabeled datasets. We verify the efficiency of our method through exhaustive objective evaluations on three large-scale datasets, as well as subjective evaluation with two web-based user studies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    MOCHI enables registration-free training of multi-view 3D face reconstruction by enforcing topological consistency via a pseudo-linear inverse kinematic solver, using synthetic-data-trained 2D landmarks for alignment,...

  2. FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images

    cs.CV 2026-06 conditional novelty 6.0 of 10

    A Transformer 3D-Gaussian model reconstructs incremental, animatable 4D head avatars from sparse portraits via alternating attention, sparse-to-dense UV densification, and residual motion refinement.

  3. FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    FFAvatar uses a Transformer-based 3D Gaussian model with alternating attention and sparse-to-dense learning to enable feed-forward, incremental reconstruction of animatable 4D head avatars from sparse portrait images.

  4. DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    DyaPlex introduces a dual-tower Transformer that adds a streaming motion pathway to a frozen full-duplex speech model using dyadic token interleaving and time-aligned RoPE for synchronized multimodal dyadic interaction.

  5. Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation

    cs.LG 2026-01 conditional novelty 6.0 of 10

    A causal diffusion-forcing model generates interactive head-avatar motion with 500ms motion-generation latency and learns expressive reactions via DPO with synthetic negative samples.

  6. STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

    cs.CV 2025-12 conditional novelty 6.0 of 10

    STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.

  7. Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

    cs.GR 2025-09 conditional novelty 6.0 of 10

    Think2Sing uses LLM-generated, time-aligned motion subtitles and a motion-intensity proxy to guide diffusion-based 3D head animation from singing audio and lyrics.

  8. Combining Facial Videos and Biosignals for Stress Estimation During Driving

    cs.CV 2026-01 unverdicted novelty 4.0 of 10

    Fusing 3D facial motion descriptors with physiological signals via cross-modal attention improves stress detection AUROC from 52.7% to 92.0% on driving data.

Pith tools