Pith. sign in

REVIEW 1 cited by

3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.10848 v2 pith:ISJZFJLR submitted 2024-09-17 cs.CV cs.AIcs.LGcs.MMcs.SDeess.AS

classification cs.CVcs.AIcs.LGcs.MMcs.SDeess.AS
keywords vertexactionfacialapproachcontrolanimationaudio-drivendfacepolicy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial movements due to their frame-by-frame vertex generation approach, we propose 3DFacePolicy, a pioneer work that introduces a novel definition of vertex trajectory changes across consecutive frames through the concept of "action". By predicting action sequences for each vertex that encode frame-to-frame movements, we reformulate vertex generation approach into an action-based control paradigm. Specifically, we leverage a robotic control mechanism, diffusion policy, to predict action sequences conditioned on both audio and vertex states. Extensive experiments on VOCASET and BIWI datasets demonstrate that our approach significantly outperforms state-of-the-art methods and is particularly expert in dynamic, expressive and naturally smooth facial animations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QuMAB: Query-based Multi-Annotator Behavior Modeling with Reliability under Sparse Labels

    cs.MM 2025-07 conditional novelty 6.0 of 10

    QuMAB models each annotator with a lightweight query in a cross-attention network, reconstructs missing labels, and reports accuracy gains over aggregation baselines on two new dense-label datasets.

Pith tools