Pith. sign in

REVIEW 2 cited by

DiffusionTalker: Personalization and Acceleration for Speech-Driven 3D Face Diffuser

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.16565 v2 pith:2IBB3ENR submitted 2023-11-28 cs.CV cs.SDeess.AS

classification cs.CVcs.SDeess.AS
keywords animationfaciallearningmethodsmodelspeech-drivenaccelerationaudio
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech-driven 3D facial animation has been an attractive task in both academia and industry. Traditional methods mostly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the non-deterministic fact of speech-driven 3D face animation and employ the diffusion model for the task. However, personalizing facial animation and accelerating animation generation are still two major limitations of existing diffusion-based methods. To address the above limitations, we propose DiffusionTalker, a diffusion-based method that utilizes contrastive learning to personalize 3D facial animation and knowledge distillation to accelerate 3D animation generation. Specifically, to enable personalization, we introduce a learnable talking identity to aggregate knowledge in audio sequences. The proposed identity embeddings extract customized facial cues across different people in a contrastive learning manner. During inference, users can obtain personalized facial animation based on input audio, reflecting a specific talking style. With a trained diffusion model with hundreds of steps, we distill it into a lightweight model with 8 steps for acceleration. Extensive experiments are conducted to demonstrate that our method outperforms state-of-the-art methods. The code will be released.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A 3D facial animation framework that disentangles content and emotion and predicts frame-wise emotion intensity from audio plus text for dynamic expressions.

  2. A Comprehensive Review of Human Error in Risk-Informed Decision Making: Integrating Human Reliability Assessment, Artificial Intelligence, and Human Performance Models

    cs.HC 2025-06 unverdicted novelty 2.0 of 10

    A review of human error research concluding that integrating AI and cognitive models into human reliability assessment can markedly improve predictive fidelity, but data scarcity and opacity remain barriers.

Pith tools