Pith. sign in

REVIEW

Speech Driven Video Editing via an Audio-Conditioned Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.04474 v3 pith:CXJFPNCF submitted 2023-01-10 cs.CV cs.LGcs.SDeess.AS

classification cs.CVcs.LGcs.SDeess.AS
keywords diffusionmodelvideoeditingdenoisingend-to-endfacialmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Taking inspiration from recent developments in visual generative tasks using diffusion models, we propose a method for end-to-end speech-driven video editing using a denoising diffusion model. Given a video of a talking person, and a separate auditory speech recording, the lip and jaw motions are re-synchronized without relying on intermediate structural representations such as facial landmarks or a 3D face model. We show this is possible by conditioning a denoising diffusion model on audio mel spectral features to generate synchronised facial motion. Proof of concept results are demonstrated on both single-speaker and multi-speaker video editing, providing a baseline model on the CREMA-D audiovisual data set. To the best of our knowledge, this is the first work to demonstrate and validate the feasibility of applying end-to-end denoising diffusion models to the task of audio-driven video editing.

Discussion (0). Continue with ORCID to comment.

Pith tools