Pith. sign in

REVIEW 1 cited by

Speech inpainting: Context-based speech synthesis guided by video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00489 v1 pith:2QX5QOOO submitted 2023-06-01 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords speechaudiovisualaudio-visualcontentcorruptedinformationinpainting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal or even restore missing audio information. Specifically, this paper focuses on the problem of audio-visual speech inpainting, which is the task of synthesizing the speech in a corrupted audio segment in a way that it is consistent with the corresponding visual content and the uncorrupted audio context. We present an audio-visual transformer-based deep learning model that leverages visual cues that provide information about the content of the corrupted audio. It outperforms the previous state-of-the-art audio-visual model and audio-only baselines. We also show how visual features extracted with AV-HuBERT, a large audio-visual transformer for speech recognition, are suitable for synthesizing speech.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A multi-codebook discrete diffusion model with coarse-to-fine RVQ generation and span-localized guidance improves speech inpainting and editing on RealEdit.

Pith tools