Pith. sign in

REVIEW 3 cited by

The Conversation: Deep Audio-Visual Speech Enhancement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.04121 v2 pith:5G4ORXGG submitted 2018-04-11 cs.CV cs.SD

classification cs.CVcs.SD
keywords speakersspeechaudio-visualdeepenhancementenvironmentsseparateable
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose a deep audio-visual speech enhancement network that is able to separate a speaker's voice given lip regions in the corresponding video, by predicting both the magnitude and the phase of the target signal. The method is applicable to speakers unheard and unseen during training, and for unconstrained environments. We demonstrate strong quantitative and qualitative results, isolating extremely challenging real-world examples.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment

    cs.SD 2025-01 conditional novelty 6.0 of 10

    MUTUD trains audiovisual speech models with both modalities but lets them run with audio only, recovering much of the multimodal benefit at a fraction of the compute.

  2. SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera

    cs.SD 2024-12 conditional novelty 5.0 of 10

    Adding depth maps from an RGB-D camera to multiview microphone-array signals markedly improves 3D localization and classification of visually invisible sound sources in simulated indoor scenes.

  3. ASAudio: A Survey of Advanced Spatial Audio Research

    eess.AS 2025-08 unverdicted novelty 3.0 of 10

    A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.

Pith tools