REVIEW 3 cited by
The Conversation: Deep Audio-Visual Speech Enhancement
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose a deep audio-visual speech enhancement network that is able to separate a speaker's voice given lip regions in the corresponding video, by predicting both the magnitude and the phase of the target signal. The method is applicable to speakers unheard and unseen during training, and for unconstrained environments. We demonstrate strong quantitative and qualitative results, isolating extremely challenging real-world examples.
Forward citations
Cited by 3 Pith papers
-
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
MUTUD trains audiovisual speech models with both modalities but lets them run with audio only, recovering much of the multimodal benefit at a fraction of the compute.
-
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
Adding depth maps from an RGB-D camera to multiview microphone-array signals markedly improves 3D localization and classification of visually invisible sound sources in simulated indoor scenes.
-
ASAudio: A Survey of Advanced Spatial Audio Research
A comprehensive survey that systematically categorizes spatial audio research by representation, task, dataset, and evaluation.
Discussion (0). Continue with ORCID to comment.