Pith. sign in

REVIEW 1 cited by

Improving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.09669 v2 pith:NZH3IASX submitted 2021-06-17 cs.SD cs.CVcs.LG

classification cs.SDcs.CVcs.LG
keywords on-screenseparationaudio-visualvideosmodelaudiosoundimprovements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a state-of-the-art audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify limitations of previous work on audio-visual on-screen sound separation, including the simplicity and coarse resolution of spatio-temporal attention, and poor convergence of the audio separation model. Our proposed model addresses these issues using cross-modal and self-attention modules that capture audio-visual dependencies at a finer resolution over time, and by unsupervised pre-training of audio separation model. These improvements allow the model to generalize to a much wider set of unseen videos. We also show a robust way to further improve the generalization capability of our models by calibrating the probabilities of our audio-visual on-screen classifier, using only a small amount of in-domain videos labeled for their on-screen presence. For evaluation and semi-supervised training, we collected human annotations of on-screen audio from a large database of in-the-wild videos (YFCC100m). Our results show marked improvements in on-screen separation performance, in more general conditions than previous methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation

    eess.AS 2025-01 conditional novelty 5.0 of 10

    VoiceFormer fuses text, video, and audio in a transformer to separate a target speaker, and stays robust when audio and video are misaligned by up to 200 ms.

Pith tools