Pith. sign in

REVIEW

TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.12306 v1 pith:466OMRFF submitted 2023-09-21 cs.CV cs.SDeess.AS

classification cs.CVcs.SDeess.AS
keywords lossactivecontrastivedetectioneffectiveexistingimprovinglearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures while learning effective representations has been less explored. In this work, we propose TalkNCE, a novel talk-aware contrastive loss. The loss is only applied to part of the full segments where a person on the screen is actually speaking. This encourages the model to learn effective representations through the natural correspondence of speech and facial movements. Our loss can be jointly optimized with the existing objectives for training ASD models without the need for additional supervision or training data. The experiments demonstrate that our loss can be easily integrated into the existing ASD frameworks, improving their performance. Our method achieves state-of-the-art performances on AVA-ActiveSpeaker and ASW datasets.

Discussion (0). Sign in to comment.

Pith tools