Pith. sign in

REVIEW

Margin-Mixup: A Method for Robust Speaker Verification in Multi-Speaker Audio

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03515 v1 pith:YT3FF4WJ submitted 2023-04-07 eess.AS cs.SD

classification eess.AScs.SD
keywords speakerverificationaudiomargin-mixupmulti-speakerrobustsystemsassumption
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper is concerned with the task of speaker verification on audio with multiple overlapping speakers. Most speaker verification systems are designed with the assumption of a single speaker being present in a given audio segment. However, in a real-world setting this assumption does not always hold. In this paper, we demonstrate that current speaker verification systems are not robust against audio with noticeable speaker overlap. To alleviate this issue, we propose margin-mixup, a simple training strategy that can easily be adopted by existing speaker verification pipelines to make the resulting speaker embeddings robust against multi-speaker audio. In contrast to other methods, margin-mixup requires no alterations to regular speaker verification architectures, while attaining better results. On our multi-speaker test set based on VoxCeleb1, the proposed margin-mixup strategy improves the EER on average with 44.4% relative to our state-of-the-art speaker verification baseline systems.

Discussion (0). Sign in to comment.

Pith tools