Pith. sign in

REVIEW 1 cited by

Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.03965 v2 pith:4VQGMQPF submitted 2022-04-08 eess.AS cs.SD

classification eess.AScs.SD
keywords large-marginpldaspeakerback-endscosinecross-entropylossesscoring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of large-margin softmax cross-entropy losses in training deep speaker embedding neural networks has triggered a gradual shift from parametric back-ends to a simpler cosine similarity measure for speaker verification. Popular parametric back-ends include the probabilistic linear discriminant analysis (PLDA) and its variants. This paper investigates the properties of margin-based cross-entropy losses leading to such a shift and aims to find scoring back-ends best suited for speaker verification. In addition, we revisit the pre-processing techniques which have been widely used in the past and assess their effectiveness on large-margin embeddings. Experiments on the state-of-the-art ECAPA-TDNN networks trained with various large-margin softmax cross-entropy losses show a substantial increment in intra-speaker compactness making the conventional PLDA superfluous. In this regard, we found that constraining the within-speaker covariance matrix could improve the performance of the PLDA. It is demonstrated through a series of experiments on the VoxCeleb-1 and SITW core-core test sets with 40.8% equal error rate (EER) reduction and 35.1% minimum detection cost (minDCF) reduction. It also outperforms cosine scoring consistently with reductions in EER and minDCF by 10.9% and 4.9%, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

    eess.AS 2026-08 conditional novelty 4.0 of 10

    Because voices vary with mood, health, age, speaking style, and recording conditions, the voiceprint metaphor is scientifically misleading and voice evidence should be expressed as calibrated degrees of support.

Pith tools