Pith. sign in

REVIEW 5 cited by

In defence of metric learning for speaker recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.11982 v2 pith:F6QCWKIX submitted 2020-03-26 eess.AS cs.SD

classification eess.AScs.SD
keywords recognitionspeakerlearningmetriclossmethodsobjectiveoutperform
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level representation that has small intra-speaker and large inter-speaker distance. A popular belief in speaker recognition is that networks trained with classification objectives outperform metric learning methods. In this paper, we present an extensive evaluation of most popular loss functions for speaker recognition on the VoxCeleb dataset. We demonstrate that the vanilla triplet loss shows competitive performance compared to classification-based losses, and those trained with our proposed metric learning objective outperform state-of-the-art methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    Speaker recognition networks form hierarchical clusters in latent space that can be matched to semantic classes using new HCCM algorithm and quantified by Liebig's score.

  2. Non-Adaptive Adversarial Face Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A non-adaptive batch of 100 face queries against a face recognition API can produce synthetic faces that impersonate a target identity with a chosen attribute, exceeding 93% success against AWS CompareFaces at its def...

  3. Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation

    eess.AS 2026-06 unverdicted novelty 5.0 of 10

    The paper introduces Modified RISE-eval to evaluate GradCAM and LayerCAM attention maps on speaker recognition networks and reports distinct advantages for each method under different conditions.

  4. Explainable AI in Speaker Recognition -- Making Latent Representations Understandable

    eess.AS 2026-04 conditional novelty 4.0 of 10

    Hierarchical clustering of speaker recognition embeddings reveals that the network organizes representations by gender at upper levels and by gender-nationality conjunctions at lower levels, interpreted via a proposed...

  5. Addressing malware family concept drift with triplet autoencoder

    cs.CR 2025-07 reject novelty 4.0 of 10

    Combining a triplet autoencoder with DBSCAN centroid thresholds improves detection of unseen malware families in two datasets, but the temporal evaluation protocol and unreported hyperparameters undermine the results.

Pith tools