REVIEW 5 cited by
In defence of metric learning for speaker recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The objective of this paper is 'open-set' speaker recognition of unseen speakers, where ideal embeddings should be able to condense information into a compact utterance-level representation that has small intra-speaker and large inter-speaker distance. A popular belief in speaker recognition is that networks trained with classification objectives outperform metric learning methods. In this paper, we present an extensive evaluation of most popular loss functions for speaker recognition on the VoxCeleb dataset. We demonstrate that the vanilla triplet loss shows competitive performance compared to classification-based losses, and those trained with our proposed metric learning objective outperform state-of-the-art methods.
Forward citations
Cited by 5 Pith papers
-
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
Speaker recognition networks form hierarchical clusters in latent space that can be matched to semantic classes using new HCCM algorithm and quantified by Liebig's score.
-
Non-Adaptive Adversarial Face Generation
A non-adaptive batch of 100 face queries against a face recognition API can produce synthetic faces that impersonate a target identity with a chosen attribute, exceeding 93% success against AWS CompareFaces at its def...
-
Explainable AI in Speaker Recognition -- Attention Map Visualisation and Evaluation
The paper introduces Modified RISE-eval to evaluate GradCAM and LayerCAM attention maps on speaker recognition networks and reports distinct advantages for each method under different conditions.
-
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
Hierarchical clustering of speaker recognition embeddings reveals that the network organizes representations by gender at upper levels and by gender-nationality conjunctions at lower levels, interpreted via a proposed...
-
Addressing malware family concept drift with triplet autoencoder
Combining a triplet autoencoder with DBSCAN centroid thresholds improves detection of unseen malware families in two datasets, but the temporal evaluation protocol and unreported hyperparameters undermine the results.
Discussion (0). Sign in to comment.