Pith. sign in

REVIEW 2 cited by

The IDLAB VoxCeleb Speaker Recognition Challenge 2020 System Description

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.12468 v1 pith:S62BZ3HK submitted 2020-10-23 eess.AS cs.SD

classification eess.AScs.SD
keywords supervisedtrainingspeakersystemsystemstracksunsupervisedverification
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this technical report we describe the IDLAB top-scoring submissions for the VoxCeleb Speaker Recognition Challenge 2020 (VoxSRC-20) in the supervised and unsupervised speaker verification tracks. For the supervised verification tracks we trained 6 state-of-the-art ECAPA-TDNN systems and 4 Resnet34 based systems with architectural variations. On all models we apply a large margin fine-tuning strategy, which enables the training procedure to use higher margin penalties by using longer training utterances. In addition, we use quality-aware score calibration which introduces quality metrics in the calibration system to generate more consistent scores across varying levels of utterance conditions. A fusion of all systems with both enhancements applied led to the first place on the open and closed supervised verification tracks. The unsupervised system is trained through contrastive learning. Subsequent pseudo-label generation by iterative clustering of the training embeddings allows the use of supervised techniques. This procedure led to the winning submission on the unsupervised track, and its performance is closing in on supervised training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Clustering-based hard negative sampling for supervised contrastive speaker verification

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Clustering speaker voiceprints and packing training batches with within-cluster negative pairs improves supervised contrastive speaker verification by up to 18% relative EER on VoxCeleb.

  2. Enhancing Self-Supervised Speaker Verification Using Similarity-Connected Graphs and GCN

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A GCN-based similarity graph refinement step improves DINO pseudo-label clustering for self-supervised speaker verification, reporting 1.57% EER on VoxCeleb1-O.

Pith tools