REVIEW 2 cited by
The JHU submission to VoxSRC-21: Track 3
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This technical report describes Johns Hopkins University speaker recognition system submitted to Voxceleb Speaker Recognition Challenge 2021 Track 3: Self-supervised speaker verification (closed). Our overall training process is similar to the proposed one from the first place team in the last year's VoxSRC2020 challenge. The main difference is a recently proposed non-contrastive self-supervised method in computer vision (CV), distillation with no labels (DINO), is used to train our initial model, which outperformed the last year's contrastive learning based on momentum contrast (MoCo). Also, this requires only a few iterations in the iterative clustering stage, where pseudo labels for supervised embedding learning are updated based on the clusters of the embeddings generated from a model that is continually fine-tuned over iterations. In the final stage, Res2Net50 is trained on the final pseudo labels from the iterative clustering stage. This is our best submitted model to the challenge, showing 1.89, 6.50, and 6.89 in EER(%) in voxceleb1 test o, VoxSRC-21 validation, and test trials, respectively.
Forward citations
Cited by 2 Pith papers
-
Clustering-based hard negative sampling for supervised contrastive speaker verification
Clustering speaker voiceprints and packing training batches with within-cluster negative pairs improves supervised contrastive speaker verification by up to 18% relative EER on VoxCeleb.
-
Enhancing Self-Supervised Speaker Verification Using Similarity-Connected Graphs and GCN
A GCN-based similarity graph refinement step improves DINO pseudo-label clustering for self-supervised speaker verification, reporting 1.57% EER on VoxCeleb1-O.
Discussion (0). Continue with ORCID to comment.