Pith. sign in

REVIEW 1 cited by

UNISOUND System for VoxCeleb Speaker Recognition Challenge 2023

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.12526 v1 pith:ZWZTB52H submitted 2023-08-24 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords challengetracksystemplacerecognitionscorespeakersubmission
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This report describes the UNISOUND submission for Track1 and Track2 of VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC 2023). We submit the same system on Track 1 and Track 2, which is trained with only VoxCeleb2-dev. Large-scale ResNet and RepVGG architectures are developed for the challenge. We propose a consistency-aware score calibration method, which leverages the stability of audio voiceprints in similarity score by a Consistency Measure Factor (CMF). CMF brings a huge performance boost in this challenge. Our final system is a fusion of six models and achieves the first place in Track 1 and second place in Track 2 of VoxSRC 2023. The minDCF of our submission is 0.0855 and the EER is 1.5880%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification

    eess.AS 2024-12 conditional novelty 4.0 of 10

    A combination of reversible residual blocks and 8-bit optimizer state quantization trains deep speaker embedding extractors with up to 16.2x less GPU memory and comparable accuracy.

Pith tools