Pith. sign in

REVIEW 2 cited by

ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02167 v1 pith:UME5KXGE submitted 2024-06-04 eess.AS eess.SP

classification eess.ASeess.SP
keywords short-durationspeakertrialcomputationaleres2netv2featuremodelperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker verification systems experience significant performance degradation when tasked with short-duration trial recordings. To address this challenge, a multi-scale feature fusion approach has been proposed to effectively capture speaker characteristics from short utterances. Constrained by the model's size, a robust backbone Enhanced Res2Net (ERes2Net) combining global and local feature fusion demonstrates sub-optimal performance in short-duration speaker verification. To further improve the short-duration feature extraction capability of ERes2Net, we expand the channel dimension within each stage. However, this modification also increases the number of model parameters and computational complexity. To alleviate this problem, we propose an improved ERes2NetV2 by pruning redundant structures, ultimately reducing both the model parameters and its computational cost. A range of experiments conducted on the VoxCeleb datasets exhibits the superiority of ERes2NetV2, which achieves EER of 0.61% for the full-duration trial, 0.98% for the 3s-duration trial, and 1.48% for the 2s-duration trial on VoxCeleb1-O, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Fusing zero-shot TTS-generated speech embeddings with original short-utterance embeddings reduces speaker-verification EER by 10-16% on VoxCeleb1 without retraining.

  2. Marco-Voice Technical Report

    cs.CL 2025-08 reject novelty 4.0 of 10

    Marco-Voice is a TTS system combining voice cloning and emotional speech generation via speaker-emotion disentanglement, contrastive learning, and a new Mandarin emotional dataset, with claimed quality gains over Cosy...

Pith tools