REVIEW 9 cited by
CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Time delay neural network (TDNN) has been proven to be efficient for speaker verification. One of its successful variants, ECAPA-TDNN, achieved state-of-the-art performance at the cost of much higher computational complexity and slower inference speed. This makes it inadequate for scenarios with demanding inference rate and limited computational resources. We are thus interested in finding an architecture that can achieve the performance of ECAPA-TDNN and the efficiency of vanilla TDNN. In this paper, we propose an efficient network based on context-aware masking, namely CAM++, which uses densely connected time delay neural network (D-TDNN) as backbone and adopts a novel multi-granularity pooling to capture contextual information at different levels. Extensive experiments on two public benchmarks, VoxCeleb and CN-Celeb, demonstrate that the proposed architecture outperforms other mainstream speaker verification systems with lower computational cost and faster inference speed.
Forward citations
Cited by 9 Pith papers
-
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.
-
HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training
A two-stage audio-visual deepfake detector, HOLA, uses 1.81M pre-training samples and hierarchical cross-modal fusion modules to achieve first place and near-perfect AUC on AV-Deepfake1M++ video-level detection.
-
Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models
A multi-stage pipeline combining curriculum learning, chain-of-thought data, and RL with verifiable rewards achieves 11.57% WER on the MLC-SLM multilingual ASR test set, versus a 20.17% baseline.
-
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
A speaker-aware progressive OSD model using WavLM, Campplus, and VAD-gated masking reports 82.76% F1 on AMI, above the listed prior best of 79.21%.
-
VoxAging: Continuously Tracking Speaker Aging with a Large-Scale Longitudinal Dataset in English and Mandarin
A 293-speaker longitudinal dataset with weekly samples over up to 17 years is introduced and used to show that speaker verification error grows with age, especially for female and middle-aged speakers.
-
NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
A hierarchical NVAE plug-in generates diverse pseudo-speaker embeddings that raise ASV EER above 38% on FACodec/CosyVoice2 with a controllable privacy-utility trade-off.
-
X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
An open, modular cascaded system (streaming ASR + MT + prompt-conditioned TTS) preserves speaker identity in long-form multi-speaker translation, at higher latency and slightly lower translation quality than proprietary APIs.
-
Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
A Bi-LSTM variant of ECAPA-TDNN's Res2Block cuts speaker-verification EER by 23% on VoxCeleb1-O at nearly the same parameter count.
-
DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
A dual-branch pooling method, DRASP, combines global statistics with segment-level attention and improves MOS prediction correlation with human ratings.
Discussion (0). Continue with ORCID to comment.