Pith. sign in

REVIEW 1 cited by

Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.15585 v1 pith:BAZ25TCR submitted 2024-08-28 cs.SD eess.AS

classification cs.SDeess.AS
keywords whisperspeakermodelaggregationapproachbaselineblocksencoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, Whisper, a large-scale pre-trained model for automatic speech recognition, is proposed to apply to speaker verification. A partial multi-scale feature aggregation (PMFA) approach is proposed based on a subset of Whisper encoder blocks to derive highly discriminative speaker embeddings.Experimental results demonstrate that using the middle to later blocks of the Whisper encoder keeps more speaker information. On the VoxCeleb1 and CN-Celeb1 datasets, our system achieves 1.42% and 8.23% equal error rates (EERs) respectively, receiving 0.58% and 1.81% absolute EER reductions over the ECAPA-TDNN baseline, and 0.46% and 0.97% over the ResNet34 baseline. Furthermore, our results indicate that using Whisper models trained on multilingual data can effectively enhance the model's robustness across languages. Finally, the low-rank adaptation approach is evaluated, which reduces the trainable model parameters by approximately 45 times while only slightly increasing EER by 0.2%.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Missing-Token Prompted Reliability-Aware Fusion for Robust Polyglot Speaker Identification

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    MRAF framework uses missing-token prompting and reliability-aware cross-attention fusion to achieve 100% accuracy on some POLY-SIM 2026 tasks and competitive results on missing-face cases.

Pith tools