Pith. sign in

REVIEW 2 cited by

Decomposed Temporal Dynamic CNN: Efficient Time-Adaptive Network for Text-Independent Speaker Verification Explained with Speaker Activation Map

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.15277 v2 pith:L6KURH27 submitted 2022-03-29 eess.AS cs.SD

classification eess.AScs.SD
keywords speakerdynamicinformationverificationactivationdtdy-cnnsdtdy-resnet-34kernel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To extract accurate speaker information for text-independent speaker verification, temporal dynamic CNNs (TDY-CNNs) adapting kernels to each time bin was proposed. However, model size of TDY-CNN is too large and the adaptive kernel's degree of freedom is limited. To address these limitations, we propose decomposed temporal dynamic CNNs (DTDY-CNNs) which forms time-adaptive kernel by combining static kernel with dynamic residual based on matrix decomposition. Proposed DTDY-ResNet-34(x0.50) using attentive statistical pooling without data augmentation shows EER of 0.96%, which is better than other state-of-the-art methods. DTDY-CNNs are successful upgrade of TDY-CNNs, reducing the model size by 64% and enhancing the performance. We showed that DTDY-CNNs extract more accurate frame-level speaker embeddings as well compared to TDY-CNNs. Detailed behaviors of DTDY-ResNet-34(x0.50) on extraction of speaker information were analyzed using speaker activation map (SAM) produced by modified gradient-weighted class activation mapping (Grad-CAM) for speaker verification. DTDY-ResNet-34(x0.50) effectively extracts speaker information from not only formant frequencies but also high frequency information of unvoiced phonemes, thus explaining its outstanding performance on text-independent speaker verification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Frequency Dynamic Convolutions for Sound Event Detection

    eess.AS 2025-06 conditional novelty 4.0 of 10

    A family of frequency-adaptive convolutions improves CRNN sound event detection on DESED by up to 10.98% in PSDS1, with a lighter TFD variant matching the best score.

  2. Towards Understanding of Frequency Dependence on Sound Event Detection

    eess.AS 2025-02 conditional novelty 4.0 of 10

    Frequency-dependent augmentation and convolution improve sound event detection through complementary mechanisms, as shown by class-wise, Grad-CAM, and PCA analyses.

Pith tools