Pith. sign in

REVIEW 1 cited by

DS-TDNN: Dual-stream Time-delay Neural Network with Global-aware Filter for Speaker Verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11020 v3 pith:7J3YNF54 submitted 2023-03-20 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords ds-tdnnlayerspeakercontextglobalverificationcalledcomplexity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Conventional time-delay neural networks (TDNNs) struggle to handle long-range context, their ability to represent speaker information is therefore limited in long utterances. Existing solutions either depend on increasing model complexity or try to balance between local features and global context to address this issue. To effectively leverage the long-term dependencies of audio signals and constrain model complexity, we introduce a novel module called Global-aware Filter layer (GF layer) in this work, which employs a set of learnable transform-domain filters between a 1D discrete Fourier transform and its inverse transform to capture global context. Additionally, we develop a dynamic filtering strategy and a sparse regularization method to enhance the performance of the GF layer and prevent overfitting. Based on the GF layer, we present a dual-stream TDNN architecture called DS-TDNN for automatic speaker verification (ASV), which utilizes two unique branches to extract both local and global features in parallel and employs an efficient strategy to fuse different-scale information. Experiments on the Voxceleb and SITW databases demonstrate that the DS-TDNN achieves a relative improvement of 10\% together with a relative decline of 20\% in computational cost over the ECAPA-TDNN in speaker verification task. This improvement will become more evident as the utterance's duration grows. Furthermore, the DS-TDNN also beats popular deep residual models and attention-based systems on utterances of arbitrary length.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification

    eess.AS 2025-09 conditional novelty 5.0 of 10

    A Bi-LSTM variant of ECAPA-TDNN's Res2Block cuts speaker-verification EER by 23% on VoxCeleb1-O at nearly the same parameter count.

Pith tools