Pith. sign in

REVIEW 1 cited by

Phoneme-aware and Channel-wise Attentive Learning for Text DependentSpeaker Verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.13514 v1 pith:UY3JYJCH submitted 2021-06-25 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords learningattentivechannel-wisephoneme-awarespeakerproposedframe-levelfurther
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes a multi-task learning network with phoneme-aware and channel-wise attentive learning strategies for text-dependent Speaker Verification (SV). In the proposed structure, the frame-level multi-task learning along with the segment-level adversarial learning is adopted for speaker embedding extraction. The phoneme-aware attentive pooling is exploited on frame-level features in the main network for speaker classifier, with the corresponding posterior probability for the phoneme distribution in the auxiliary subnet. Further, the introduction of Squeeze and Excitation (SE-block) performs dynamic channel-wise feature recalibration, which improves the representational ability. The proposed method exploits speaker idiosyncrasies associated with pass-phrases, and is further improved by the phoneme-aware attentive pooling and SE-block from temporal and channel-wise aspects, respectively. The experiments conducted on RSR2015 Part 1 database confirm that the proposed system achieves outstanding results for textdependent SV.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024

    cs.SD 2024-11 conditional novelty 3.0 of 10

    An ASR content gate plus concatenated wav2vec-BERT and ReDimNet speaker embeddings achieved normalized min-DCF 0.0452 and rank 2 on the TDSV 2024 text-dependent speaker verification challenge.

Pith tools