Pith. sign in

REVIEW 1 cited by

Introducing ECAPA-TDNN and Wav2Vec2.0 Embeddings to Stuttering Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.01564 v1 pith:OCN2W2CQ submitted 2022-04-04 cs.SD cs.LGeess.AS

Introducing ECAPA-TDNN and Wav2Vec2.0 Embeddings to Stuttering Detection

classification cs.SD cs.LGeess.AS
keywords embeddingsdatasetsdetectionstutteringtaskstrainedwav2vec2audio
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The adoption of advanced deep learning (DL) architecture in stuttering detection (SD) tasks is challenging due to the limited size of the available datasets. To this end, this work introduces the application of speech embeddings extracted with pre-trained deep models trained on massive audio datasets for different tasks. In particular, we explore audio representations obtained using emphasized channel attention, propagation, and aggregation-time-delay neural network (ECAPA-TDNN) and Wav2Vec2.0 model trained on VoxCeleb and LibriSpeech datasets respectively. After extracting the embeddings, we benchmark with several traditional classifiers, such as a k-nearest neighbor, Gaussian naive Bayes, and neural network, for the stuttering detection tasks. In comparison to the standard SD system trained only on the limited SEP-28k dataset, we obtain a relative improvement of 16.74% in terms of overall accuracy over baseline. Finally, we have shown that combining two embeddings and concatenating multiple layers of Wav2Vec2.0 can further improve SD performance up to 1% and 2.64% respectively.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Gender Gap Analysis in News and Talk Online Radio Broadcast

    cs.CY 2026-06 conditional novelty 6.0

    Male speakers account for 77% of speaking time on U.S. news and talk radio, with female shares never exceeding 35.8% in any topic and falling to 10.8% in talk-show segments.