Pith. sign in

REVIEW 3 cited by

Very Deep Self-Attention Networks for End-to-End Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.13377 v2 pith:INUIC377 submitted 2019-04-30 cs.CL cs.LGcs.SDeess.AS

classification cs.CLcs.LGcs.SDeess.AS
keywords end-to-endmodelsnetworkspreviousdeephybridsystemstransformer
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, end-to-end sequence-to-sequence models for speech recognition have gained significant interest in the research community. While previous architecture choices revolve around time-delay neural networks (TDNN) and long short-term memory (LSTM) recurrent neural networks, we propose to use self-attention via the Transformer architecture as an alternative. Our analysis shows that deep Transformer networks with high learning capacity are able to exceed performance from previous end-to-end approaches and even match the conventional hybrid systems. Moreover, we trained very deep models with up to 48 Transformer layers for both encoder and decoders combined with stochastic residual connections, which greatly improve generalizability and training efficiency. The resulting models outperform all previous end-to-end ASR approaches on the Switchboard benchmark. An ensemble of these models achieve 9.9% and 17.7% WER on Switchboard and CallHome test sets respectively. This finding brings our end-to-end models to competitive levels with previous hybrid systems. Further, with model ensembling the Transformers can outperform certain hybrid systems, which are more complicated in terms of both structure and training procedure.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Depth-scaled initialization and merged attention let deep Transformer machine translation models train successfully and beat 6-layer baselines by about 1 BLEU.

  2. Adapting Foundation ASR Models to Dysarthric Speech: A Case Study

    cs.CL 2026-06 unverdicted novelty 4.0 of 10

    Fine-tuning Whisper on over 100 hours of personalized dysarthric speech data reduces WER to 9.7% for a single speaker.

  3. From Silent Signals to Natural Language: A Dual-Stage Transformer-LLM Approach

    cs.CL 2025-09 conditional novelty 3.0 of 10

    A transformer ASR with GPT-2 post-correction reduces word error rate for EMG-based silent speech recognition from 36% to 30% on the Digital Voicing test set.

Pith tools