Pith. sign in

REVIEW 2 cited by

Two-Pass End-to-End Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.10992 v1 pith:FSZFHIM7 submitted 2019-08-29 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modelconventionallatencyqualityrecognitionrnn-tspeechstreaming
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The requirements for many applications of state-of-the-art speech recognition systems include not only low word error rate (WER) but also low latency. Specifically, for many use-cases, the system must be able to decode utterances in a streaming fashion and faster than real-time. Recently, a streaming recurrent neural network transducer (RNN-T) end-to-end (E2E) model has shown to be a good candidate for on-device speech recognition, with improved WER and latency metrics compared to conventional on-device models [1]. However, this model still lags behind a large state-of-the-art conventional model in quality [2]. On the other hand, a non-streaming E2E Listen, Attend and Spell (LAS) model has shown comparable quality to large conventional models [3]. This work aims to bring the quality of an E2E streaming model closer to that of a conventional system by incorporating a LAS network as a second-pass component, while still abiding by latency constraints. Our proposed two-pass model achieves a 17%-22% relative reduction in WER compared to RNN-T alone and increases latency by a small fraction over RNN-T.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models

    eess.AS 2025-08 conditional novelty 6.0 of 10

    A hybrid decoding scheme with a fast TDT draft decoder and selective transformer patches matches baseline word error rate while cutting decoder latency roughly threefold.

  2. Non-Intrusive Automatic Speech Recognition Refinement: A Survey

    eess.AS 2025-08 accept novelty 4.0 of 10

    A survey that classifies non-intrusive ASR refinement methods into five categories, reviews domain adaptation and evaluation datasets, proposes standardized metrics, and identifies future research directions.

Pith tools