Pith. sign in

REVIEW 2 cited by

Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.06030 v2 pith:OKE4F2UE submitted 2020-10-12 cs.CL cs.AIcs.LGcs.SDeess.AS

classification cs.CLcs.AIcs.LGcs.SDeess.AS
keywords streamingdual-modefull-contextaccuracylatencyrecognitionspeechstate-of-the-art
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible, while full-context ASR waits for the completion of a full speech utterance before emitting completed hypotheses. In this work, we propose a unified framework, Dual-mode ASR, to train a single end-to-end ASR model with shared weights for both streaming and full-context speech recognition. We show that the latency and accuracy of streaming ASR significantly benefit from weight sharing and joint training of full-context ASR, especially with inplace knowledge distillation during the training. The Dual-mode ASR framework can be applied to recent state-of-the-art convolution-based and transformer-based ASR networks. We present extensive experiments with two state-of-the-art ASR networks, ContextNet and Conformer, on two datasets, a widely used public dataset LibriSpeech and a large-scale dataset MultiDomain. Experiments and ablation studies demonstrate that Dual-mode ASR not only simplifies the workflow of training and deploying streaming and full-context ASR models, but also significantly improves both emission latency and recognition accuracy of streaming ASR. With Dual-mode ASR, we achieve new state-of-the-art streaming ASR results on both LibriSpeech and MultiDomain in terms of accuracy and latency.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    A single speech encoder trained via ASR-aware distillation with variable attention masking performs competitively in both streaming and full-context modes at 200M and 2B scale.

  2. Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A Temporal Alignment Buffer with minimum-KL delay selection lets Delayed-KD reach 5.42% CER on AISHELL-1 at 40 ms latency, matching U2++ at 320 ms.

Pith tools