Pith. sign in

REVIEW 5 cited by

Decoder-only Architecture for Streaming End-to-end Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16107 v2 pith:A73MXEHC submitted 2024-06-23 eess.AS cs.CL

classification eess.AScs.CL
keywords decoder-onlyspeechstreamingblockwisepromptsarchitecturedecodermodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Decoder-only language models (LMs) have been successfully adopted for speech-processing tasks including automatic speech recognition (ASR). The LMs have ample expressiveness and perform efficiently. This efficiency is a suitable characteristic for streaming applications of ASR. In this work, we propose to use a decoder-only architecture for blockwise streaming ASR. In our approach, speech features are compressed using CTC output and context embedding using blockwise speech subnetwork, and are sequentially provided as prompts to the decoder. The decoder estimates the output tokens promptly at each block. To this end, we also propose a novel training scheme using random-length prefix prompts to make the model robust to the truncated prompts caused by blockwise processing. An experimental comparison shows that our proposed decoder-only streaming ASR achieves 8% relative word error rate reduction in the LibriSpeech test-other set while being twice as fast as the baseline model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.

  2. MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MFLA adds finite look-ahead attention plus a CIF-based token counter to Whisper, enabling streaming recognition with a wait-k latency-quality trade-off.

  3. SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation

    cs.CL 2025-04 conditional novelty 6.0 of 10

    An offline-trained speech LLM with boundary-aware CIF speech prompts and test-time wait-k decoding achieves better quality-latency trade-offs in simultaneous speech-to-speech translation than StreamSpeech on CVSS-C.

  4. Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A streaming ASR architecture that adds a frozen LLM as the internal language model of a factorized transducer, with vocabulary adaptation and weak-to-strong LM swapping, reports large WER reductions.

  5. Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Across controlled ASR and speech translation experiments, dense feature prepending does not outperform cross-attention in quality and is slightly slower and more memory hungry.

Pith tools