Pith. sign in

Exploring RWKV for Memory Efficient and Low Latency Streaming ASR

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Recently, self-attention-based transformers and conformers have been introduced as alternatives to RNNs for ASR acoustic modeling. Nevertheless, the full-sequence attention mechanism is non-streamable and computationally expensive, thus requiring modifications, such as chunking and caching, for efficient streaming ASR. In this paper, we propose to apply RWKV, a variant of linear attention transformer, to streaming ASR. RWKV combines the superior performance of transformers and the inference efficiency of RNNs, which is well-suited for streaming ASR scenarios where the budget for latency and memory is restricted. Experiments on varying scales (100h - 10000h) demonstrate that RWKV-Transducer and RWKV-Boundary-Aware-Transducer achieve comparable to or even better accuracy compared with chunk conformer transducer, with minimal latency and inference memory cost.

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2024 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

A Survey of RWKV

cs.CL · 2024-12-19 · conditional · novelty 3.0

A review of the RWKV architecture, its versions, applications, benchmarks, and open-source ecosystem; it presents no new experimental results.

citing papers explorer

Showing 1 of 1 citing paper.

  • A Survey of RWKV cs.CL · 2024-12-19 · conditional · none · ref 166 · internal anchor

    A review of the RWKV architecture, its versions, applications, benchmarks, and open-source ecosystem; it presents no new experimental results.