Pith. sign in

REVIEW 1 cited by

Augmenting conformers with structured state-space sequence models for online speech recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08551 v2 pith:PHFEKLX2 submitted 2023-09-15 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords modelsonlineaugmentingconformerscontextconvolutionleftmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Online speech recognition, where the model only accesses context to the left, is an important and challenging use case for ASR systems. In this work, we investigate augmenting neural encoders for online ASR by incorporating structured state-space sequence models (S4), a family of models that provide a parameter-efficient way of accessing arbitrarily long left context. We performed systematic ablation studies to compare variants of S4 models and propose two novel approaches that combine them with convolutions. We found that the most effective design is to stack a small S4 using real-valued recurrent weights with a local convolution, allowing them to work complementarily. Our best model achieves WERs of 4.01%/8.53% on test sets from Librispeech, outperforming Conformers with extensively tuned convolution.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions

    cs.LG 2025-01 conditional novelty 7.0 of 10

    Treating state-space layers as tensor networks with CNN-style connectivity and optimized contraction orders yields hybrid SSM networks that outperform homogeneous SSMs on raw audio tasks and enable competitive streami...

Pith tools