Pith. sign in

REVIEW 3 cited by

First-Pass Large Vocabulary Continuous Speech Recognition using Bi-Directional Recurrent DNNs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1408.2873 v2 pith:PGNWK44Z submitted 2014-08-12 cs.CL cs.LGcs.NE

classification cs.CLcs.LGcs.NE
keywords networkrecognitionspeechfirst-passneuralsystemsapproachbi-directional
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present a method to perform first-pass large vocabulary continuous speech recognition using only a neural network and language model. Deep neural network acoustic models are now commonplace in HMM-based speech recognition systems, but building such systems is a complex, domain-specific task. Recent work demonstrated the feasibility of discarding the HMM sequence modeling framework by directly predicting transcript text from audio. This paper extends this approach in two ways. First, we demonstrate that a straightforward recurrent neural network architecture can achieve a high level of accuracy. Second, we propose and evaluate a modified prefix-search decoding algorithm. This approach to decoding enables first-pass speech recognition with a language model, completely unaided by the cumbersome infrastructure of HMM-based systems. Experiments on the Wall Street Journal corpus demonstrate fairly competitive word error rates, and the importance of bi-directional network recurrence.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CANDLE: CTC-based Arabic Noisy-character Deduplication using a Lightweight Encoder

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    CANDLE uses CTC on lightweight character encoders for Arabic noise deduplication, reporting 5.37% SER on benchmarks and up to 12.8% tokenizer fertility reduction.

  2. Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency

    eess.AS 2024-12 conditional novelty 6.0 of 10

    A streaming CTC keyword-spotting decoder with cross-layer cosine-similarity refinement reports a 6.8% absolute recall gain over graph-based decoding on the Hey Snips dataset.

  3. Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Delayed fusion scores partial ASR hypotheses with a pre-trained LLM only after pruning and at word boundaries, giving lower word error rates than N-best rescoring without retraining the ASR model.

Pith tools