REVIEW 19 cited by
Listen, Attend and Spell
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly. Our system has two components: a listener and a speller. The listener is a pyramidal recurrent network encoder that accepts filter bank spectra as inputs. The speller is an attention-based recurrent network decoder that emits characters as outputs. The network produces character sequences without making any independence assumptions between the characters. This is the key improvement of LAS over previous end-to-end CTC models. On a subset of the Google voice search task, LAS achieves a word error rate (WER) of 14.1% without a dictionary or a language model, and 10.3% with language model rescoring over the top 32 beams. By comparison, the state-of-the-art CLDNN-HMM model achieves a WER of 8.0%.
Forward citations
Cited by 19 Pith papers
-
Structure Before Collapse: Transient semantic geometry in next-token prediction
Semantic geometry emerges transiently early in next-token prediction training before collapsing to Neural Collapse symmetry in synthetic settings with latent semantic factors.
-
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization
PairAlign learns compact variable-length token sequences for audio via self-alignment on paired content-preserving views, achieving 55% fewer archive tokens than VQ while preserving edit-distance retrieval at 12.71 tokens/s.
-
Learning to See Inside Opaque Liquid Containers using Speckle Vibrometry
A speckle-vibrometry rig with a transformer can remotely classify the liquid level inside opaque containers from surface vibrations excited by sound.
-
Generative Testing of Automated Speech Recognition Systems
Phoneme-level latent interpolation in a TTS model yields ~98% black-box ASR failures with higher naturalness than waveform attacks and quality competitive with white-box PGD.
-
Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context
An LLM-based revision method with phonetic-semantic context reduces named entity word error rate by up to 30% relative on a new 45-hour MIT classroom speech dataset.
-
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
A text-like 'unit language' mined from discrete speech units via n-gram modeling, plus task-prompt multi-task training, improves textless speech-to-speech translation to near text-supervised performance.
-
Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers
A transformer speech encoder can learn to internally rearrange audio information into text order, enabling a lightweight decoder trained with simple cross-entropy to nearly match RNN-Transducer accuracy with faster inference.
-
Cross-Attention End-to-End ASR for Two-Party Conversations
End-to-end ASR model with speaker-specific cross-attention for two-party conversations outperforms standard models on the Switchboard corpus.
-
NIESR: Nuisance Invariant End-to-end Speech Recognition
NIESR applies unsupervised adversarial invariance induction to end-to-end ASR, reporting 5.48-14.44% relative error reductions on WSJ0, CHiME3, and TIMIT without nuisance factor labels.
-
Self Multi-Head Attention for Speaker Recognition
Self multi-head attention applied after CNN encoding of spectrograms outperforms temporal and statistical pooling for speaker verification on VoxCeleb1 with 18% relative EER reduction.
-
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition
A 4B-parameter LLM ASR system with five multi-token-prediction branches reports 2.97% CER Chinese, 3.68% WER English, 3.70% long-form WER, and a 0.0053 real-time factor, but the acceptance-rate calculation and ablatio...
-
StepAudio 2.5 Technical Report
StepAudio 2.5 is a unified audio-language foundation model that reaches state-of-the-art results on ASR, TTS, and realtime interaction by using task-tailored RLHF on a shared backbone.
-
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
A gradient-sensitive gating network plus multi-stage dropout fuses FBanks and HuBERT features, matching BLEU while cutting MuST-C training epochs by roughly 1.24x.
-
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
Across controlled ASR and speech translation experiments, dense feature prepending does not outperform cross-attention in quality and is slightly slower and more memory hungry.
-
Two-Pass End-to-End Speech Recognition
A two-pass speech recognizer, with a streaming RNN-T first pass and a LAS attention-based rescoring second pass sharing an encoder, reduces word error rate by 17 to 22 percent relative to RNN-T alone at under 200 ms a...
-
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
MedASR is an open-source 105M-parameter ASR model achieving 58% relative WER reduction versus Whisper Large-v3 on medical dictation.
-
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
A Temporal Alignment Buffer with minimum-KL delay selection lets Delayed-KD reach 5.42% CER on AISHELL-1 at 40 ms latency, matching U2++ at 320 ms.
-
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Cosine similarity between t-vector speaker embeddings in a streaming transducer speech translation model detects speaker changes (F1 up to 0.68) and classifies gender (0.989 accuracy).
-
Hierarchical Sequence to Sequence Voice Conversion with Limited Data
Hierarchical seq2seq model for parallel voice conversion pretrained as autoencoder on single-speaker data then adapted to limited multispeaker data, using mel spectrograms converted via wavenet vocoder.
Discussion (0). Continue with ORCID to comment.