A synthetic data pipeline with TTS, voice conversion, and novel L1-to-L2 accent conversion improves Whisper ASR word error rates on the ATCO2 corpus when used alone or mixed with real data.
hub
Deep Speech: Scaling up end-to-end speech recognition
17 Pith papers cite this work. Polarity classification is still indexing.
abstract
We present a state-of-the-art speech recognition system developed using end-to-end deep learning. Our architecture is significantly simpler than traditional speech systems, which rely on laboriously engineered processing pipelines; these traditional systems also tend to perform poorly when used in noisy environments. In contrast, our system does not need hand-designed components to model background noise, reverberation, or speaker variation, but instead directly learns a function that is robust to such effects. We do not need a phoneme dictionary, nor even the concept of a "phoneme." Key to our approach is a well-optimized RNN training system that uses multiple GPUs, as well as a set of novel data synthesis techniques that allow us to efficiently obtain a large amount of varied data for training. Our system, called Deep Speech, outperforms previously published results on the widely studied Switchboard Hub5'00, achieving 16.0% error on the full test set. Deep Speech also handles challenging noisy environments better than widely used, state-of-the-art commercial speech systems.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
Introduces a feature-vocoder adversarial attack on ASR using SSL representations that reports +26.6 WER black-box transfer and +36.2 WER defense resistance over baselines.
A gauge-covariant stochastic neural field theory is introduced that derives the maximal Lyapunov exponent and amplification factor, showing finite-width effects as perturbative corrections to dressed kernels that leave the marginality condition unchanged for fixed kernel geometry.
SpeechParaling-Bench is a new evaluation framework for paralinguistic-aware speech generation that reveals major limitations in current large audio-language models.
Deep learning generalization error follows power-law scaling with training set size across multiple domains, with model size scaling sublinearly with data size.
Mixed precision training uses FP16 for most computations, FP32 master weights for accumulation, and loss scaling to enable accurate training of large DNNs with halved memory usage.
End-to-end Conformer decoder achieves 23.8% character error rate on intracortical recordings for speech without external language models.
SWIM scales Whisper ASR to 20 concurrent multilingual clients via buffer merging, achieving ~2.4s delay at 5 clients versus 3.4s for single-client baselines while preserving accuracy.
DPOP is a new loss function that prevents DPO from lowering preferred response likelihoods and outperforms standard DPO on diverse datasets, MT-Bench, and enables Smaug-72B to exceed 80% on the Open LLM Leaderboard.
Decouples prosody alignment via pre-computed phoneme timestamps and adds VAE to achieve robust fine-grained prosody transfer in single-speaker neural TTS from unseen speakers.
Gated fusion of fastText and BERT embeddings into an end-to-end ASR model captures multi-sentence conversational context and lowers word error rate on the Switchboard corpus.
A multimodal adversarial attack using stage-sampled image nullification and cross-attention flattening degrades lip-sync and facial dynamics in Hallo-based talking-head generation.
MLLM-enabled video translation is usefully framed as three roles—Semantic Reasoner, Expressive Performer, and Visual Synthesizer—rather than a cascade of ASR, MT, TTS, and lip-sync.
Direct prompting scales more consistently than CoT prompting for speech-to-text translation as the amount of S2TT data increases.
On-policy distillation from a Qwen-ASR teacher improves a 0.6B Ark-ASR model over supervised fine-tuning and a same-scale baseline on four of five ASR benchmarks using 100k hours of speech.
The authors propose and test a data augmentation framework based on deepfake audio to improve training of speech-to-text transcription models.
A survey of spatial speech perception systems covering sound source localization, directional enhancement, and automatic speech recognition methods and their integration.
citing papers explorer
-
Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition
A synthetic data pipeline with TTS, voice conversion, and novel L1-to-L2 accent conversion improves Whisper ASR word error rates on the ATCO2 corpus when used alone or mixed with real data.
-
Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition
Introduces a feature-vocoder adversarial attack on ASR using SSL representations that reports +26.6 WER black-box transfer and +36.2 WER defense resistance over baselines.
-
Gauge-covariant stochastic neural fields: Stability and finite-width effects
A gauge-covariant stochastic neural field theory is introduced that derives the maximal Lyapunov exponent and amplification factor, showing finite-width effects as perturbative corrections to dressed kernels that leave the marginality condition unchanged for fixed kernel geometry.
-
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
SpeechParaling-Bench is a new evaluation framework for paralinguistic-aware speech generation that reveals major limitations in current large audio-language models.
-
Deep Learning Scaling is Predictable, Empirically
Deep learning generalization error follows power-law scaling with training set size across multiple domains, with model size scaling sublinearly with data size.
-
Mixed Precision Training
Mixed precision training uses FP16 for most computations, FP32 master weights for accumulation, and loss scaling to enable accurate training of large DNNs with halved memory usage.
-
End-to-End Intracortical Speech Decoding from Neural Activity
End-to-end Conformer decoder achieves 23.8% character error rate on intracortical recordings for speech without external language models.
-
Sink or SWIM: Tackling Real-Time ASR at Scale
SWIM scales Whisper ASR to 20 concurrent multilingual clients via buffer merging, achieving ~2.4s delay at 5 clients versus 3.4s for single-client baselines while preserving accuracy.
-
Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
DPOP is a new loss function that prevents DPO from lowering preferred response likelihoods and outperforms standard DPO on diverse datasets, MT-Bench, and enables Smaug-72B to exceed 80% on the Open LLM Leaderboard.
-
Fine-grained robust prosody transfer for single-speaker neural text-to-speech
Decouples prosody alignment via pre-computed phoneme timestamps and adds VAE to achieve robust fine-grained prosody transfer in single-speaker neural TTS from unseen speakers.
-
Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion
Gated fusion of fastText and BERT embeddings into an end-to-end ASR model captures multi-sentence conversational context and lowers word error rate on the Switchboard corpus.
-
SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation
A multimodal adversarial attack using stage-sampled image nullification and cross-attention flattening degrades lip-sync and facial dynamics in Hallo-based talking-head generation.
-
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
MLLM-enabled video translation is usefully framed as three roles—Semantic Reasoner, Expressive Performer, and Visual Synthesizer—rather than a cascade of ASR, MT, TTS, and lip-sync.
-
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
Direct prompting scales more consistently than CoT prompting for speech-to-text translation as the amount of S2TT data increases.
-
Data-Efficient On-Policy Distillation for Automatic Speech Recognition
On-policy distillation from a Qwen-ASR teacher improves a 0.6B Ark-ASR model over supervised fine-tuning and a same-scale baseline on four of five ASR benchmarks using 100k hours of speech.
-
Deepfake audio as a data augmentation technique for training automatic speech to text transcription models
The authors propose and test a data augmentation framework based on deepfake audio to improve training of speech-to-text transcription models.
-
Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition
A survey of spatial speech perception systems covering sound source localization, directional enhancement, and automatic speech recognition methods and their integration.