REVIEW 32 cited by
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present SpecAugment, a simple data augmentation method for speech recognition. SpecAugment is applied directly to the feature inputs of a neural network (i.e., filter bank coefficients). The augmentation policy consists of warping the features, masking blocks of frequency channels, and masking blocks of time steps. We apply SpecAugment on Listen, Attend and Spell networks for end-to-end speech recognition tasks. We achieve state-of-the-art performance on the LibriSpeech 960h and Swichboard 300h tasks, outperforming all prior work. On LibriSpeech, we achieve 6.8% WER on test-other without the use of a language model, and 5.8% WER with shallow fusion with a language model. This compares to the previous state-of-the-art hybrid system of 7.5% WER. For Switchboard, we achieve 7.2%/14.6% on the Switchboard/CallHome portion of the Hub5'00 test set without the use of a language model, and 6.8%/14.1% with shallow fusion, which compares to the previous state-of-the-art hybrid system at 8.3%/17.3% WER.
Forward citations
Cited by 32 Pith papers
-
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
SQuTR is a large bilingual benchmark of 37,317 synthesized spoken queries under clean/low/medium/high noise, showing that retrieval quality steadily degrades as noise increases.
-
A Non-autoregressive Model for Joint STT and TTS
A joint non-autoregressive model handles both STT and TTS in one framework, beating its own STT baseline and matching its TTS baseline with extra unpaired data and iterative refinement.
-
Physiological Noise Augmentation Improves Non-Invasive Brain-to-Speech
PNA decomposes MEG recordings via ICA, isolates artifact components using EOG/ECG references, and re-injects scaled artifacts into clean data to train decoders that are invariant to physiological noise, improving imag...
-
DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
DHAuDS is a new audio benchmark that corrupts four existing datasets with dynamically varying and diverse acoustic noise, and evaluates three classifiers under test-time adaptation.
-
Text Reinforcement for Multimodal Time Series Forecasting
Reinforcement learning trains an LLM to generate improved text from time series, improving multimodal forecasting on Time-MMD.
-
Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification
Replacing square spectrogram patches with full-frequency temporal patches plus patch-aligned masking improves audio classification accuracy and reduces compute for Transformer and Mamba models.
-
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
LLaSO releases a 3.8B speech-language model, 25.5M training instances, and an evaluation benchmark, claiming a normalized score of 0.72.
-
Scaling and Distilling Transformer Models for sEMG
Vanilla transformers on the emg2qwerty dataset improve cross-user typing accuracy up to 109M parameters, and simple logit distillation recovers most of the gain in a 2.2M-parameter student.
-
Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics
Output-space adversarial training improved clean-data performance and adversarial robustness of two bird sound classifiers across seven soundscape test sets, and stabilized prototype-based explanations.
-
Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World
Adversarial audio noise, optimized with audio augmentations, can trigger and distort the behavior of audio-based LLMs both digitally and when played through the air.
-
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
SSLAM pre-trains audio transformers on partially mixed audio clips with a source retention loss, improving polyphonic sound tagging while keeping monophonic benchmark scores.
-
Measuring Diversity in Synthetic Datasets
DCScore measures dataset diversity as the sum of self-classification probabilities under a softmax similarity matrix, and the paper shows it tracks generation temperature, human judgment, and LLM rankings.
-
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
A voice-conversion augmentation that preserves far-field acoustics while transplanting near-field speaker identity improves FFSVC2020 verification in the training phase, but the headline test-time results use the test...
-
EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
EdgeSpot-4 lifts 10-shot accuracy at 1% false-alarm rate from 73.7% to 82.0% on Google Speech Commands with 29.4M MACs and 128k parameters.
-
Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
A Bi-LSTM variant of ECAPA-TDNN's Res2Block cuts speaker-verification EER by 23% on VoxCeleb1-O at nearly the same parameter count.
-
Beyond Words: Interjection Classification for Improved Human-Computer Interaction
Interjection classification over five speakers improves when pitch, tempo, and background-noise augmentation is added, but absolute accuracy stays below 60% on unseen speakers.
-
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
AD-AVSR combines dual-stream audio encoding, audio-guided visual refinement, visual-guided noise suppression, and thresholded audio-visual pair selection to improve audio-visual speech recognition word error rates und...
-
Improving Stereo 3D Sound Event Localization and Detection: Perceptual Features, Stereo-specific Data Augmentation, and Distance Normalization
A CRNN with mid-side intensity, spatial coherence, stereo channel swapping, FilterAugment, frequency shifting, and distance normalization improves stereo 3D SELD on STARSS23.
-
From Sharpness to Better Generalization for Speech Deepfake Detection
Sharpness, a measure of loss sensitivity to weight perturbations, correlates with speech deepfake detection error on unseen data, and Sharpness-Aware Minimization reduces both sharpness and error in most settings.
-
Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer
A variational Bayesian method adapts deep acoustic models by estimating distributions over hidden features, with a Gaussian mean-field variant for parallel data and an empirical Bayes variant for non-parallel data, an...
-
Unifying Prediction and Explanation in Time-Series Transformers via Shapley-based Pretraining
A time-series transformer trained with Shapley-based and contrastive pretraining outputs predictions and Shapley explanations in one forward pass, at a fraction of post-hoc explanation cost.
-
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation
Multi-prototype refinement plus latent-space augmentation reduces equal error rate for audio deepfake detection on ASVspoof 2019/2021 and In-The-Wild benchmarks.
-
LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
A graph neural network that mixes k-nearest-neighbor and fuzzy C-means cluster features outperforms transformer baselines on AudioSet, FSD50K, and ESC-50.
-
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
NIM4-ASR delivers SOTA ASR performance on public benchmarks using a 2.3B-parameter LLM with multi-stage training, real-time streaming, and million-scale hotword customization via RAG.
-
Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
A ResNet-Conformer network with shared weights and multi-scale attention, trained on synthetic and augmented real audio, improves detection and direction estimation over the DCASE 2024 baseline while distance error st...
-
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
FairASR pretrains a Conformer with InfoNCE plus a gradient-reversed supervised contrastive loss over demographic labels, reducing demographic WER gaps on FairSpeech with small overall WER cost.
-
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
Fine-tuning SeamlessM4T on 20 hours of Bhojpuri-Hindi data with tuned hyperparameters and SpecAugment reaches 36.4 dev BLEU but only 9.9 test BLEU in the IWSLT 2025 low-resource task.
-
Quantum Approaches for Dysphonia Assessment in Small Speech Datasets
Quanvolutional neural networks outperformed classical CNNs on dysphonia classification from small Mel-spectrogram datasets, though the experimental design is flawed.
-
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
Combining knowledge distillation with l0 or low-rank pruning improves compressed RNN-T ASR, and joint pruning with fine-tuning gives 8.9% and 13.4% relative WER gains over baseline.
-
Technical Report: A Practical Guide to Kaldi ASR Optimization
The paper proposes engineering tweaks to Kaldi ASR (Conformer+TDNN-F architecture, SpecAugment, Bayesian n-gram merging) but presents no experimental evidence for any claimed improvement.
-
Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets
A funnel-shaped TDNN plus 1x1 layers is reported to reach 97% language identification accuracy on ten Common Voice languages, but the evaluation is insufficient to support the claim.
-
A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
A robotics-focused review of sound source localization research, emphasizing deep learning architectures, datasets, and open challenges.
Discussion (0). Continue with ORCID to comment.