Pith. sign in

hub Mixed citations

Conformer: Convolution-augmented Transformer for Speech Recognition

Mixed citation behavior. Most common role is method (60%).

31 Pith papers citing it
385 external citations · Pith
Method 60% of classified citations
abstract

Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs). Transformer models are good at capturing content-based global interactions, while CNNs exploit local features effectively. In this work, we achieve the best of both worlds by studying how to combine convolution neural networks and transformers to model both local and global dependencies of an audio sequence in a parameter-efficient way. To this regard, we propose the convolution-augmented transformer for speech recognition, named Conformer. Conformer significantly outperforms the previous Transformer and CNN based models achieving state-of-the-art accuracies. On the widely used LibriSpeech benchmark, our model achieves WER of 2.1%/4.3% without using a language model and 1.9%/3.9% with an external language model on test/testother. We also observe competitive performance of 2.7%/6.3% with a small model of only 10M parameters.

hub tools

citation-role summary

method 3 background 2

citation-polarity summary

representative citing papers

How to Evaluate Speech Translation with Source-Aware Neural MT Metrics

cs.CL · 2025-11-05 · unverdicted · novelty 7.0

Source-aware MT metrics adapted to speech translation via ASR transcripts or back-translations as audio proxies, plus a new cross-lingual re-segmentation algorithm, improve correlation with human judgments over reference-only baselines.

TRADE: Transducer-Augmented Decoder for Speech LLM

cs.CL · 2026-06-07 · unverdicted · novelty 6.0

TRADE augments multimodal Speech LLMs with a transducer branch for streaming ASR, reporting 6.71% WER offline and 8.40% streaming on the Open ASR Leaderboard from one checkpoint.

Executable Boundary Contracts for Sound Event Traces

cs.LO · 2026-05-19 · unverdicted · novelty 6.0

Defines executable boundary contracts for sound event traces using an STL-embeddable Boolean fragment plus interval and duration clauses, then evaluates them on speech and soundscape data where they disagree with standard scores.

DGSNA: Dynamic Generative Scene-based Noise Addition method

cs.SD · 2024-11-19 · unverdicted · novelty 6.0

DGSNA dynamically generates scene-specific noise via prompt-driven language models and text-to-audio diffusion, then mixes it with speech to improve recognition and keyword spotting robustness by up to 11.32%.

UniVoice: A Unified Model for Speech and Singing Voice Generation

cs.SD · 2026-06-04 · unverdicted · novelty 5.0

UniVoice is a conditional flow matching model with a Diffusion Transformer backbone that unifies TTS and SVS via modality-specific encoders and a null melody token for speech, achieving 5.26% speech PER and 16.22% singing PER.

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

eess.AS · 2026-04-20 · conditional · novelty 4.0 · 2 refs

A 2.3B-parameter LLM-based ASR system achieves competitive recognition accuracy and reduced hallucination through a multi-stage training paradigm with asynchronous encoder updates, ASR-specialized RL, and phoneme-level RAG for hotword customization.

Non-Intrusive Automatic Speech Recognition Refinement: A Survey

eess.AS · 2025-08-10 · accept · novelty 4.0

A survey that classifies non-intrusive ASR refinement methods into five categories, reviews domain adaptation and evaluation datasets, proposes standardized metrics, and identifies future research directions.

citing papers explorer

Showing 31 of 31 citing papers.