Pith. sign in

REVIEW 23 cited by

Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04675 v2 pith:UKU6VZHO submitted 2024-07-05 eess.AS cs.SD

classification eess.AScs.SD
keywords seed-asrspeechmodelslanguagemodelrecognitionscenariosaccents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in data matching scenarios and are gradually approaching a bottleneck. In this work, we introduce Seed-ASR, a large language model (LLM) based speech recognition model. Seed-ASR is developed based on the framework of audio conditioned LLM (AcLLM), leveraging the capabilities of LLMs by inputting continuous speech representations together with contextual information into the LLM. Through stage-wise large-scale training and the elicitation of context-aware capabilities in LLM, Seed-ASR demonstrates significant improvement over end-to-end models on comprehensive evaluation sets, including multiple domains, accents/dialects and languages. Additionally, Seed-ASR can be further deployed to support specific needs in various scenarios without requiring extra language models. Compared to recently released large ASR models, Seed-ASR achieves 10%-40% reduction in word (or character, for Chinese) error rates on Chinese and English public test sets, further demonstrating its powerful performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

    cs.CV 2026-08 conditional novelty 7.0 of 10

    InteracVid delivers 454K livestream-derived context-query-response triplets, pairing real or LLM-reconstructed chat triggers with real audio-video reactions, and shows fine-tuning gains on genuine queries.

  2. Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A distilled multilingual ASR student, trained by on-policy distillation from language-specialized RL teachers, outperforms the best individual teacher and several larger open-source models.

  3. Context-Aware ASR for Mandarin Technical Lectures

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Self-built lecture glossaries from first-pass ASR raise technical-term recall across five backbones while holding or lowering CER on a new Mandarin AI/ML lecture benchmark.

  4. TRADE: Transducer-Augmented Decoder for Speech LLM

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    TRADE augments multimodal Speech LLMs with a transducer branch for streaming ASR, reporting 6.71% WER offline and 8.40% streaming on the Open ASR Leaderboard from one checkpoint.

  5. Audio Interaction Model

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    Audio-Interaction unifies offline and online audio tasks into one streaming model via the SoundFlow framework and a new 2.6M-item streaming corpus, enabling real-time instruction following and proactive responses.

  6. JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    JSPG jointly combines semantic, pinyin, and glyph retrieval with an extended Smith-Waterman algorithm to dynamically filter keyword dictionaries and improve accuracy in Chinese contextual ASR.

  7. VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models

    cs.SD 2026-05 unverdicted novelty 6.0 of 10

    VocalParse applies interleaved and Chain-of-Thought prompting to a Large Audio Language Model to jointly transcribe lyrics, melody and word-note alignments, achieving state-of-the-art results on multiple singing datasets.

  8. Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs

    eess.AS 2026-04 unverdicted novelty 6.0 of 10

    A multi-stage training method for LLM-based ASR uses new entropy allocation metrics to achieve competitive benchmark performance with 2.3B parameters while mitigating hallucinations via better encoder-LLM decoupling.

  9. LLMs and Speech: Integration vs. Combination

    eess.AS 2026-03 conditional novelty 6.0 of 10

    With matched data and sizes, CTC+LLM shallow fusion beats tight speech-LLM integration on in-domain ASR, while prefix LLMs win average WER on out-of-domain HuggingFace sets.

  10. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  11. Step-Audio 2 Technical Report

    cs.CL 2025-07 unverdicted novelty 6.0 of 10

    Step-Audio 2 integrates a latent audio encoder, reasoning-centric reinforcement learning, and discrete audio token generation into language modeling to deliver state-of-the-art performance on audio understanding and c...

  12. ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

    cs.SD 2026-07 reject novelty 5.0 of 10

    A 4B-parameter LLM ASR system with five multi-token-prediction branches reports 2.97% CER Chinese, 3.68% WER English, 3.70% long-form WER, and a 0.0053 real-time factor, but the acceptance-rate calculation and ablatio...

  13. StepAudio 2.5 Technical Report

    eess.AS 2026-05 unverdicted novelty 5.0 of 10

    StepAudio 2.5 is a unified audio-language foundation model that reaches state-of-the-art results on ASR, TTS, and realtime interaction by using task-tailored RLHF on a shared backbone.

  14. Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    The authors introduce LLM-based semantic judgment and an agentic interaction loop that improves semantic fidelity and enables iterative corrections in automatic speech recognition beyond traditional WER.

  15. From Synthesis to Clinical Assistance: A Strategy-Aware Agent Framework for Autism Intervention based on Real Clinical Dataset

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    ASDAgent generates synthetic ABA-strategy dialogues that match human therapist distributions (KL 0.083) and achieves 80% expert consistency, while its outputs improve small language models for therapeutic tasks.

  16. Cross-Learning Fine-Tuning Strategy for Dysarthric Speech Recognition Via CDSD database

    cs.SD 2025-08 conditional novelty 5.0 of 10

    Joint fine-tuning on seven dysarthric speakers' data reduced per-speaker character error rates by up to 13.15 percentage points compared to single-speaker fine-tuning on the CDSD corpus.

  17. Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems

    eess.AS 2025-08 conditional novelty 5.0 of 10

    Adding cross-utterance audio context to Conformer-Transducer ASR reduces WER/CER by 0.5 to 1.1 absolute points on four benchmarks, and a splicing-based batch scheme cuts training time by up to about 19%.

  18. Qwen2.5-Omni Technical Report

    cs.CL 2025-03 conditional novelty 5.0 of 10

    Qwen2.5-Omni presents a multimodal model with block-wise encoders, TMRoPE position embeddings, and a Thinker-Talker architecture that enables simultaneous text and streaming speech generation while matching text perfo...

  19. Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

    cs.CL 2026-07 unverdicted novelty 4.0 of 10

    JSTIP interleaves speech and text sequences during pretraining on 38k hours of ASR data to improve entity accuracy over ASR-only and simple joint-training baselines while matching performance from domain text.

  20. NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

    eess.AS 2026-04 conditional novelty 4.0 of 10

    A 2.3B-parameter LLM-based ASR system achieves competitive recognition accuracy and reduced hallucination through a multi-stage training paradigm with asynchronous encoder updates, ASR-specialized RL, and phoneme-leve...

  21. NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

    eess.AS 2026-04 unverdicted novelty 4.0 of 10

    NIM4-ASR delivers SOTA ASR performance on public benchmarks using a 2.3B-parameter LLM with multi-stage training, real-time streaming, and million-scale hotword customization via RAG.

  22. LLMs and Speech: Integration vs. Combination

    eess.AS 2026-03 unverdicted novelty 4.0 of 10

    Tight integration of acoustic models with LLMs for ASR is ablated against shallow fusion across label units, fine-tuning strategies, LLM sizes, and joint CTC decoding to mitigate hallucinations.

  23. Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition

    cs.SD 2025-09 reject novelty 4.0 of 10

    An LLM-based ASR error correction framework with noise-adaptive encoding and dynamic multi-modal fusion reports WER gains, but its fusion weights require ground-truth text at inference.

Pith tools