AISHELL-5 releases 100+ hours of real in-car multi-channel multi-speaker Mandarin speech, 40 hours of noise, and a baseline showing mainstream ASR models still fail badly on this task.
The baseline system con- sists of two primary sub-modules: speech frontend processing and automatic speech recognition (ASR)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
AISHELL-5 releases 100+ hours of real in-car multi-channel multi-speaker Mandarin speech, 40 hours of noise, and a baseline showing mainstream ASR models still fail badly on this task.