REVIEW 6 cited by
The CHiME-7 DASR Challenge: Distant Meeting Transcription with Multiple Devices in Diverse Scenarios
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The CHiME challenges have played a significant role in the development and evaluation of robust automatic speech recognition (ASR) systems. We introduce the CHiME-7 distant ASR (DASR) task, within the 7th CHiME challenge. This task comprises joint ASR and diarization in far-field settings with multiple, and possibly heterogeneous, recording devices. Different from previous challenges, we evaluate systems on 3 diverse scenarios: CHiME-6, DiPCo, and Mixer 6. The goal is for participants to devise a single system that can generalize across different array geometries and use cases with no a-priori information. Another departure from earlier CHiME iterations is that participants are allowed to use open-source pre-trained models and datasets. In this paper, we describe the challenge design, motivation, and fundamental research questions in detail. We also present the baseline system, which is fully array-topology agnostic and features multi-channel diarization, channel selection, guided source separation and a robust ASR model that leverages self-supervised speech representations (SSLR).
Forward citations
Cited by 6 Pith papers
-
Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
An LLM conditioned on speaker embeddings and utterance time boundaries jointly transcribes and timestamps overlapping multi-speaker speech.
-
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
A pseudo-supervised method that uses close-talk microphone recordings to estimate training targets lets a speech enhancement model adapt to real far-field data, cutting CER to 29.80% on MISP2023.
-
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset
Release of M3SD, a 770+ hour pseudo-labeled multi-scenario, multi-language audio-visual speaker diarization dataset, built from YouTube and Bilibili videos without manual annotation.
-
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
Training a neural speech enhancer on pseudo labels derived from aligned close-talk recordings improves far-field meeting ASR, reaching second place in the MISP-Meeting Challenge.
-
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
A challenge system combining speaker diarization, speaker embeddings, and a Qwen2.5 LLM adapter architecture reports 18.08% tcpWER on multilingual multi-speaker ASR, far below the 60.39% baseline.
-
Exploring Speaker Diarization with Mixture of Experts
A speaker diarization system that adds a shared-and-soft mixture-of-experts layer to a memory-augmented sequence-to-sequence model reports lower error rates on CHiME-6, DiPCo, and Mixer 6, with some DIHARD-III claims ...
Discussion (0). Continue with ORCID to comment.