REVIEW 3 cited by
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
T0 review · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The top MISP 2025 systems achieve DER 8.09%, CER 9.48%, and cpCER 11.56%, far outperforming the provided audio-visual baselines.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Extended reading notes
Core claim
The abstract states: 'The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.' If true, the MISP 2025 challenge produced a new state of the art on a new audio-visual meeting corpus, with the top AVDR result being an order-of-magnitude improvement over the baseline cpCER.
Load-bearing premise
The evaluation and ground-truth construction rely on manual synchronization of three independent clocks (microphone array, camera, recorder) by detecting a cup knock (Section 2.2); if this alignment is inaccurate, all diarization and recognition metrics on the evaluation set are unreliable. A second load-bearing premise is that the reported leaderboard numbers are comparable across teams despite different post-processing and external pretraining data, an assumption that is not tested statistically in the paper.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
assumptions (3)
- domain assumption Evaluation ground truth is accurate and unbiased
- domain assumption Train/dev/eval partitions have no speaker or room overlap
- domain assumption No-score collar is not applied and overlapping speech is scored
Cite this review
Pith. "Pith review of The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition." pith.science (2026). https://pith.science/paper/ZPBEVIY5
@misc{pith2026250513971,
author = {Pith},
title = {Pith review of: The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZPBEVIY5}},
note = {Machine review of arXiv:2505.13971}
}
read the original abstract
Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal, multi-device meeting transcription by incorporating video modality alongside audio. The tasks include Audio-Visual Speaker Diarization (AVSD), Audio-Visual Speech Recognition (AVSR), and Audio-Visual Diarization and Recognition (AVDR). We present the challenge's objectives, tasks, dataset, baseline systems, and solutions proposed by participants. The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.
Figures
Forward citations
Cited by 3 Pith papers
-
MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings
A new benchmark shows MLLMs underperform humans on meeting Theory-of-Mind tasks, especially detecting pseudo-consensus and hidden dissent.
-
M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset
Release of M3SD, a 770+ hour pseudo-labeled multi-scenario, multi-language audio-visual speaker diarization dataset, built from YouTube and Bilibili videos without manual annotation.
-
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
A hybrid diarization and ASR system with a CER-supervised bridging module achieved the best results in two MISP 2025 tracks.
Reference graph
Works this paper leans on
-
[1]
Introduction In recent years, the proliferation of speech-enabled applica- tions has led to increasingly complex usage scenarios, such as home environments, and professional meetings. These sce- narios present considerable challenges due to adverse acoustic conditions, including far-field audio, background noise, and re- verberation. Additionally, convers...
arXiv 2021
-
[2]
Dataset 2.1. Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios. The dataset is partitioned into three distinct subsets: 119 hours for training (Train), 3 hours for development (Dev), and 3 hours for evaluation (Eval), the latter serving as ...
-
[3]
Challenge Description 3.1. Task 1: Audio-Visual Speaker Diarization (A VSD) Audio-visual speaker diarization aims to solve the “who spoke when” problem by labeling speech timestamps with classes cor- responding to speaker identity using multi-speaker audio and video data. Training and development sets provide all audio and video recordings along with the ...
-
[4]
Recognition results and reference transcriptions belonging to the same speaker are concatenated on the timeline in a ses- sion
-
[5]
(2), whereNspk is the total number of speakers in the session
CERs between the reference and all possible speaker permu- tations of the hypothesis {si|i = 0, 1, · · ·, P Nspk Nspk } are cal- culated as Eq. (2), whereNspk is the total number of speakers in the session
-
[6]
The lowest CER as the cpCER, the process is as follows: cpCER = min {si|i=0,1,··· ,P Nspk Nspk } CERi (3)
-
[7]
Results and Analysis 4.1. A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge. The table lists the participating teams, their pro- posed methods, pre-training strategies, external datasets used, the modality employed, and the corresponding Diarization Er- ror Rate (DER) achi...
work page 2025
-
[8]
Conclusion In this paper, we presented the MISP 2025 Challenge, a multi- modal speech processing benchmark focused on meeting sce- narios. We introduced the large-scale audio-visual meeting dataset (MISP-Meeting), provided comprehensive baseline sys- tems for audio-visual speaker diarization, speech recognition, and joint diarization and recognition, and ...
work page 2025
Show all 39 references
-
[9]
The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,
D. Mostefa, N. Moreau, K. Choukri, G. Potamianos, S. M. Chu, A. Tyagi, J. R. Casas, J. Turmo, L. Cristoforetti, F. Tobia et al., “The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,” Language resources and evaluation , vol. 41, pp. 389–407, 2007
2007
-
[10]
M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,
F. Yu, S. Zhang, Y . Fu, L. Xie, S. Zheng, Z. Du, W. Huang, P. Guo, Z. Yan, B. Ma et al., “M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)....
2022
-
[11]
Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,
Y . Fu, L. Cheng, S. Lv, Y . Jv, Y . Kong, Z. Chen, Y . Hu, L. Xie, J. Wu, H. Bu et al., “Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,” arXiv preprint arXiv:2104.03603, 2021
2021 arXiv
-
[12]
Continuous speech separation: Dataset and analysis,
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y . Luo, J. Wu, X. Xiao, and J. Li, “Continuous speech separation: Dataset and analysis,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7284–7288
2020
-
[13]
Lib- rispeech: an asr corpus based on public domain audio books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015, pp. 5206–5210
2015
-
[14]
CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,
S. Watanabe, M. Mandel, J. Barker et al. , “CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,” in Proc. CHiME 2020, 2020, pp. 1–7
2020
-
[15]
The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,
H. Chen, H. Zhou, J. Du, C.-H. Lee, J. Chen, S. Watanabe, S. M. Siniscalchi, O. Scharenborg, D.-Y . Liu, B.-C. Yin et al. , “The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,” in ICASSP 2022-2022 IEEE International...
2022
-
[16]
The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,
Z. Wang, S. Wu, H. Chen, M.-K. He, J. Du, C.-H. Lee, J. Chen, S. Watanabe, S. Siniscalchi, O. Scharenborg et al. , “The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,” in ICASSP 2023-2023 IEEE International C...
2022
-
[17]
The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,
S. Wu, C. Wang, H. Chen, Y . Dai, C. Zhang, R. Wang, H. Lan, J. Du, C.-H. Lee, J. Chen et al. , “The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,” in ICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics,...
2023
-
[18]
MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,
H. Chen, C.-H. H. Yang, J.-C. Gu, S. M. Siniscalchi, and J. Du, “MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,” inProceed- ings of the 63st Annual Meeting of the Association for Compu- tational Linguistics (Volum...
2025
-
[19]
End-to-end audio-visual neural speaker diarization,
M. He, J. Du, and C.-H. Lee, “End-to-end audio-visual neural speaker diarization,” in Proc. Interspeech 2022, 2022, pp. 1461– 1465
2022
-
[20]
Conformer: Convolution- augmented transformer for speech recognition,
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu et al. , “Conformer: Convolution- augmented transformer for speech recognition,” arXiv preprint arXiv:2005.08100, 2020
2005 arXiv
-
[21]
Dover-lap: A method for com- bining overlap-aware diarization outputs,
D. Raj, L. P. Garcia-Perera, Z. Huang, S. Watanabe, D. Povey, A. Stolcke, and S. Khudanpur, “Dover-lap: A method for com- bining overlap-aware diarization outputs,” in 2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2021, pp. 881– 888
2021
-
[22]
The rich transcription 2006 spring meeting recognition evaluation,
J. G. Fiscus, J. Ajot, M. Michel et al., “The rich transcription 2006 spring meeting recognition evaluation,” in Proc. MLMI 2006 , 2006, pp. 309–322
2006
-
[23]
Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,
Y . Dai, H. Chen, J. Du et al. , “Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,” in Proc. ICME 2023, 2023, pp. 2627–2632
2023
-
[24]
Gpu-accelerated guided source separation for meeting transcription,
D. Raj, D. Povey, and S. Khudanpur, “Gpu-accelerated guided source separation for meeting transcription,” arXiv preprint arXiv:2212.05271, 2022
2022 arXiv
-
[25]
The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,
Z. Wang, S. Wu, H. Chen et al. , “The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,” in Proc. ICASSP 2023, 2023, pp. 1–5
2022
-
[26]
V oxblink2: A 100k+ speaker recognition corpus and the open-set speaker-identification benchmark,
Y . Lin, M. Cheng, F. Zhang, Y . Gao, S. Zhang, and M. Li, “V oxblink2: A 100k+ speaker recognition corpus and the open-set speaker-identification benchmark,” arXiv preprint arXiv:2407.11510, 2024
2024 arXiv
-
[27]
V oxceleb2: Deep speaker recognition,
J. S. Chung, A. Nagrani, and A. Zisserman, “V oxceleb2: Deep speaker recognition,” arXiv preprint arXiv:1806.05622, 2018
2018 arXiv
-
[28]
Kespeech: An open source speech dataset of mandarin and its eight subdialects,
Z. Tang, D. Wang, Y . Xu, J. Sun, X. Lei, S. Zhao, C. Wen, X. Tan, C. Xie, S. Zhou et al., “Kespeech: An open source speech dataset of mandarin and its eight subdialects,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Bench- marks Track (Roun...
2021
-
[29]
3d-speaker: A large-scale multi-device, multi-distance, and multi-dialect cor- pus for speech representation disentanglement,
S. Zheng, L. Cheng, Y . Chen, H. Wang, and Q. Chen, “3d-speaker: A large-scale multi-device, multi-distance, and multi-dialect cor- pus for speech representation disentanglement,” arXiv preprint arXiv:2306.15354, 2023
2023 arXiv
-
[30]
Ecapa- tdnn: Emphasized channel attention, propagation and ag- gregation in tdnn based speaker verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa- tdnn: Emphasized channel attention, propagation and ag- gregation in tdnn based speaker verification,” arXiv preprint arXiv:2005.07143, 2020
2005 arXiv
-
[31]
Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiaoet al., “Wavlm: Large-scale self- supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
-
[32]
The ami meeting corpus,
W. Kraaij, T. Hain, M. Lincoln, and W. Post, “The ami meeting corpus,” in Proc. International Conference on Methods and Tech- niques in Behavioral Research, 2005, pp. 1–4
2005
-
[33]
Lip-reading with densely connected temporal convolutional networks,
P. Ma, Y . Wang, J. Shen, S. Petridis, and M. Pantic, “Lip-reading with densely connected temporal convolutional networks,” inPro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021, pp. 2857–2866
2021
-
[34]
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning . PMLR, 2023, pp. 28 492–28 518
2023
-
[35]
Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,
B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng et al., “Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). ...
2022
-
[36]
Sequence-to-sequence neural di- arization with automatic speaker detection and representation,
M. Cheng, Y . Lin, and M. Li, “Sequence-to-sequence neural di- arization with automatic speaker detection and representation,” arXiv preprint arXiv:2411.13849, 2024
2024 arXiv
-
[37]
Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,
S. Zhao, Y . Ma, C. Ni, C. Zhang, H. Wang, T. H. Nguyen, K. Zhou, J. Q. Yip, D. Ng, and B. Ma, “Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,” in ICASSP 2024-2024 IEEE International Conference on Acoust...
2024
-
[38]
Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,
Z. Gao, S. Zhang, I. McLoughlin, and Z. Yan, “Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,” arXiv preprint arXiv:2206.08317, 2022
2022 arXiv
-
[39]
Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,
T. Yoshioka and T. Nakatani, “Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,” IEEE Transactions on Audio, Speech, and Language Pro- cessing, vol. 20, no. 10, pp. 2707–2720, 2012
2012
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.