Pith. sign in

REVIEW 3 cited by

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The top MISP 2025 systems achieve DER 8.09%, CER 9.48%, and cpCER 11.56%, far outperforming the provided audio-visual baselines.

arxiv 2505.13971 v2 pith:ZPBEVIY5 submitted 2025-05-20 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords achievedaudio-visualdiarizationchallengeerrorimprovingraterecognition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Meetings are hard for speech systems because people talk over each other and the room adds noise and echo. The MISP 2025 challenge gave research teams a new set of recorded meetings, with audio from an eight-microphone array and video from a panoramic camera. Teams had to build systems that figure out who is speaking when (diarization), what they are saying (recognition), or both at once. The organizers provided simple baseline systems and scored every submission. According to the paper, the best teams beat the baselines by wide margins: the best diarization error rate fell from 15.52% to 8.09%, the best character error rate fell from 20.10% to 9.48%, and the best combined diarization plus recognition score fell from 84.05% to 11.56%. An interesting pattern is that the winning audio-visual systems used the video only weakly; the top performers relied mostly on audio. The paper is a challenge report: it describes the data, the rules, the baselines, and the top solutions, but it does not deeply analyze why video mattered little, and it does not provide enough details to reproduce the experiments. There is also a misleading number in the abstract: it calls the 72.49 point drop in the combined score a 72.49% improvement, when the correct relative improvement is 86.25%.
Extended reading notes

Core claim

The abstract states: 'The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.' If true, the MISP 2025 challenge produced a new state of the art on a new audio-visual meeting corpus, with the top AVDR result being an order-of-magnitude improvement over the baseline cpCER.

Load-bearing premise

The evaluation and ground-truth construction rely on manual synchronization of three independent clocks (microphone array, camera, recorder) by detecting a cup knock (Section 2.2); if this alignment is inaccurate, all diarization and recognition metrics on the evaluation set are unreliable. A second load-bearing premise is that the reported leaderboard numbers are comparable across teams despite different post-processing and external pretraining data, an assumption that is not tested statistically in the paper.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper makes no parametric scientific claims; the central inputs are the collected dataset and the challenge design. The listed domain assumptions underlie the reliability of the metrics and the comparability of results. No new physical or theoretical entities are introduced.

assumptions (3)
  • domain assumption Evaluation ground truth is accurate and unbiased
    The paper assumes manual clock synchronization (cup knock) and headset microphone transcriptions with SNR > 15 dB yield sufficiently correct speaker and word labels (Section 2.2).
  • domain assumption Train/dev/eval partitions have no speaker or room overlap
    Stated in Section 2.1; ensures evaluation measures generalization, but is assumed true without independent verification.
  • domain assumption No-score collar is not applied and overlapping speech is scored
    Scoring rule stated in Section 3.1; a design choice that affects DER comparability with challenges that use collars.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition." pith.science (2026). https://pith.science/paper/ZPBEVIY5

@misc{pith2026250513971,
  author       = {Pith},
  title        = {Pith review of: The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZPBEVIY5}},
  note         = {Machine review of arXiv:2505.13971}
}
read the original abstract

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal, multi-device meeting transcription by incorporating video modality alongside audio. The tasks include Audio-Visual Speaker Diarization (AVSD), Audio-Visual Speech Recognition (AVSR), and Audio-Visual Diarization and Recognition (AVDR). We present the challenge's objectives, tasks, dataset, baseline systems, and solutions proposed by participants. The best-performing systems achieved significant improvements over the baseline: the top AVSD model achieved a Diarization Error Rate (DER) of 8.09%, improving by 7.43%; the top AVSR system achieved a Character Error Rate (CER) of 9.48%, improving by 10.62%; and the best AVDR system achieved a concatenated minimum-permutation Character Error Rate (cpCER) of 11.56%, improving by 72.49%.

Figures

Figures reproduced from arXiv: 2505.13971 by the authors.

Figure 1
Figure 1. Statistics of meeting rooms. 2. Dataset 2.1. Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios. The dataset is partitioned into three distinct subsets: 119 hours for training (Train), 3 hours for development (Dev), and 3 hours for evaluation (Eval), the latter serving as the basis for rigorous scoring… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new benchmark shows MLLMs underperform humans on meeting Theory-of-Mind tasks, especially detecting pseudo-consensus and hidden dissent.

  2. M3SD: Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset

    eess.AS 2025-06 reject novelty 5.0 of 10

    Release of M3SD, a 770+ hour pseudo-labeled multi-scenario, multi-language audio-visual speaker diarization dataset, built from YouTube and Bilibili videos without manual annotation.

  3. Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A hybrid diarization and ASR system with a CER-supervised bridging module achieved the best results in two MISP 2025 tracks.

Reference graph

Works this paper leans on

39 extracted references · 20 canonical work pages · cited by 3 Pith papers

  1. [1]

    These sce- narios present considerable challenges due to adverse acoustic conditions, including far-field audio, background noise, and re- verberation

    Introduction In recent years, the proliferation of speech-enabled applica- tions has led to increasingly complex usage scenarios, such as home environments, and professional meetings. These sce- narios present considerable challenges due to adverse acoustic conditions, including far-field audio, background noise, and re- verberation. Additionally, convers...

  2. [2]

    Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios

    Dataset 2.1. Statics The MISP-Meeting dataset[10] comprises a total of 125 hours of synchronized audio and video data, meticulously curated to reflect real-world meeting scenarios. The dataset is partitioned into three distinct subsets: 119 hours for training (Train), 3 hours for development (Dev), and 3 hours for evaluation (Eval), the latter serving as ...

  3. [3]

    who spoke when

    Challenge Description 3.1. Task 1: Audio-Visual Speaker Diarization (A VSD) Audio-visual speaker diarization aims to solve the “who spoke when” problem by labeling speech timestamps with classes cor- responding to speaker identity using multi-speaker audio and video data. Training and development sets provide all audio and video recordings along with the ...

  4. [4]

    Recognition results and reference transcriptions belonging to the same speaker are concatenated on the timeline in a ses- sion

  5. [5]

    (2), whereNspk is the total number of speakers in the session

    CERs between the reference and all possible speaker permu- tations of the hypothesis {si|i = 0, 1, · · ·, P Nspk Nspk } are cal- culated as Eq. (2), whereNspk is the total number of speakers in the session

  6. [6]

    The lowest CER as the cpCER, the process is as follows: cpCER = min {si|i=0,1,··· ,P Nspk Nspk } CERi (3)

  7. [7]

    A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge

    Results and Analysis 4.1. A VSD Table 2 presents a summary of the methods and results for Track 1 (Audio-Visual Speaker Diarization) in the MISP 2025 Challenge. The table lists the participating teams, their pro- posed methods, pre-training strategies, external datasets used, the modality employed, and the corresponding Diarization Er- ror Rate (DER) achi...

  8. [8]

    Conclusion In this paper, we presented the MISP 2025 Challenge, a multi- modal speech processing benchmark focused on meeting sce- narios. We introduced the large-scale audio-visual meeting dataset (MISP-Meeting), provided comprehensive baseline sys- tems for audio-visual speaker diarization, speech recognition, and joint diarization and recognition, and ...

Show all 39 references
  1. [9]

    The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,

    D. Mostefa, N. Moreau, K. Choukri, G. Potamianos, S. M. Chu, A. Tyagi, J. R. Casas, J. Turmo, L. Cristoforetti, F. Tobia et al., “The chil audiovisual corpus for lecture and meeting analysis in- side smart rooms,” Language resources and evaluation , vol. 41, pp. 389–407, 2007

  2. [10]

    M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

    F. Yu, S. Zhang, Y . Fu, L. Xie, S. Zheng, Z. Du, W. Huang, P. Guo, Z. Yan, B. Ma et al., “M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)....

  3. [11]

    Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,

    Y . Fu, L. Cheng, S. Lv, Y . Jv, Y . Kong, Z. Chen, Y . Hu, L. Xie, J. Wu, H. Bu et al., “Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,” arXiv preprint arXiv:2104.03603, 2021

  4. [12]

    Continuous speech separation: Dataset and analysis,

    Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y . Luo, J. Wu, X. Xiao, and J. Li, “Continuous speech separation: Dataset and analysis,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7284–7288

  5. [13]

    Lib- rispeech: an asr corpus based on public domain audio books,

    V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- rispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015, pp. 5206–5210

  6. [14]

    CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,

    S. Watanabe, M. Mandel, J. Barker et al. , “CHiME-6 Chal- lenge: Tackling Multispeaker Speech Recognition for Unseg- mented Recordings,” in Proc. CHiME 2020, 2020, pp. 1–7

  7. [15]

    The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,

    H. Chen, H. Zhou, J. Du, C.-H. Lee, J. Chen, S. Watanabe, S. M. Siniscalchi, O. Scharenborg, D.-Y . Liu, B.-C. Yin et al. , “The first multimodal information based speech processing (misp) chal- lenge: Data, tasks, baselines and results,” in ICASSP 2022-2022 IEEE International...

  8. [16]

    The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,

    Z. Wang, S. Wu, H. Chen, M.-K. He, J. Du, C.-H. Lee, J. Chen, S. Watanabe, S. Siniscalchi, O. Scharenborg et al. , “The mul- timodal information based speech processing (misp) 2022 chal- lenge: Audio-visual diarization and recognition,” in ICASSP 2023-2023 IEEE International C...

  9. [17]

    The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,

    S. Wu, C. Wang, H. Chen, Y . Dai, C. Zhang, R. Wang, H. Lan, J. Du, C.-H. Lee, J. Chen et al. , “The multimodal information based speech processing (misp) 2023 challenge: Audio-visual target speaker extraction,” in ICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics,...

  10. [18]

    MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,

    H. Chen, C.-H. H. Yang, J.-C. Gu, S. M. Siniscalchi, and J. Du, “MISP-Meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,” inProceed- ings of the 63st Annual Meeting of the Association for Compu- tational Linguistics (Volum...

  11. [19]

    End-to-end audio-visual neural speaker diarization,

    M. He, J. Du, and C.-H. Lee, “End-to-end audio-visual neural speaker diarization,” in Proc. Interspeech 2022, 2022, pp. 1461– 1465

  12. [20]

    Conformer: Convolution- augmented transformer for speech recognition,

    A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu et al. , “Conformer: Convolution- augmented transformer for speech recognition,” arXiv preprint arXiv:2005.08100, 2020

  13. [21]

    Dover-lap: A method for com- bining overlap-aware diarization outputs,

    D. Raj, L. P. Garcia-Perera, Z. Huang, S. Watanabe, D. Povey, A. Stolcke, and S. Khudanpur, “Dover-lap: A method for com- bining overlap-aware diarization outputs,” in 2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2021, pp. 881– 888

  14. [22]

    The rich transcription 2006 spring meeting recognition evaluation,

    J. G. Fiscus, J. Ajot, M. Michel et al., “The rich transcription 2006 spring meeting recognition evaluation,” in Proc. MLMI 2006 , 2006, pp. 309–322

  15. [23]

    Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,

    Y . Dai, H. Chen, J. Du et al. , “Improving audio-visual speech recognition by lip-subword correlation based visual pre-training and cross-modal fusion encoder,” in Proc. ICME 2023, 2023, pp. 2627–2632

  16. [24]

    Gpu-accelerated guided source separation for meeting transcription,

    D. Raj, D. Povey, and S. Khudanpur, “Gpu-accelerated guided source separation for meeting transcription,” arXiv preprint arXiv:2212.05271, 2022

  17. [25]

    The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,

    Z. Wang, S. Wu, H. Chen et al. , “The multimodal information based speech processing (misp) 2022 challenge: Audio-visual di- arization and recognition,” in Proc. ICASSP 2023, 2023, pp. 1–5

  18. [26]

    V oxblink2: A 100k+ speaker recognition corpus and the open-set speaker-identification benchmark,

    Y . Lin, M. Cheng, F. Zhang, Y . Gao, S. Zhang, and M. Li, “V oxblink2: A 100k+ speaker recognition corpus and the open-set speaker-identification benchmark,” arXiv preprint arXiv:2407.11510, 2024

  19. [27]

    V oxceleb2: Deep speaker recognition,

    J. S. Chung, A. Nagrani, and A. Zisserman, “V oxceleb2: Deep speaker recognition,” arXiv preprint arXiv:1806.05622, 2018

  20. [28]

    Kespeech: An open source speech dataset of mandarin and its eight subdialects,

    Z. Tang, D. Wang, Y . Xu, J. Sun, X. Lei, S. Zhao, C. Wen, X. Tan, C. Xie, S. Zhou et al., “Kespeech: An open source speech dataset of mandarin and its eight subdialects,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Bench- marks Track (Roun...

  21. [29]

    3d-speaker: A large-scale multi-device, multi-distance, and multi-dialect cor- pus for speech representation disentanglement,

    S. Zheng, L. Cheng, Y . Chen, H. Wang, and Q. Chen, “3d-speaker: A large-scale multi-device, multi-distance, and multi-dialect cor- pus for speech representation disentanglement,” arXiv preprint arXiv:2306.15354, 2023

  22. [30]

    Ecapa- tdnn: Emphasized channel attention, propagation and ag- gregation in tdnn based speaker verification,

    B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa- tdnn: Emphasized channel attention, propagation and ag- gregation in tdnn based speaker verification,” arXiv preprint arXiv:2005.07143, 2020

  23. [31]

    Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiaoet al., “Wavlm: Large-scale self- supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022

  24. [32]

    The ami meeting corpus,

    W. Kraaij, T. Hain, M. Lincoln, and W. Post, “The ami meeting corpus,” in Proc. International Conference on Methods and Tech- niques in Behavioral Research, 2005, pp. 1–4

  25. [33]

    Lip-reading with densely connected temporal convolutional networks,

    P. Ma, Y . Wang, J. Shen, S. Petridis, and M. Pantic, “Lip-reading with densely connected temporal convolutional networks,” inPro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021, pp. 2857–2866

  26. [34]

    Robust speech recognition via large-scale weak supervision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning . PMLR, 2023, pp. 28 492–28 518

  27. [35]

    Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,

    B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng et al., “Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). ...

  28. [36]

    Sequence-to-sequence neural di- arization with automatic speaker detection and representation,

    M. Cheng, Y . Lin, and M. Li, “Sequence-to-sequence neural di- arization with automatic speaker detection and representation,” arXiv preprint arXiv:2411.13849, 2024

  29. [37]

    Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,

    S. Zhao, Y . Ma, C. Ni, C. Zhang, H. Wang, T. H. Nguyen, K. Zhou, J. Q. Yip, D. Ng, and B. Ma, “Mossformer2: Com- bining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation,” in ICASSP 2024-2024 IEEE International Conference on Acoust...

  30. [38]

    Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,

    Z. Gao, S. Zhang, I. McLoughlin, and Z. Yan, “Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,” arXiv preprint arXiv:2206.08317, 2022

  31. [39]

    Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,

    T. Yoshioka and T. Nakatani, “Generalization of multi-channel linear prediction methods for blind mimo impulse response short- ening,” IEEE Transactions on Audio, Speech, and Language Pro- cessing, vol. 20, no. 10, pp. 2707–2720, 2012

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.