REVIEW 2 major objections 5 minor 1 cited by
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A modular, geometry-agnostic cascade for multi-talker distant ASR cuts time-constrained word error by 63% over the challenge baseline.
desk verdict Solid systems paper with credible ablations; the headline 63% number is real arithmetic on the reported tables, but the undisclosed post-challenge eval-set selection keeps the top-line result from being fully convincing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Modified EEND-VC diarization frontend: an end-to-end diarization network produces local speaker activity per 30-second chunk, then guided source separation (a clustering-based mask estimator) is applied to each detected local speaker before speaker embeddings (ECAPA-TDNN) are extracted and clustered. This segmentation-first order is what lets the system exploit distributed microphones without knowing array geometry. Two supporting mechanisms carry the result: multi-channel speaker counting that pools embeddings from 15-second subchunks within microphone groups found by hierarchical clustering, using normalized maximum eigengap; and the speech-enhancement chain built on microphone selection by envelope variance and the C50 clarity index, a lower bound of 15 microphones, WPE dereverberation, and the spatial-prediction multichannel Wiener filter beamformer. The ASR stage combines Whisper-based and WavLM-based recognizers through ROVER voting with language-model rescoring, which the ablations show consistently lowers error.
What would settle it
Run the full system on a recording with a known four-speaker count but one quiet or far-field speaker: if the pooled-embedding count comes out wrong and tcpWER spikes because that speaker is split or missed, the counting assumption is falsified. The paper's own data specify the expected signature: with oracle counts, macro tcpWER drops from 20.78% to 19.21%, and a single overcount error on one session pushes that session's tcpWER to about 85%.
Extended reading notes
Core claim
The central claim is that the ordering of operations inside diarization matters more than any single model. The paper shows that inserting guided source separation between local end-to-end segmentation and global embedding clustering produces cleaner speaker embeddings, which then make target-speaker voice activity detection, beamformed enhancement, and ASR all work better. It further claims that speaker counting should be decoupled from clustering and done by pooling embeddings across 15-second subchunks and across groups of microphones with similar transfer functions, then reading the count from the normalized maximum eigengap. With that pipeline, the best post-challenge system reaches 20.78% macro tcpWER on the evaluation set, and when the true speaker count is given the same system reaches 19.21%, which locates the remaining bottleneck in speaker counting rather than in ASR.
Load-bearing premise
The load-bearing premise is that speaker embeddings pooled across microphones and 15-second subchunks remain comparable enough for a single eigengap-based count; if microphone transfer functions differ too much within the 120-second similarity window, or if the count is off by one, the whole diarization-then-enhancement chain degrades sharply, as the oracle-count results and Figure 10 show.
Editorial extensions
If this is right
- A geometry-agnostic system can serve as a strong starting point for meeting transcription: on the NOTSOFAR-1 subset the post-challenge system reaches 12.95% session-wise macro tcpWER, close to systems built specifically for that room.
- Correcting speaker counting alone moves macro tcpWER from 20.78% to 19.21% on the evaluation set, so improving the count is a direct route to further gains.
- The segmentation-first pipeline means the system can be retrained or swapped module-by-module: each component is ablated independently, and each ablation points to a specific next step, such as better embeddings for the highest-confusion dataset.
- Complementary ASR backends combined with ROVER reduce insertions and substitutions even when they increase deletions slightly, so system combination is worth the extra compute.
Reading between the lines
- My inference: because the paper's microphone-group pooling assumes embeddings from different channels are comparable, the counting module should be tested in a setup with strongly mismatched microphone placements; that test would probe the weakest assumption directly.
- The reported oracle-count gain suggests a cheap extension: feed the pipeline a manual or externally known attendee count, which the authors note is plausible for meetings, and most of the remaining tcpWER gap to room-tuned systems should close.
- One untested extension follows from the clustering design: replacing eigengap-based counting with a dedicated classifier trained on the pooled embeddings could remove the single largest error source, since Figure 10 shows one overcount pushing a session to about 85% tcpWER.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes NTT's microphone-array-geometry-independent multi-talker distant ASR system for the CHiME-8 DASR task. The system follows a diarization-first cascade: EEND-VC-based pre-diarization with a novel multi-channel speaker-counting stage, TS-VAD refinement, GSS-based speech enhancement with a new microphone-subset-selection rule and SP-MWF beamforming, and four ASR backends (Whisper-L, Whisper-M, NeMo Transducer, WavLM Transducer) combined with LM rescoring and ROVER. The submitted systems (rows 1-3 of Table 2) and post-challenge systems (rows 4-5) are evaluated on CHiME-6, DiPCo, Mixer 6, and NOTSOFAR-1. The headline claim is that the post-challenge system 5 (DIA3 + ASR5) achieves 20.78% macro tcpWER on the evaluation set versus 56.50% for the baseline, a 63% relative improvement, and outperforms the STCON geometry-independent system on NOTSOFAR-1 (12.95% vs. 13.80%). The paper also reports an extensive ablation study in Appendix A covering diarization, speaker counting, speech enhancement, and ASR components.
Significance. If the post-challenge results are accepted as out-of-sample numbers, the paper demonstrates a substantial practical result: a modular, geometry-agnostic cascade of EEND-VC, TS-VAD, GSS/SP-MWF, and foundation-model ASR can match or beat much more specialized meeting-transcription systems on at least one benchmark. The paper is strong in its level of technical detail: the system is described end-to-end, the ablation tables (A.5-A.12) provide per-module evidence, and the reported numbers are internally consistent. In particular, the ablation shows that GSS before embedding extraction reduces DER by about 6 points (Table A.5), the proposed multi-channel group-wise counting improves speaker counting accuracy from 52.7% to 89.6% on the development set (Table A.6), SP-MWF improves over R1-MWF in reference-microphone selection (Table A.7), and ROVER combination gives consistent gains (Table A.10). The main significance is conditional on the evaluation-set selection protocol, because the headline 63% figure is a post-challenge number and the manuscript does not disclose how many post-challenge configurations were tried on the evaluation set before the final one was selected.
major comments (2)
- [§7.3, Tables 2 and 3] The headline 63% relative improvement and the NOTSOFAR-1 outperformance claim rest on the post-challenge rows (systems 4 and 5) of Table 2 and on Table 3, but the manuscript does not state how many post-challenge system configurations were evaluated on the CHiME-8 evaluation set before DIA3 with ASR5 was selected, nor whether the DIA3, ASR5, and post-processing hyperparameters (e.g., θts-vad=0.30 and τmerge=1.5 in §3.5.1, and the ROVER/LM weighting in §5.3) were frozen before evaluation labels were inspected. Because the same evaluation set is used both to select and to report the best configuration, the 20.78% macro tcpWER is not yet established as an out-of-sample property. The paper's acknowledgment of development-set over-tuning in §7.3 is not a substitute for disclosing whether the evaluation set was used for model selection. The authors should report the evaluation-set selection protocol, give results for all post-challenge configurations tried, or explicitly separate the challenge-submitted numbers from the exploratory post-challenge numbers in the abstract and conclusion.
- [§3.2, Eq. (14), Figure 10] The speaker-counting module fixes the number of clusters used by vector clustering, and the paper's own analysis shows that this is the system's main fragility. Figure 10 reports that a single overcount error in one Mixer 6 session drives the session tcpWER to about 85%, and Table 2 shows that replacing estimated counts with oracle counts (DIA4) improves macro tcpWER from 20.78% to 19.21% (system 5 vs. 7). The group-wise counting procedure in Algorithm 1 pools embeddings across microphones within a Tcorr=120 second similarity window under the assumption that enhanced embeddings are comparable across channels (Eq. (14)); if microphone transfer functions differ substantially within that window, the count can fail. Since the paper claims a geometry-independent system for diverse distributed arrays, the manuscript should either quantify the conditions under which this pooling assumption holds, or present the oracle-count results as a separate upper-bound analysis rather than as part of the main claimed system performance.
minor comments (5)
- [Throughout] There are several typos and spacing errors, including 'Multi-T alker' in the highlights, 'V ector clustering' in §3.5.1, 'fined-tuned' in §5.2.2, 'Diairization' in Appendix A.1.1, 'applyingdata selection' in §5.2.4, 'ASR:: Several backends has been developed' in §6, and 'As show in figure 6' in §3.4. Please proofread the manuscript.
- [§2 and §5.3] Section 2 and Figure 9 describe 'four ASR backends', while §5.3 and Table A.10 describe six ASR models (the four main models plus ASR1' and ASR4' variants) and 12 hypotheses for ROVER. Please align the counts to avoid confusing the reader.
- [Figure A.13] The caption/text says 'The three graphs are identical and differ only by their labels'; this is confusing because each subplot uses a different hyperparameter on the color/point legend. Please clarify that the axes are identical but each panel varies one post-processing parameter.
- [§4.3.5] The sentence 'The flooring coefficient 0 ≤ δ ≤ 1 is set to −9 (dB), or δ ≈ 0.355' should read 'set to −9 dB, i.e., δ ≈ 0.355' to be mathematically precise.
- [§5.1] The phrase 'processed with GSS for the Oracle segmentation' is ambiguous; it should specify whether 'Oracle' refers to oracle diarization boundaries or to a particular segmentation protocol.
Circularity Check
No significant circularity: the headline numbers are external benchmark measurements, and the cited self-work is configuration continuity rather than load-bearing derivation.
full rationale
The paper is an empirical challenge system description. Its central claim, a 63% relative macro tcpWER improvement over the CHiME-8 baseline, is computed from measured evaluation-set results in Table 2: baseline 56.50% versus the post-challenge DIA3+ASR5 system at 20.78%. This is an external benchmark comparison, not a quantity derived from the model's own fitted parameters or equations. No equation in the paper reduces to its own input: the speaker-counting module (Algorithm 1, Eq. (14)) is evaluated separately via SCA/SCE in Table A.6, the diarization output is evaluated via DER, and the ASR backends are evaluated via tcpWER against the challenge baseline and the STCON system. The post-challenge configuration is explicitly labelled as such in Table 2, and the paper acknowledges possible development-set over-tuning in Section 7.3: "This may indicate over-tuning of the parameters on the development set." The self-citations (e.g., [19], [36], [37]) are used for model configuration and training-scheme continuity, not to justify the central empirical claim; the relevant components are ablated on the development set against external baselines. The undisclosed selection protocol for the post-challenge configuration is a reproducibility concern, but without evidence that evaluation labels were used for selection it is not a circularity step under the stated rules.
Assumptions & free parameters
free parameters (12)
- Chunk length Tc =
30 s
- Subchunk length Tsubc =
15 s
- Minimum speech duration Tmin =
0.75 s
- Inter-microphone correlation window Tcorr =
120 s
- Microphone grouping threshold thetamic =
0.05
- NME p-value search range =
[1, floor(N/10)]
- TS-VAD post-processing parameters =
theta_tsvad=0.30, tau_offset=0.0 s, tau_merge=1.5 s
- Minimum microphone count Kmin =
15
- Microphone top ratio K1 =
0.65 x M
- LM rescoring weights beta1, beta2 =
not specified
- Curriculum learning CER threshold and loss scale =
30% CER, 1e-3 loss weight
- VoxCeleb contrastive data selection threshold =
not specified
assumptions (6)
- domain assumption Speech and noise are uncorrelated, E[xs xn^H] = 0.
- domain assumption The target source covariance Rs can be treated as rank-1 for beamformer equivalence results.
- domain assumption GSS with cACGMM masks is reliable when guided by EEND local activities.
- domain assumption Enhanced speaker embeddings from different microphones and subchunks are comparable enough for spectral clustering.
- domain assumption Development-set hyperparameter choices transfer to the evaluation set.
- domain assumption Pretrained foundation models transfer to CHiME-8 conditions after fine-tuning on limited in-domain data.
Cite this review
Pith. "Pith review of Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge." pith.science (2026). https://pith.science/paper/7XB2225L
@misc{pith2026250209859,
author = {Pith},
title = {Pith review of: Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge},
year = {2026},
howpublished = {\url{https://pith.science/paper/7XB2225L}},
note = {Machine review of arXiv:2502.09859}
}
read the original abstract
In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker counting, diarization, and ASR. It handles various recording conditions, from diner parties to professional meetings and from two to eight speakers. We perform diarization first, followed by speech enhancement, and then ASR as the challenge baseline. However, we introduced several key refinements. First, we derived a powerful speaker diarization relying on end-to-end speaker diarization with vector clustering (EEND-VC), multi-channel speaker counting using enhanced embeddings from EEND-VC, and target-speaker voice activity detection (TS-VAD). For speech enhancement, we introduced a novel microphone selection rule to better select the most relevant microphones among the distributed microphones and investigated improvements to beamforming. Finally, for ASR, we developed several models exploiting Whisper and WavLM speech foundation models. We present the results we submitted to the challenge and updated results we obtained afterward. Our strongest system achieves a 63% relative macro tcpWER improvement over the baseline and outperforms the challenge best results on the NOTSOFAR-1 meeting evaluation data among geometry-independent systems.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
MOVER: Combining Multiple Meeting Recognition Systems
MOVER is a five-stage combination method that fuses diarization and ASR outputs from multiple meeting recognition systems, improving tcpWER by 9.55% and 8.51% relative on CHiME-8 DASR and NOTSOFAR-1.
Reference graph
Works this paper leans on
-
[1]
Haeb-Umbach, J
R. Haeb-Umbach, J. Heymann, L. Drude, S. Watanabe, M. Delcroix, T. Nakatani, Far-field automatic speech recognition, Proc. IEEE 109 (2) (2021) 124–148
2021
-
[2]
Cornell, M
S. Cornell, M. Wiesner, et al., The CHiME-7 DASR challenge: Distant meeting transcription with multiple devices in diverse scenarios, in: Proc. CHiME, 2023, pp. 1–6
2023
-
[3]
https://www.chimechallenge.org/challenges/chime8/task1/index, accessed 11 December 2024 (2024)
2024
-
[4]
Barker, E
J. Barker, E. Vincent, N. Ma, H. Christensen, P. Green, The PASCAL CHiME speech separation and recognition challenge, Comput. Speech Lang. 27 (3) (2013) 621–633
2013
-
[5]
Vincent, J
E. Vincent, J. Barker, S. Watanabe, J. Le Roux, F. Nesta, M. Matas- soni, The second ‘CHiME’ speech separation and recognition challenge: Datasets, tasks and baselines, in: Proc. ICASSP, 2013, pp. 126–130
2013
-
[6]
Barker, R
J. Barker, R. Marxer, E. Vincent, S. Watanabe, The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines, in: Proc. ASRU, 2015, pp. 504–511
2015
-
[7]
Barker, S
J. Barker, S. Watanabe, E. Vincent, J. Trmal, The fifth ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines, in: Proc. Interspeech, 2018, pp. 1561–1565
2018
-
[8]
Watanabe, M
S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V. Manohar, D. Povey, D. Raj, D. Snyder, A. S. Subrama- nian, J. Trmal, B. B. Yair, C. Boeddeker, Z. Ni, Y. Fujita, S. Horiguchi, N. Kanda, T. Yoshioka, N. Ryant, CHiME-6 challenge: Tackling multi- speaker speech recognition for unsegmented recordings, in: Proc. CHiME, 2020, pp. 1–7
2020
Show all 89 references
-
[9]
Vinnikov, A
A. Vinnikov, A. Ivry, A. Hurvitz, I. Abramovski, S. Koubi, I. Gurvich, S. Peer, X. Xiao, B. M. Elizalde, N. Kanda, X. Wang, S. Shaer, S. Yagev, Y. Asher, S. Sivasankaran, Y. Gong, M. Tang, H. Wang, E. Krupka, NOTSOF AR-1 challenge: New datasets, baseline, and tasks for distant...
2024
-
[10]
D. Raj, P. Denisov, Z. Chen, H. Erdogan, Z. Huang, M. He, S. Watanabe, J. Du, T. Yoshioka, Y. Luo, N. Kanda, J. Li, S. Wisdom, J. R. Hershey, Integration of speech separation, diarization, and recognition for multi- speaker meetings: System description, comparison, and analysi...
2021
-
[11]
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, X. Xiao, J. Li, Continuous speech separation: Dataset and analysis, in: Proc. ICASSP, 2020, pp. 7284–7288
2020
-
[12]
Yoshioka, I
T. Yoshioka, I. Abramovski, C. Aksoylar, Z. Chen, M. David, D. Dimi- triadis, Y. Gong, I. Gurvich, X. Huang, Y. Huang, A. Hurvitz, L. Jiang, S. Koubi, E. Krupka, I. Leichter, C. Liu, P. Parthasarathy, A. Vinnikov, L. Wu, X. Xiao, W. Xiong, H. Wang, Z. Wang, J. Zhang, Y. Zhao, ...
2019
-
[13]
von Neumann, C
T. von Neumann, C. Boeddeker, T. Cord-Landwehr, M. Delcroix, R. Haeb- Umbach, Meeting recognition with continuous speech separation and transcription-supported diarization, in: Proc. ICASSP Workshops, IEEE, 2024, pp. 775–779
2024
-
[14]
Boeddecker, J
C. Boeddecker, J. Heitkaemper, J. Schmalenstroeer, L. Drude, J. Heymann, R. Haeb-Umbach, Front-end processing for the CHiME-5 dinner party sce- nario, in: Proc. CHiME, 2018, pp. 35–40
2018
-
[15]
N. Ito, S. Araki, T. Nakatani, Complex angular central Gaussian mixture model for directional statistics in mask-based microphone array signal pro- cessing, in: Proc. EUSIPCO, 2016, pp. 1153–1157
2016
-
[16]
Medennikov, M
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Ko- renevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, A. Laptev, A. Romanenko, The STC system for the CHiME-6 challenge, in: Proc. CHiME, 2020, pp. 36–41
2020
-
[17]
R. Wan, M. He, J. Du, H. Zhou, S. Niu, H. Chen, Y. Yue, G. Yang, S. Wu, L. Sun, Y. Tu, H. Tang, S. Qian, T. Gao, M. Wang, G. Wan, J. Pan, J. Gao, C.-H. Lee, The USTC-NERCSLIP systems for CHiME-7 challenge, in: Proc. CHiME, 2023, pp. 13–18
2023
-
[18]
Prisyach, Y
T. Prisyach, Y. Khokhlov, M. Korenevsky, A. Mitrofanov, T. Timofeeva, I. Odegov, R. Nasretdinov, I. Lezhenin, D. Miroshnichenko, A. Karelin, M. Mitrofanova, R. Svechnikov, S. Novoselov, A. Romanenko, STCON sys- tem for the CHiME-7 challenge, in: Proc. CHiME, 2023, pp. 87–92
2023
-
[19]
N. Kamo, N. Tawara, K. Matsuura, T. Ashihara, T. Moriya, A. Ogawa, H. Sato, T. Ochiai, A. Ando, R. Ikeshita, T. Kano, M. Delcroix, T. Nakatani, T. Asami, S. Araki, NTT multi-speaker ASR system for the DASR task of CHiME-7 challenge, in: Proc. CHiME, 2023, pp. 45–50. 49
2023
-
[20]
Mitrofanov, T
A. Mitrofanov, T. Prisyach, T. Timofeeva, S. Novoselov, M. Korenevsky, Y. Khokhlov, A. Akulov, A. Anikin, R. Khalili, I. Lezhenin, A. Melnikov, D. Miroshnichenko, N. Mamaev, I. Odegov, O. Rudnitskaya, A. Roma- nenko, STCON system for the CHiME-8 challenge, in: Proc. CHiME, 202...
2024
-
[21]
N. Kamo, N. Tawara, A. Ando, T. Kano, H. Sato, R. Ikeshita, T. Moriya, S. Horiguchi, K. Matsuura, A. Ogawa, A. Plaquet, T. Ashihara, T. Ochiai, M. Mimura, M. Delcroix, T. Nakatani, T. Asami, S. Araki, NTT multi- speaker ASR system for the DASR task of CHiME-8 challenge, in: Pr...
2024
-
[22]
S. Niu, R. Wang, J. Du, G. Yang, Y. Tu, S. Wu, S. Qian, H. Wu, H. Xu, X. Zhang, G. Zhong, X. Yu, J. Chen, M. Wang, D. Cai, T. Gao, G. Wan, F. Ma, J. Pan, J. Gao, The USTC-NERCSLIP systems for the CHiME-8 NOTSOF AR-1 challenge, in: Proc. CHiME, 2024, pp. 31–36
2024
-
[24]
Kinoshita, M
K. Kinoshita, M. Delcroix, N. Tawara, Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds, in: Proc. ICASSP, 2021, pp. 7198–7202
2021
-
[25]
Bredin, A
H. Bredin, A. Laurent, End-to-end speaker segmentation for overlap-aware resegmentation, in: Proc. Interspeech, 2021, pp. 3111–3115
2021
-
[26]
Medennikov, M
I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Ko- renevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, A. Laptev, A. Romanenko, Target-speaker voice activity de- tection: a novel approach for multi-speaker diarization in a dinner party...
2020
-
[27]
M.-K. He, J. Du, Q.-F. Liu, C.-H. Lee, ANSD-MA-MSE: Adaptive neu- ral speaker diarization using memory-aware multi-speaker embedding, IEEE/ACM Trans. Audio Speech Lang. Process. 31 (2023) 1561–1573
2023
-
[28]
Yoshioka, T
T. Yoshioka, T. Nakatani, Generalization of multi-channel linear prediction methods for blind MIMO impulse response shortening, IEEE Trans. Audio Speech Lang. Process. 20 (10) (2012) 2707–2720
2012
-
[29]
Nakatani, T
T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, B.-H. Juang, Speech dereverberation based on variance-normalized delayed linear prediction, IEEE Trans. Audio Speech Lang. Process. 18 (7) (2010) 1717–1731
2010
-
[30]
Benesty, J
J. Benesty, J. Chen, Y. Huang, Noncausal (Frequency-Domain) Optimal Filters, Springer Berlin Heidelberg, Berlin, Heidelberg, 2008, Ch. 6, pp. 115–137. 50
2008
-
[31]
Cornelis, M
B. Cornelis, M. Moonen, J. Wouters, Performance analysis of multichan- nel Wiener filter-based noise reduction in hearing aids under second or- der statistics estimation errors, IEEE Trans. Audio Speech Lang. Process. 19 (5) (2011) 1368–1381
2011
-
[32]
J. G. Fiscus, A post-processing system to yield reduced word error rates: Recognizer Output Voting Error Reduction (ROVER), in: Proc. ASRU, 1997, pp. 347–354
1997
-
[33]
D. Raj, P. Garcia, Z. Huang, S. Watanabe, D. Povey, A. Stolcke, S. Khu- danpur, DOVER-Lap: A method for combining overlap-aware diarization outputs, in: Proc. SLT, 2021, pp. 881–888
2021
-
[34]
Horiguchi, Y
S. Horiguchi, Y. Takashima, P. Garcia, S. Watanabe, Y. Kawaguchi, Multi- channel end-to-end neural diarization with distributed microphones, in: Proc. ICASSP, 2022, pp. 7332–7336
2022
-
[35]
Z. Lu, Y. Wang, Y. Zhang, W. Han, Z. Chen, P. Haghani, Unsupervised data selection via discrete speech representation for ASR, in: Proc. Inter- speech, 2022, pp. 3393–3397
2022
-
[36]
Tawara, M
N. Tawara, M. Delcroix, A. Ando, A. Ogawa, NTT speaker diarization sys- tem for CHiME-7: Multi-domain, multi-microphone end-to-end and vector clustering diarization, in: Proc. ICASSP, 2024, pp. 11281–11285
2024
-
[37]
Tawara, A
N. Tawara, A. Ando, S. Horiguchi, M. Delcroix, Multi-channel speaker counting for EEND-VC-based speaker diarization on multi-domain conver- sation, in: Proc. ICASSP, 2025
2025
-
[38]
von Neumann, C
T. von Neumann, C. Boeddeker, M. Delcroix, R. Haeb-Umbach, MeetEval: A toolkit for computation of word error rates for meeting transcription systems, in: Proc. CHiME, 2023, pp. 27–32
2023
-
[39]
Medennikov, M
I. Medennikov, M. Korenevsky, et al., Target-speaker voice activity de- tection: a novel approach for multi-speaker diarization in a dinner party scenario, in: Proc. Interspeech, 2020, pp. 274–278
2020
-
[40]
Kinoshita, M
K. Kinoshita, M. Delcroix, N. Tawara, Advances in integration of end-to- end neural and clustering-based diarization for real conversational speech, in: Proc. Interspeech, 2021, pp. 3565–3569
2021
-
[41]
Desplanques, J
B. Desplanques, J. Thienpondt, K. Demuynck, ECAPA-TDNN: Empha- sized channel attention, propagation and aggregation in TDNN based speaker verification, in: Proc. Interspeech, 2020, pp. 3830–3834
2020
-
[42]
J. H. Ward Jr., Hierarchical grouping to optimize an objective function, J. Am. Stat. Assoc. 58 (301) (1963) 236–244
1963
-
[43]
J. Shi, J. Malik, Normalized cuts and image segmentation, IEEE Trans. Pattern Anal. Mach. Intell. 22 (8) (2000) 888–905. 51
2000
-
[44]
A. Y. Ng, M. I. Jordan, Y. Weiss, On spectral clustering: Analysis and an algorithm, in: Proc. NIPS, 2001, pp. 849–856
2001
-
[45]
Bredin, pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe, in: Proc
H. Bredin, pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe, in: Proc. Interspeech, 2023, pp. 1983–1987
2023
-
[46]
T. J. Park, H. Huang, A. Juki´ c, K. Dhawan, K. C. Puvvada, N. R. Koluguri, N. Karpov, A. Laptev, J. Balam, B. Ginsburg, The CHiME-7 challenge: System description and performance of NeMo team’s DASR system, in: Proc. CHiME, 2023, pp. 57–62
2023
-
[47]
T. J. Park, K. J. Han, M. Kumar, S. Narayanan, Auto-tuning spectral clus- tering for speaker diarization using normalized maximum eigengap, IEEE Signal Processing Letters 27 (2020) 381–385
2020
-
[48]
Ishiguro, T
K. Ishiguro, T. Yamada, S. Araki, T. Nakatani, H. Sawada, Probabilis- tic speaker diarization with bag-of-words representations of speaker angle information, IEEE Trans. Audio Speech Lang. Process. 20 (2) (2011) 447– 460
2011
-
[49]
Boeddeker, A
C. Boeddeker, A. S. Subramanian, G. Wichern, R. Haeb-Umbach, J. Le Roux, TS-SEP: Joint diarization and separation conditioned on esti- mated speaker embeddings, IEEE/ACM Trans. Audio, Speech Lang. Pro- cess. 32 (2024) 1185—-1197
2024
-
[50]
Yoshioka, X
T. Yoshioka, X. Wang, D. Wang, M. Tang, Z. Zhu, Z. Chen, N. Kanda, Vararray: Array-geometry-agnostic continuous speech separation, in: Proc. ICASSP, 2022, pp. 6027–6031
2022
-
[51]
D. Wang, Z. Chen, T. Yoshioka, Neural speech separation using spa- tially distributed microphones, in: Interspeech 2020, 2020, pp. 339–343. doi:10.21437/Interspeech.2020-1089
2020 doi
-
[52]
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, X. Yu, F. Wei, WavLM: Large-scale self-supervised pre-training for full stack speech processing, IEEE J. Sel. Top. Signal Process...
2022
-
[53]
https://huggingface.co/speechbrain/spkrec-ecapa-voxceleb, accessed 11 December 2024
2024
-
[54]
Nagrani, J
A. Nagrani, J. S. Chung, W. Xie, A. Zisserman, VoxCeleb: Large-scale speaker verification in the wild, Comput. Speech Lang. 60 (2020) 101027
2020
-
[55]
G. Yang, M. He, S. Niu, R. Wang, Y. Yue, S. Qian, S. Wu, J. Du, C.-H. Lee, Neural speaker diarization using memory-aware multi-speaker embed- ding with sequence-to-sequence architecture, in: Proc. ICASSP, 2024, pp. 11626–11630. 52
2024
-
[56]
Cornell, T
S. Cornell, T. Park, S. Huang, C. Boeddeker, X. Chang, M. Maciejewski, M. Wiesner, P. Garcia, S. Watanabe, The CHiME-8 DASR challenge for generalizable and array agnostic distant automatic speech recognition and diarization, in: Proc. CHiME, 2024, pp. 1–6
2024
-
[57]
Yamashita, S
N. Yamashita, S. Horiguchi, T. Homma, Improving the naturalness of sim- ulated conversations for end-to-end neural diarization, in: Proc. Odyssey, 2022, pp. 133–140
2022
-
[58]
Panayotov, G
V. Panayotov, G. Chen, D. Povey, S. Khudanpur, LibriSpeech: an ASR corpus based on public domain audio books, in: Proc. ICASSP, 2015, pp. 5206–5210
2015
-
[59]
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, S. Khudanpur, A study on data augmentation of reverberant speech for robust speech recognition, in: Proc. ICASSP, 2017, pp. 5220–5224
2017
-
[60]
Snyder, G
D. Snyder, G. Chen, D. Povey, MUSAN: A music, speech, and noise corpus, arXiv:1510.08484 (2015)
2015 arXiv
-
[61]
Fujita, N
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, S. Watanabe, End-to- end neural speaker diarization with permutation-free objectives, in: Proc. Interspeech, 2019, pp. 4300–4304
2019
-
[62]
M. Wolf, C. Nadeu, Channel selection measures for multi-microphone speech recognition, Speech Commun. 57 (2014) 170–180
2014
-
[63]
Souden, J
M. Souden, J. Benesty, S. Affes, On optimal frequency-domain multichan- nel linear filtering for noise reduction, IEEE Trans. Audio Speech Lang. Process. 18 (2) (2009) 260–276
2009
-
[64]
T. C. Lawin-Ore, S. Doclo, Reference microphone selection for MWF-based noise reduction using distributed microphone arrays, in: Speech Commu- nication; 10. ITG Symposium, 2012, pp. 1–4
2012
-
[65]
Erdogan, J
H. Erdogan, J. R. Hershey, S. Watanabe, M. I. Mandel, J. L. Roux, Improved MVDR beamforming using single-channel mask prediction net- works, in: Proc. Interspeech, 2016, pp. 1981–1985
2016
-
[66]
Warsitz, R
E. Warsitz, R. Haeb-Umbach, Blind acoustic beamforming based on gener- alized eigenvalue decomposition, IEEE Trans. Audio Speech Lang. Process. 15 (5) (2007) 1529–1539
2007
-
[67]
P. A. Naylor, N. D. Gaubitch, Speech dereverberation, Springer Science & Business Media, 2010
2010
-
[68]
Lavechin, M
M. Lavechin, M. M´ etais, H. Titeux, A. Boissonnet, J. Copet, M. Rivi` ere, E. Bergelson, A. Cristia, E. Dupoux, H. Bredin, Brouhaha: Multi-task training for voice activity detection, speech-to-noise ratio, and C50 room acoustics estimation, in: Proc. ASRU, 2023, pp. 1–7. 53
2023
-
[69]
Gannot, E
S. Gannot, E. Vincent, S. Markovich-Golan, A. Ozerov, A consolidated perspective on multimicrophone speech enhancement and source separa- tion, IEEE/ACM Transactions on Audio, Speech, and Language Processing 25 (4) (2017) 692–730
2017
-
[70]
Doclo, M
S. Doclo, M. Moonen, GSVD-based optimal filtering for single and multimi- crophone speech enhancement, IEEE Trans. Signal Process. 50 (9) (2002) 2230–2244
2002
-
[71]
Spriet, M
A. Spriet, M. Moonen, J. Wouters, Spatially pre-processed speech distor- tion weighted multi-channel Wiener filtering for noise reduction, Signal Process. 84 (12) (2004) 2367–2387
2004
-
[72]
Doclo, A
S. Doclo, A. Spriet, J. Wouters, M. Moonen, Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction, Speech Commun. 49 (7) (2007) 636–656
2007
-
[73]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in NIPS, 2017, pp. 5998–6008
2017
-
[74]
Graves, Sequence transduction with recurrent neural networks, in: ICML Workshop on Representation Learning, 2012
A. Graves, Sequence transduction with recurrent neural networks, in: ICML Workshop on Representation Learning, 2012
2012
-
[75]
Radford, J
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large-scale weak supervision, in: Proc. ICML, 2023, pp. 28492–28518
2023
-
[76]
Sennrich, B
R. Sennrich, B. Haddow, A. Birch, Neural machine translation of rare words with subword units, in: Proc. ACL, 2016, pp. 1715–1725
2016
-
[77]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language models are unsupervised multitask learners, OpenAI blog 1 (8) (2019) 9
2019
-
[78]
Kuchaiev, J
O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V. Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, J. M. Cohen, NeMo: A toolkit for building AI applications using neural modules, arXiv:1909.09577 (2019)
2019 arXiv
-
[79]
Gulati, J
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, R. Pang, Conformer: Convolution-augmented transformer for speech recognition, in: Proc. Interspeech, 2020, pp. 5036– 5040
2020
-
[80]
Rekesh, N
D. Rekesh, N. R. Koluguri, S. Kriman, S. Majumdar, V. Noroozi, H. Huang, O. Hrinchuk, K. Puvvada, A. Kumar, J. Balam, B. Ginsburg, Fast Con- former with linearly scalable attention for efficient speech recognition, in: Proc. ASRU, 2023, pp. 1–8. 54
2023
-
[81]
K. Kim, F. Wu, Y. Peng, J. Pan, P. Sridhar, K. J. Han, S. Watanabe, E-Branchformer: Branchformer with enhanced merging for speech recogni- tion, in: Proc. SLT, 2023, pp. 84–91
2023
-
[82]
Ogawa, N
A. Ogawa, N. Tawara, M. Delcroix, S. Araki, Lattice rescoring based on large ensemble of complementary neural language models, in: Proc. ICASSP, 2022, pp. 6517–6520
2022
-
[83]
Ogawa, N
A. Ogawa, N. Kamo, K. Matsuura, T. Ashihara, T. Moriya, T. Kano, N. Tawara, M. Delcroix, Applying LLMs for rescoring N-best ASR hy- potheses of casual conversations: Effects of domain adaptation and context carry-over, arXiv:2406.18972 (2024)
2024 arXiv
-
[84]
Povey, A
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, K. Vesely, The Kaldi speech recognition toolkit, in: Proc. ASRU, 2011
2011
-
[85]
Andrusenko, R
A. Andrusenko, R. Nasretdinov, A. Romanenko, Uconv-conformer: High reduction of input sequence length for end-to-end speech recognition, in: Proc. ICASSP, 2023, pp. 1–5
2023
-
[86]
Watanabe, T
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. En- rique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, T. Ochiai, ESPnet: End-to-end speech processing toolkit, in: Proc. Inter- speech, 2018, pp. 2207–2211
2018
-
[87]
Z. Yao, L. Guo, X. Yang, W. Kang, F. Kuang, Y. Yang, Z. Jin, L. Lin, D. Povey, Zipformer: A faster and better encoder for automatic speech recognition, in: Proc. ICLR, 2024
2024
-
[88]
Hirano, M
Y. Hirano, M. Nguyen, K. Azuma, J. M. Saragih, S. Sakti, The NAIST system for the CHiME-8 NOTSOF AR-1 task, in: Proc. CHiME, 2024, pp. 59–63
2024
-
[89]
Huang, Y
K. Huang, Y. Li, Z. Wang, H. Wang, W. Rao, Z. Sun, Z. Tang, S. Huang, Y. Wang, T. Yu, L. Xie, S. dong Shang, The NPU-TEA system for the CHiME-8 NOTSOF AR-1 challenge, in: Proc. CHiME, 2024, pp. 45–48
2024
-
[90]
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, Q. V. Le, SpecAugment: A simple data augmentation method for auto- matic speech recognition, in: Proc. Interspeech, 2019, pp. 2613–2617. 55
2019
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.