Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A modular, geometry-agnostic cascade for multi-talker distant ASR cuts time-constrained word error by 63% over the challenge baseline.

desk verdict Solid systems paper with credible ablations; the headline 63% number is real arithmetic on the reported tables, but the undisclosed post-challenge eval-set selection keeps the top-line result from being fully convincing. read the letter →

arxiv 2502.09859 v2 pith:7XB2225L submitted 2025-02-14 eess.AS eess.SP

classification eess.ASeess.SP
keywords multi-talkerdistantASRspeakerdiarizationcountingguidedsourceseparationbeamformingspeechfoundationmodelssystemcombinationCHiME-8DASR
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a system paper for multi-talker distant automatic speech recognition: transcribing a meeting or dinner-party recording while also deciding who spoke when. The authors claim that a diarization-first cascade — speaker counting, diarization, speech enhancement, then ASR — can be made robust across unknown microphone array geometries without sacrificing accuracy. On the CHiME-8 DASR evaluation set, their strongest configuration achieves a 63% relative improvement in macro tcpWER over the challenge baseline (20.78% versus 56.50%), and it reports the best result among geometry-independent systems on the NOTSOFAR-1 meeting subset. The broader point is that modular enhancement and recognized speech foundation models can close most of the gap to systems specialized for a single room geometry.

What carries the argument

Modified EEND-VC diarization frontend: an end-to-end diarization network produces local speaker activity per 30-second chunk, then guided source separation (a clustering-based mask estimator) is applied to each detected local speaker before speaker embeddings (ECAPA-TDNN) are extracted and clustered. This segmentation-first order is what lets the system exploit distributed microphones without knowing array geometry. Two supporting mechanisms carry the result: multi-channel speaker counting that pools embeddings from 15-second subchunks within microphone groups found by hierarchical clustering, using normalized maximum eigengap; and the speech-enhancement chain built on microphone selection by envelope variance and the C50 clarity index, a lower bound of 15 microphones, WPE dereverberation, and the spatial-prediction multichannel Wiener filter beamformer. The ASR stage combines Whisper-based and WavLM-based recognizers through ROVER voting with language-model rescoring, which the ablations show consistently lowers error.

What would settle it

Run the full system on a recording with a known four-speaker count but one quiet or far-field speaker: if the pooled-embedding count comes out wrong and tcpWER spikes because that speaker is split or missed, the counting assumption is falsified. The paper's own data specify the expected signature: with oracle counts, macro tcpWER drops from 20.78% to 19.21%, and a single overcount error on one session pushes that session's tcpWER to about 85%.

Watch

Extended reading notes

Core claim

The central claim is that the ordering of operations inside diarization matters more than any single model. The paper shows that inserting guided source separation between local end-to-end segmentation and global embedding clustering produces cleaner speaker embeddings, which then make target-speaker voice activity detection, beamformed enhancement, and ASR all work better. It further claims that speaker counting should be decoupled from clustering and done by pooling embeddings across 15-second subchunks and across groups of microphones with similar transfer functions, then reading the count from the normalized maximum eigengap. With that pipeline, the best post-challenge system reaches 20.78% macro tcpWER on the evaluation set, and when the true speaker count is given the same system reaches 19.21%, which locates the remaining bottleneck in speaker counting rather than in ASR.

Load-bearing premise

The load-bearing premise is that speaker embeddings pooled across microphones and 15-second subchunks remain comparable enough for a single eigengap-based count; if microphone transfer functions differ too much within the 120-second similarity window, or if the count is off by one, the whole diarization-then-enhancement chain degrades sharply, as the oracle-count results and Figure 10 show.

Editorial extensions

If this is right

  • A geometry-agnostic system can serve as a strong starting point for meeting transcription: on the NOTSOFAR-1 subset the post-challenge system reaches 12.95% session-wise macro tcpWER, close to systems built specifically for that room.
  • Correcting speaker counting alone moves macro tcpWER from 20.78% to 19.21% on the evaluation set, so improving the count is a direct route to further gains.
  • The segmentation-first pipeline means the system can be retrained or swapped module-by-module: each component is ablated independently, and each ablation points to a specific next step, such as better embeddings for the highest-confusion dataset.
  • Complementary ASR backends combined with ROVER reduce insertions and substitutions even when they increase deletions slightly, so system combination is worth the extra compute.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: because the paper's microphone-group pooling assumes embeddings from different channels are comparable, the counting module should be tested in a setup with strongly mismatched microphone placements; that test would probe the weakest assumption directly.
  • The reported oracle-count gain suggests a cheap extension: feed the pipeline a manual or externally known attendee count, which the authors note is plausible for meetings, and most of the remaining tcpWER gap to room-tuned systems should close.
  • One untested extension follows from the clustering design: replacing eigengap-based counting with a dedicated classifier trained on the pooled embeddings could remove the single largest error source, since Figure 10 shows one overcount pushing a session to about 85% tcpWER.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper describes NTT's microphone-array-geometry-independent multi-talker distant ASR system for the CHiME-8 DASR task. The system follows a diarization-first cascade: EEND-VC-based pre-diarization with a novel multi-channel speaker-counting stage, TS-VAD refinement, GSS-based speech enhancement with a new microphone-subset-selection rule and SP-MWF beamforming, and four ASR backends (Whisper-L, Whisper-M, NeMo Transducer, WavLM Transducer) combined with LM rescoring and ROVER. The submitted systems (rows 1-3 of Table 2) and post-challenge systems (rows 4-5) are evaluated on CHiME-6, DiPCo, Mixer 6, and NOTSOFAR-1. The headline claim is that the post-challenge system 5 (DIA3 + ASR5) achieves 20.78% macro tcpWER on the evaluation set versus 56.50% for the baseline, a 63% relative improvement, and outperforms the STCON geometry-independent system on NOTSOFAR-1 (12.95% vs. 13.80%). The paper also reports an extensive ablation study in Appendix A covering diarization, speaker counting, speech enhancement, and ASR components.

Significance. If the post-challenge results are accepted as out-of-sample numbers, the paper demonstrates a substantial practical result: a modular, geometry-agnostic cascade of EEND-VC, TS-VAD, GSS/SP-MWF, and foundation-model ASR can match or beat much more specialized meeting-transcription systems on at least one benchmark. The paper is strong in its level of technical detail: the system is described end-to-end, the ablation tables (A.5-A.12) provide per-module evidence, and the reported numbers are internally consistent. In particular, the ablation shows that GSS before embedding extraction reduces DER by about 6 points (Table A.5), the proposed multi-channel group-wise counting improves speaker counting accuracy from 52.7% to 89.6% on the development set (Table A.6), SP-MWF improves over R1-MWF in reference-microphone selection (Table A.7), and ROVER combination gives consistent gains (Table A.10). The main significance is conditional on the evaluation-set selection protocol, because the headline 63% figure is a post-challenge number and the manuscript does not disclose how many post-challenge configurations were tried on the evaluation set before the final one was selected.

major comments (2)
  1. [§7.3, Tables 2 and 3] The headline 63% relative improvement and the NOTSOFAR-1 outperformance claim rest on the post-challenge rows (systems 4 and 5) of Table 2 and on Table 3, but the manuscript does not state how many post-challenge system configurations were evaluated on the CHiME-8 evaluation set before DIA3 with ASR5 was selected, nor whether the DIA3, ASR5, and post-processing hyperparameters (e.g., θts-vad=0.30 and τmerge=1.5 in §3.5.1, and the ROVER/LM weighting in §5.3) were frozen before evaluation labels were inspected. Because the same evaluation set is used both to select and to report the best configuration, the 20.78% macro tcpWER is not yet established as an out-of-sample property. The paper's acknowledgment of development-set over-tuning in §7.3 is not a substitute for disclosing whether the evaluation set was used for model selection. The authors should report the evaluation-set selection protocol, give results for all post-challenge configurations tried, or explicitly separate the challenge-submitted numbers from the exploratory post-challenge numbers in the abstract and conclusion.
  2. [§3.2, Eq. (14), Figure 10] The speaker-counting module fixes the number of clusters used by vector clustering, and the paper's own analysis shows that this is the system's main fragility. Figure 10 reports that a single overcount error in one Mixer 6 session drives the session tcpWER to about 85%, and Table 2 shows that replacing estimated counts with oracle counts (DIA4) improves macro tcpWER from 20.78% to 19.21% (system 5 vs. 7). The group-wise counting procedure in Algorithm 1 pools embeddings across microphones within a Tcorr=120 second similarity window under the assumption that enhanced embeddings are comparable across channels (Eq. (14)); if microphone transfer functions differ substantially within that window, the count can fail. Since the paper claims a geometry-independent system for diverse distributed arrays, the manuscript should either quantify the conditions under which this pooling assumption holds, or present the oracle-count results as a separate upper-bound analysis rather than as part of the main claimed system performance.
minor comments (5)
  1. [Throughout] There are several typos and spacing errors, including 'Multi-T alker' in the highlights, 'V ector clustering' in §3.5.1, 'fined-tuned' in §5.2.2, 'Diairization' in Appendix A.1.1, 'applyingdata selection' in §5.2.4, 'ASR:: Several backends has been developed' in §6, and 'As show in figure 6' in §3.4. Please proofread the manuscript.
  2. [§2 and §5.3] Section 2 and Figure 9 describe 'four ASR backends', while §5.3 and Table A.10 describe six ASR models (the four main models plus ASR1' and ASR4' variants) and 12 hypotheses for ROVER. Please align the counts to avoid confusing the reader.
  3. [Figure A.13] The caption/text says 'The three graphs are identical and differ only by their labels'; this is confusing because each subplot uses a different hyperparameter on the color/point legend. Please clarify that the axes are identical but each panel varies one post-processing parameter.
  4. [§4.3.5] The sentence 'The flooring coefficient 0 ≤ δ ≤ 1 is set to −9 (dB), or δ ≈ 0.355' should read 'set to −9 dB, i.e., δ ≈ 0.355' to be mathematically precise.
  5. [§5.1] The phrase 'processed with GSS for the Oracle segmentation' is ambiguous; it should specify whether 'Oracle' refers to oracle diarization boundaries or to a particular segmentation protocol.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline numbers are external benchmark measurements, and the cited self-work is configuration continuity rather than load-bearing derivation.

full rationale

The paper is an empirical challenge system description. Its central claim, a 63% relative macro tcpWER improvement over the CHiME-8 baseline, is computed from measured evaluation-set results in Table 2: baseline 56.50% versus the post-challenge DIA3+ASR5 system at 20.78%. This is an external benchmark comparison, not a quantity derived from the model's own fitted parameters or equations. No equation in the paper reduces to its own input: the speaker-counting module (Algorithm 1, Eq. (14)) is evaluated separately via SCA/SCE in Table A.6, the diarization output is evaluated via DER, and the ASR backends are evaluated via tcpWER against the challenge baseline and the STCON system. The post-challenge configuration is explicitly labelled as such in Table 2, and the paper acknowledges possible development-set over-tuning in Section 7.3: "This may indicate over-tuning of the parameters on the development set." The self-citations (e.g., [19], [36], [37]) are used for model configuration and training-scheme continuity, not to justify the central empirical claim; the relevant components are ablated on the development set against external baselines. The undisclosed selection protocol for the post-challenge configuration is a reproducibility concern, but without evidence that evaluation labels were used for selection it is not a circularity step under the stated rules.

Assumptions & free parameters 12 free parameters · 6 assumptions · 0 invented entities

The paper is an empirical systems paper, so the ledger is dominated by development-set-fitted hyperparameters and standard domain assumptions about audio and speaker embeddings. No new physical or learned entities are postulated. The most consequential free parameters are the speaker counting and microphone selection thresholds because they directly control the major failure mode identified by the authors.

free parameters (12)
  • Chunk length Tc = 30 s
    Chunk size for EEND segmentation and vector clustering; determined using the development set (Section 3.5.1).
  • Subchunk length Tsubc = 15 s
    Subchunk size for speaker counting embeddings; half the chunk length, tuned on the development set (Section 3.2).
  • Minimum speech duration Tmin = 0.75 s
    Embeddings from subchunks with less total speech are discarded (Algorithm 1).
  • Inter-microphone correlation window Tcorr = 120 s
    Window used to compute microphone similarity in Eq. (12); set based on preliminary experiments (Section 3.2).
  • Microphone grouping threshold thetamic = 0.05
    AHC stopping threshold for microphone grouping (Section 3.2).
  • NME p-value search range = [1, floor(N/10)]
    Reduced from the baseline range [1, floor(N/4)] to reduce speaker overcounting; tuned on the development set (Section 3.5.1).
  • TS-VAD post-processing parameters = theta_tsvad=0.30, tau_offset=0.0 s, tau_merge=1.5 s
    Grid-searched on the development set to optimize tcpWER rather than DER (Appendix A.1.4).
  • Minimum microphone count Kmin = 15
    Lower bound on selected microphones in the proposed subset selection; empirically determined on the development set (Section 4.2.2).
  • Microphone top ratio K1 = 0.65 x M
    Fraction of microphones ranked by EV and C50 whose intersection is selected; empirical (Section 4.2.2).
  • LM rescoring weights beta1, beta2 = not specified
    Weights in Eq. (37); tuned on the development set (Section 5.3).
  • Curriculum learning CER threshold and loss scale = 30% CER, 1e-3 loss weight
    Whisper fine-tuning uses self-generated references when CER exceeds 30% and down-weights those losses; based on preliminary experiments (Section 5.2.1).
  • VoxCeleb contrastive data selection threshold = not specified
    Threshold reduces VoxCeleb to a quarter of its size; the value is not reported (Section 5.2.4).
assumptions (6)
  • domain assumption Speech and noise are uncorrelated, E[xs xn^H] = 0.
    Used to derive the SDW-MWF solution in Eq. (25) from the optimization in Eq. (24). Standard but not exactly true in reverberant recordings.
  • domain assumption The target source covariance Rs can be treated as rank-1 for beamformer equivalence results.
    Equation (28) uses Rs = Ps z z^H to show SP-MWF equals R1-MWF under rank-1; the paper notes reverberation can make Rs higher rank but uses the beamformers anyway (Section 4.3.3).
  • domain assumption GSS with cACGMM masks is reliable when guided by EEND local activities.
    The speech enhancement and speaker embedding extraction depend on GSS masks (Section 4.1); the assumption is partially validated by ablations but not proven for all conditions.
  • domain assumption Enhanced speaker embeddings from different microphones and subchunks are comparable enough for spectral clustering.
    Algorithm 1 pools embeddings across microphone groups and subchunks before NME counting; if channel variability breaks comparability, the speaker count estimates fail.
  • domain assumption Development-set hyperparameter choices transfer to the evaluation set.
    Many parameters are tuned on the CHiME-8 development set; the paper itself observes higher SCE on evaluation and suggests possible over-tuning (Section 7.3).
  • domain assumption Pretrained foundation models transfer to CHiME-8 conditions after fine-tuning on limited in-domain data.
    Whisper, WavLM, ECAPA-TDNN, and VoxCeleb-based models are used with only 70 hours of CHiME-8 training data; the paper assumes this is sufficient (Sections 5.1 and 5.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge." pith.science (2026). https://pith.science/paper/7XB2225L

@misc{pith2026250209859,
  author       = {Pith},
  title        = {Pith review of: Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7XB2225L}},
  note         = {Machine review of arXiv:2502.09859}
}
read the original abstract

In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker counting, diarization, and ASR. It handles various recording conditions, from diner parties to professional meetings and from two to eight speakers. We perform diarization first, followed by speech enhancement, and then ASR as the challenge baseline. However, we introduced several key refinements. First, we derived a powerful speaker diarization relying on end-to-end speaker diarization with vector clustering (EEND-VC), multi-channel speaker counting using enhanced embeddings from EEND-VC, and target-speaker voice activity detection (TS-VAD). For speech enhancement, we introduced a novel microphone selection rule to better select the most relevant microphones among the distributed microphones and investigated improvements to beamforming. Finally, for ASR, we developed several models exploiting Whisper and WavLM speech foundation models. We present the results we submitted to the challenge and updated results we obtained afterward. Our strongest system achieves a 63% relative macro tcpWER improvement over the baseline and outperforms the challenge best results on the NOTSOFAR-1 meeting evaluation data among geometry-independent systems.

Figures

Figures reproduced from arXiv: 2502.09859 by the authors.

Figure 1
Figure 1. Overview of proposed multi-talker DASR system. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Schematic diagram of proposed speaker diarization system. Double-lined arrows [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Schematic diagram of EEND-VC applied to m-th channel audio. For simplicity, this shows an example with three distinct speakers (represented by red, green, and blue waveforms in the bottom visualization). The first chunk ym,1 contains three active speakers in total (red, green, and blue speakers), while the i-th chunk ym,i only contains two active speakers (blue and red) [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Schematic diagram of multi-channel speaker counting module. [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: Outline of i-vector-based Seq2Seq TS-VAD. [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Postprocessing to obtain utterance boundaries from speaker activity. The blue [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: Processing flow of proposed SE system, which follows the same structure as the [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Processing flow of SE system used for speaker diarization. Unlike the SE system, [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Proposed ASR backend. and combines their output with system combination. The details of the ASR models, the training data used for their training, the language model (LM) used for rescoring, and the system combination approach are described below. 5.1. Training data fo…
Figure 10
Figure 10. Figure 10: Impact of speaker count errors on tcpWER of the different sessions of the evaluation [PITH_FULL_IMAGE:figures/full_fig_p037_10.png]
Figure 11
Figure 11. Figure 11: Distribution of ASR and diarization errors for different datasets of the evaluation [PITH_FULL_IMAGE:figures/full_fig_p039_11.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOVER: Combining Multiple Meeting Recognition Systems

    eess.AS 2025-08 unverdicted novelty 5.0 of 10

    MOVER is a five-stage combination method that fuses diarization and ASR outputs from multiple meeting recognition systems, improving tcpWER by 9.55% and 8.51% relative on CHiME-8 DASR and NOTSOFAR-1.

Reference graph

Works this paper leans on

89 extracted references · 77 canonical work pages · cited by 1 Pith paper

  1. [1]

    Haeb-Umbach, J

    R. Haeb-Umbach, J. Heymann, L. Drude, S. Watanabe, M. Delcroix, T. Nakatani, Far-field automatic speech recognition, Proc. IEEE 109 (2) (2021) 124–148

  2. [2]

    Cornell, M

    S. Cornell, M. Wiesner, et al., The CHiME-7 DASR challenge: Distant meeting transcription with multiple devices in diverse scenarios, in: Proc. CHiME, 2023, pp. 1–6

  3. [3]

    https://www.chimechallenge.org/challenges/chime8/task1/index, accessed 11 December 2024 (2024)

  4. [4]

    Barker, E

    J. Barker, E. Vincent, N. Ma, H. Christensen, P. Green, The PASCAL CHiME speech separation and recognition challenge, Comput. Speech Lang. 27 (3) (2013) 621–633

  5. [5]

    Vincent, J

    E. Vincent, J. Barker, S. Watanabe, J. Le Roux, F. Nesta, M. Matas- soni, The second ‘CHiME’ speech separation and recognition challenge: Datasets, tasks and baselines, in: Proc. ICASSP, 2013, pp. 126–130

  6. [6]

    Barker, R

    J. Barker, R. Marxer, E. Vincent, S. Watanabe, The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines, in: Proc. ASRU, 2015, pp. 504–511

  7. [7]

    Barker, S

    J. Barker, S. Watanabe, E. Vincent, J. Trmal, The fifth ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines, in: Proc. Interspeech, 2018, pp. 1561–1565

  8. [8]

    Watanabe, M

    S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V. Manohar, D. Povey, D. Raj, D. Snyder, A. S. Subrama- nian, J. Trmal, B. B. Yair, C. Boeddeker, Z. Ni, Y. Fujita, S. Horiguchi, N. Kanda, T. Yoshioka, N. Ryant, CHiME-6 challenge: Tackling multi- speaker speech recognition for unsegmented recordings, in: Proc. CHiME, 2020, pp. 1–7

Show all 89 references
  1. [9]

    Vinnikov, A

    A. Vinnikov, A. Ivry, A. Hurvitz, I. Abramovski, S. Koubi, I. Gurvich, S. Peer, X. Xiao, B. M. Elizalde, N. Kanda, X. Wang, S. Shaer, S. Yagev, Y. Asher, S. Sivasankaran, Y. Gong, M. Tang, H. Wang, E. Krupka, NOTSOF AR-1 challenge: New datasets, baseline, and tasks for distant...

  2. [10]

    D. Raj, P. Denisov, Z. Chen, H. Erdogan, Z. Huang, M. He, S. Watanabe, J. Du, T. Yoshioka, Y. Luo, N. Kanda, J. Li, S. Wisdom, J. R. Hershey, Integration of speech separation, diarization, and recognition for multi- speaker meetings: System description, comparison, and analysi...

  3. [11]

    Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, X. Xiao, J. Li, Continuous speech separation: Dataset and analysis, in: Proc. ICASSP, 2020, pp. 7284–7288

  4. [12]

    Yoshioka, I

    T. Yoshioka, I. Abramovski, C. Aksoylar, Z. Chen, M. David, D. Dimi- triadis, Y. Gong, I. Gurvich, X. Huang, Y. Huang, A. Hurvitz, L. Jiang, S. Koubi, E. Krupka, I. Leichter, C. Liu, P. Parthasarathy, A. Vinnikov, L. Wu, X. Xiao, W. Xiong, H. Wang, Z. Wang, J. Zhang, Y. Zhao, ...

  5. [13]

    von Neumann, C

    T. von Neumann, C. Boeddeker, T. Cord-Landwehr, M. Delcroix, R. Haeb- Umbach, Meeting recognition with continuous speech separation and transcription-supported diarization, in: Proc. ICASSP Workshops, IEEE, 2024, pp. 775–779

  6. [14]

    Boeddecker, J

    C. Boeddecker, J. Heitkaemper, J. Schmalenstroeer, L. Drude, J. Heymann, R. Haeb-Umbach, Front-end processing for the CHiME-5 dinner party sce- nario, in: Proc. CHiME, 2018, pp. 35–40

  7. [15]

    N. Ito, S. Araki, T. Nakatani, Complex angular central Gaussian mixture model for directional statistics in mask-based microphone array signal pro- cessing, in: Proc. EUSIPCO, 2016, pp. 1153–1157

  8. [16]

    Medennikov, M

    I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Ko- renevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, A. Laptev, A. Romanenko, The STC system for the CHiME-6 challenge, in: Proc. CHiME, 2020, pp. 36–41

  9. [17]

    R. Wan, M. He, J. Du, H. Zhou, S. Niu, H. Chen, Y. Yue, G. Yang, S. Wu, L. Sun, Y. Tu, H. Tang, S. Qian, T. Gao, M. Wang, G. Wan, J. Pan, J. Gao, C.-H. Lee, The USTC-NERCSLIP systems for CHiME-7 challenge, in: Proc. CHiME, 2023, pp. 13–18

  10. [18]

    Prisyach, Y

    T. Prisyach, Y. Khokhlov, M. Korenevsky, A. Mitrofanov, T. Timofeeva, I. Odegov, R. Nasretdinov, I. Lezhenin, D. Miroshnichenko, A. Karelin, M. Mitrofanova, R. Svechnikov, S. Novoselov, A. Romanenko, STCON sys- tem for the CHiME-7 challenge, in: Proc. CHiME, 2023, pp. 87–92

  11. [19]

    N. Kamo, N. Tawara, K. Matsuura, T. Ashihara, T. Moriya, A. Ogawa, H. Sato, T. Ochiai, A. Ando, R. Ikeshita, T. Kano, M. Delcroix, T. Nakatani, T. Asami, S. Araki, NTT multi-speaker ASR system for the DASR task of CHiME-7 challenge, in: Proc. CHiME, 2023, pp. 45–50. 49

  12. [20]

    Mitrofanov, T

    A. Mitrofanov, T. Prisyach, T. Timofeeva, S. Novoselov, M. Korenevsky, Y. Khokhlov, A. Akulov, A. Anikin, R. Khalili, I. Lezhenin, A. Melnikov, D. Miroshnichenko, N. Mamaev, I. Odegov, O. Rudnitskaya, A. Roma- nenko, STCON system for the CHiME-8 challenge, in: Proc. CHiME, 202...

  13. [21]

    N. Kamo, N. Tawara, A. Ando, T. Kano, H. Sato, R. Ikeshita, T. Moriya, S. Horiguchi, K. Matsuura, A. Ogawa, A. Plaquet, T. Ashihara, T. Ochiai, M. Mimura, M. Delcroix, T. Nakatani, T. Asami, S. Araki, NTT multi- speaker ASR system for the DASR task of CHiME-8 challenge, in: Pr...

  14. [22]

    S. Niu, R. Wang, J. Du, G. Yang, Y. Tu, S. Wu, S. Qian, H. Wu, H. Xu, X. Zhang, G. Zhong, X. Yu, J. Chen, M. Wang, D. Cai, T. Gao, G. Wan, F. Ma, J. Pan, J. Gao, The USTC-NERCSLIP systems for the CHiME-8 NOTSOF AR-1 challenge, in: Proc. CHiME, 2024, pp. 31–36

  15. [24]

    Kinoshita, M

    K. Kinoshita, M. Delcroix, N. Tawara, Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds, in: Proc. ICASSP, 2021, pp. 7198–7202

  16. [25]

    Bredin, A

    H. Bredin, A. Laurent, End-to-end speaker segmentation for overlap-aware resegmentation, in: Proc. Interspeech, 2021, pp. 3111–3115

  17. [26]

    Medennikov, M

    I. Medennikov, M. Korenevsky, T. Prisyach, Y. Khokhlov, M. Ko- renevskaya, I. Sorokin, T. Timofeeva, A. Mitrofanov, A. Andrusenko, I. Podluzhny, A. Laptev, A. Romanenko, Target-speaker voice activity de- tection: a novel approach for multi-speaker diarization in a dinner party...

  18. [27]

    M.-K. He, J. Du, Q.-F. Liu, C.-H. Lee, ANSD-MA-MSE: Adaptive neu- ral speaker diarization using memory-aware multi-speaker embedding, IEEE/ACM Trans. Audio Speech Lang. Process. 31 (2023) 1561–1573

  19. [28]

    Yoshioka, T

    T. Yoshioka, T. Nakatani, Generalization of multi-channel linear prediction methods for blind MIMO impulse response shortening, IEEE Trans. Audio Speech Lang. Process. 20 (10) (2012) 2707–2720

  20. [29]

    Nakatani, T

    T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, B.-H. Juang, Speech dereverberation based on variance-normalized delayed linear prediction, IEEE Trans. Audio Speech Lang. Process. 18 (7) (2010) 1717–1731

  21. [30]

    Benesty, J

    J. Benesty, J. Chen, Y. Huang, Noncausal (Frequency-Domain) Optimal Filters, Springer Berlin Heidelberg, Berlin, Heidelberg, 2008, Ch. 6, pp. 115–137. 50

  22. [31]

    Cornelis, M

    B. Cornelis, M. Moonen, J. Wouters, Performance analysis of multichan- nel Wiener filter-based noise reduction in hearing aids under second or- der statistics estimation errors, IEEE Trans. Audio Speech Lang. Process. 19 (5) (2011) 1368–1381

  23. [32]

    J. G. Fiscus, A post-processing system to yield reduced word error rates: Recognizer Output Voting Error Reduction (ROVER), in: Proc. ASRU, 1997, pp. 347–354

  24. [33]

    D. Raj, P. Garcia, Z. Huang, S. Watanabe, D. Povey, A. Stolcke, S. Khu- danpur, DOVER-Lap: A method for combining overlap-aware diarization outputs, in: Proc. SLT, 2021, pp. 881–888

  25. [34]

    Horiguchi, Y

    S. Horiguchi, Y. Takashima, P. Garcia, S. Watanabe, Y. Kawaguchi, Multi- channel end-to-end neural diarization with distributed microphones, in: Proc. ICASSP, 2022, pp. 7332–7336

  26. [35]

    Z. Lu, Y. Wang, Y. Zhang, W. Han, Z. Chen, P. Haghani, Unsupervised data selection via discrete speech representation for ASR, in: Proc. Inter- speech, 2022, pp. 3393–3397

  27. [36]

    Tawara, M

    N. Tawara, M. Delcroix, A. Ando, A. Ogawa, NTT speaker diarization sys- tem for CHiME-7: Multi-domain, multi-microphone end-to-end and vector clustering diarization, in: Proc. ICASSP, 2024, pp. 11281–11285

  28. [37]

    Tawara, A

    N. Tawara, A. Ando, S. Horiguchi, M. Delcroix, Multi-channel speaker counting for EEND-VC-based speaker diarization on multi-domain conver- sation, in: Proc. ICASSP, 2025

  29. [38]

    von Neumann, C

    T. von Neumann, C. Boeddeker, M. Delcroix, R. Haeb-Umbach, MeetEval: A toolkit for computation of word error rates for meeting transcription systems, in: Proc. CHiME, 2023, pp. 27–32

  30. [39]

    Medennikov, M

    I. Medennikov, M. Korenevsky, et al., Target-speaker voice activity de- tection: a novel approach for multi-speaker diarization in a dinner party scenario, in: Proc. Interspeech, 2020, pp. 274–278

  31. [40]

    Kinoshita, M

    K. Kinoshita, M. Delcroix, N. Tawara, Advances in integration of end-to- end neural and clustering-based diarization for real conversational speech, in: Proc. Interspeech, 2021, pp. 3565–3569

  32. [41]

    Desplanques, J

    B. Desplanques, J. Thienpondt, K. Demuynck, ECAPA-TDNN: Empha- sized channel attention, propagation and aggregation in TDNN based speaker verification, in: Proc. Interspeech, 2020, pp. 3830–3834

  33. [42]

    J. H. Ward Jr., Hierarchical grouping to optimize an objective function, J. Am. Stat. Assoc. 58 (301) (1963) 236–244

  34. [43]

    J. Shi, J. Malik, Normalized cuts and image segmentation, IEEE Trans. Pattern Anal. Mach. Intell. 22 (8) (2000) 888–905. 51

  35. [44]

    A. Y. Ng, M. I. Jordan, Y. Weiss, On spectral clustering: Analysis and an algorithm, in: Proc. NIPS, 2001, pp. 849–856

  36. [45]

    Bredin, pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe, in: Proc

    H. Bredin, pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe, in: Proc. Interspeech, 2023, pp. 1983–1987

  37. [46]

    T. J. Park, H. Huang, A. Juki´ c, K. Dhawan, K. C. Puvvada, N. R. Koluguri, N. Karpov, A. Laptev, J. Balam, B. Ginsburg, The CHiME-7 challenge: System description and performance of NeMo team’s DASR system, in: Proc. CHiME, 2023, pp. 57–62

  38. [47]

    T. J. Park, K. J. Han, M. Kumar, S. Narayanan, Auto-tuning spectral clus- tering for speaker diarization using normalized maximum eigengap, IEEE Signal Processing Letters 27 (2020) 381–385

  39. [48]

    Ishiguro, T

    K. Ishiguro, T. Yamada, S. Araki, T. Nakatani, H. Sawada, Probabilis- tic speaker diarization with bag-of-words representations of speaker angle information, IEEE Trans. Audio Speech Lang. Process. 20 (2) (2011) 447– 460

  40. [49]

    Boeddeker, A

    C. Boeddeker, A. S. Subramanian, G. Wichern, R. Haeb-Umbach, J. Le Roux, TS-SEP: Joint diarization and separation conditioned on esti- mated speaker embeddings, IEEE/ACM Trans. Audio, Speech Lang. Pro- cess. 32 (2024) 1185—-1197

  41. [50]

    Yoshioka, X

    T. Yoshioka, X. Wang, D. Wang, M. Tang, Z. Zhu, Z. Chen, N. Kanda, Vararray: Array-geometry-agnostic continuous speech separation, in: Proc. ICASSP, 2022, pp. 6027–6031

  42. [51]

    D. Wang, Z. Chen, T. Yoshioka, Neural speech separation using spa- tially distributed microphones, in: Interspeech 2020, 2020, pp. 339–343. doi:10.21437/Interspeech.2020-1089

  43. [52]

    S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, X. Yu, F. Wei, WavLM: Large-scale self-supervised pre-training for full stack speech processing, IEEE J. Sel. Top. Signal Process...

  44. [53]

    https://huggingface.co/speechbrain/spkrec-ecapa-voxceleb, accessed 11 December 2024

  45. [54]

    Nagrani, J

    A. Nagrani, J. S. Chung, W. Xie, A. Zisserman, VoxCeleb: Large-scale speaker verification in the wild, Comput. Speech Lang. 60 (2020) 101027

  46. [55]

    G. Yang, M. He, S. Niu, R. Wang, Y. Yue, S. Qian, S. Wu, J. Du, C.-H. Lee, Neural speaker diarization using memory-aware multi-speaker embed- ding with sequence-to-sequence architecture, in: Proc. ICASSP, 2024, pp. 11626–11630. 52

  47. [56]

    Cornell, T

    S. Cornell, T. Park, S. Huang, C. Boeddeker, X. Chang, M. Maciejewski, M. Wiesner, P. Garcia, S. Watanabe, The CHiME-8 DASR challenge for generalizable and array agnostic distant automatic speech recognition and diarization, in: Proc. CHiME, 2024, pp. 1–6

  48. [57]

    Yamashita, S

    N. Yamashita, S. Horiguchi, T. Homma, Improving the naturalness of sim- ulated conversations for end-to-end neural diarization, in: Proc. Odyssey, 2022, pp. 133–140

  49. [58]

    Panayotov, G

    V. Panayotov, G. Chen, D. Povey, S. Khudanpur, LibriSpeech: an ASR corpus based on public domain audio books, in: Proc. ICASSP, 2015, pp. 5206–5210

  50. [59]

    T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, S. Khudanpur, A study on data augmentation of reverberant speech for robust speech recognition, in: Proc. ICASSP, 2017, pp. 5220–5224

  51. [60]

    Snyder, G

    D. Snyder, G. Chen, D. Povey, MUSAN: A music, speech, and noise corpus, arXiv:1510.08484 (2015)

  52. [61]

    Fujita, N

    Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, S. Watanabe, End-to- end neural speaker diarization with permutation-free objectives, in: Proc. Interspeech, 2019, pp. 4300–4304

  53. [62]

    M. Wolf, C. Nadeu, Channel selection measures for multi-microphone speech recognition, Speech Commun. 57 (2014) 170–180

  54. [63]

    Souden, J

    M. Souden, J. Benesty, S. Affes, On optimal frequency-domain multichan- nel linear filtering for noise reduction, IEEE Trans. Audio Speech Lang. Process. 18 (2) (2009) 260–276

  55. [64]

    T. C. Lawin-Ore, S. Doclo, Reference microphone selection for MWF-based noise reduction using distributed microphone arrays, in: Speech Commu- nication; 10. ITG Symposium, 2012, pp. 1–4

  56. [65]

    Erdogan, J

    H. Erdogan, J. R. Hershey, S. Watanabe, M. I. Mandel, J. L. Roux, Improved MVDR beamforming using single-channel mask prediction net- works, in: Proc. Interspeech, 2016, pp. 1981–1985

  57. [66]

    Warsitz, R

    E. Warsitz, R. Haeb-Umbach, Blind acoustic beamforming based on gener- alized eigenvalue decomposition, IEEE Trans. Audio Speech Lang. Process. 15 (5) (2007) 1529–1539

  58. [67]

    P. A. Naylor, N. D. Gaubitch, Speech dereverberation, Springer Science & Business Media, 2010

  59. [68]

    Lavechin, M

    M. Lavechin, M. M´ etais, H. Titeux, A. Boissonnet, J. Copet, M. Rivi` ere, E. Bergelson, A. Cristia, E. Dupoux, H. Bredin, Brouhaha: Multi-task training for voice activity detection, speech-to-noise ratio, and C50 room acoustics estimation, in: Proc. ASRU, 2023, pp. 1–7. 53

  60. [69]

    Gannot, E

    S. Gannot, E. Vincent, S. Markovich-Golan, A. Ozerov, A consolidated perspective on multimicrophone speech enhancement and source separa- tion, IEEE/ACM Transactions on Audio, Speech, and Language Processing 25 (4) (2017) 692–730

  61. [70]

    Doclo, M

    S. Doclo, M. Moonen, GSVD-based optimal filtering for single and multimi- crophone speech enhancement, IEEE Trans. Signal Process. 50 (9) (2002) 2230–2244

  62. [71]

    Spriet, M

    A. Spriet, M. Moonen, J. Wouters, Spatially pre-processed speech distor- tion weighted multi-channel Wiener filtering for noise reduction, Signal Process. 84 (12) (2004) 2367–2387

  63. [72]

    Doclo, A

    S. Doclo, A. Spriet, J. Wouters, M. Moonen, Frequency-domain criterion for the speech distortion weighted multichannel Wiener filter for robust noise reduction, Speech Commun. 49 (7) (2007) 636–656

  64. [73]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in NIPS, 2017, pp. 5998–6008

  65. [74]

    Graves, Sequence transduction with recurrent neural networks, in: ICML Workshop on Representation Learning, 2012

    A. Graves, Sequence transduction with recurrent neural networks, in: ICML Workshop on Representation Learning, 2012

  66. [75]

    Radford, J

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large-scale weak supervision, in: Proc. ICML, 2023, pp. 28492–28518

  67. [76]

    Sennrich, B

    R. Sennrich, B. Haddow, A. Birch, Neural machine translation of rare words with subword units, in: Proc. ACL, 2016, pp. 1715–1725

  68. [77]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language models are unsupervised multitask learners, OpenAI blog 1 (8) (2019) 9

  69. [78]

    Kuchaiev, J

    O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V. Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, J. M. Cohen, NeMo: A toolkit for building AI applications using neural modules, arXiv:1909.09577 (2019)

  70. [79]

    Gulati, J

    A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, R. Pang, Conformer: Convolution-augmented transformer for speech recognition, in: Proc. Interspeech, 2020, pp. 5036– 5040

  71. [80]

    Rekesh, N

    D. Rekesh, N. R. Koluguri, S. Kriman, S. Majumdar, V. Noroozi, H. Huang, O. Hrinchuk, K. Puvvada, A. Kumar, J. Balam, B. Ginsburg, Fast Con- former with linearly scalable attention for efficient speech recognition, in: Proc. ASRU, 2023, pp. 1–8. 54

  72. [81]

    K. Kim, F. Wu, Y. Peng, J. Pan, P. Sridhar, K. J. Han, S. Watanabe, E-Branchformer: Branchformer with enhanced merging for speech recogni- tion, in: Proc. SLT, 2023, pp. 84–91

  73. [82]

    Ogawa, N

    A. Ogawa, N. Tawara, M. Delcroix, S. Araki, Lattice rescoring based on large ensemble of complementary neural language models, in: Proc. ICASSP, 2022, pp. 6517–6520

  74. [83]

    Ogawa, N

    A. Ogawa, N. Kamo, K. Matsuura, T. Ashihara, T. Moriya, T. Kano, N. Tawara, M. Delcroix, Applying LLMs for rescoring N-best ASR hy- potheses of casual conversations: Effects of domain adaptation and context carry-over, arXiv:2406.18972 (2024)

  75. [84]

    Povey, A

    D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, K. Vesely, The Kaldi speech recognition toolkit, in: Proc. ASRU, 2011

  76. [85]

    Andrusenko, R

    A. Andrusenko, R. Nasretdinov, A. Romanenko, Uconv-conformer: High reduction of input sequence length for end-to-end speech recognition, in: Proc. ICASSP, 2023, pp. 1–5

  77. [86]

    Watanabe, T

    S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. En- rique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, T. Ochiai, ESPnet: End-to-end speech processing toolkit, in: Proc. Inter- speech, 2018, pp. 2207–2211

  78. [87]

    Z. Yao, L. Guo, X. Yang, W. Kang, F. Kuang, Y. Yang, Z. Jin, L. Lin, D. Povey, Zipformer: A faster and better encoder for automatic speech recognition, in: Proc. ICLR, 2024

  79. [88]

    Hirano, M

    Y. Hirano, M. Nguyen, K. Azuma, J. M. Saragih, S. Sakti, The NAIST system for the CHiME-8 NOTSOF AR-1 task, in: Proc. CHiME, 2024, pp. 59–63

  80. [89]

    Huang, Y

    K. Huang, Y. Li, Z. Wang, H. Wang, W. Rao, Z. Sun, Z. Tang, S. Huang, Y. Wang, T. Yu, L. Xie, S. dong Shang, The NPU-TEA system for the CHiME-8 NOTSOF AR-1 challenge, in: Proc. CHiME, 2024, pp. 45–48

  81. [90]

    D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, Q. V. Le, SpecAugment: A simple data augmentation method for auto- matic speech recognition, in: Proc. Interspeech, 2019, pp. 2613–2617. 55

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.