REVIEW 4 major objections 6 minor 1 cited by
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read SQ-Whisper adapts Whisper to transcribe one voice from overlapping speech.
desk verdict Genuinely new speaker-query adaptation for Whisper with solid ablations, but the headline 15% claim is mis-stated and the comparison to TS-HuBERT may be inflated by pretraining overlap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the SQ-Former, a small Transformer adaptor with 16 trainable query vectors. In each block, the queries first attend to enrolled target-speaker features, then cross-attend to the Whisper encoder's mixture representation; the output is a fixed-length speaker prompt appended to the encoder input and inserted between special tokens in the decoder. An accompanying speaker contrastive loss pulls a prompt toward the enrollment of the same speaker and away from other speakers in the batch, which the authors show is worth roughly 5% absolute WER.
What would settle it
Hold out a subset of speakers during training and test with enrollment clips from those unseen speakers; if SQ-Whisper still separates them well, the model generalizes beyond memorized voices, but if WER degrades sharply or the speaker prompts stop separating, the claimed mechanism fails.
Extended reading notes
Core claim
The paper's central claim is that Whisper, trained only on single-speaker audio, can be steered to recognize overlapped speech by injecting a learned speaker prompt into both the encoder and the decoder. The prompt is produced by an SQ-Former module that lets trainable queries attend to the enrollment speech to absorb the target voice, then cross-attends to the mixture to find the parts of the acoustic representation belonging to that voice. A speaker contrastive loss makes the resulting prompts cluster by speaker identity, and the paper reports up to 15% and 10% relative WER reductions over TS-HuBERT on Libri2Mix and WSJ0-2Mix respectively. Prompting both the encoder and the decoder gives further gains, showing the speaker prompt is useful at both the acoustic and the linguistic decoding stages.
Load-bearing premise
The load-bearing premise is that a clean enrollment utterance of the target speaker is available at test time and that this speaker is actually present in the mixture; when enrollment is mismatched, the reported word error rate jumps from about 20% to nearly 72%.
Editorial extensions
If this is right
- With matched enrollment, SQ-Whisper outperforms the prior TS-HuBERT baseline by up to 15% relative WER on Libri2Mix and 10% on WSJ0-2Mix.
- Adding speed perturbation and the larger Train-360 set yields state-of-the-art WERs of 14.6% on Libri2Mix Test and 4.4% on WSJ0-2Mix Test.
- The method works with LoRA, freezing most of Whisper's weights: LoRA SQ-Whisper with 40.76M trainable parameters beats TS-HuBERT with 105.18M trainable parameters.
- The encoder-side and decoder-side prompts are complementary; removing the contrastive loss or using only one prompt degrades performance.
- On the AMI meeting corpus, SQ-Whisper reaches 22.0% WER, close to an SOT system pre-trained on 900k hours of multi-speaker data despite being adapted from single-speaker Whisper.
- Because SQ-Former is a generic Transformer adaptor, the authors argue the method can be transplanted to other Transformer-based speech foundation models.
- The learned prompts appear to encode speaker identity rather than just acoustic content: with mismatched enrollment the model collapses to a fixed output, so enrollment reliability is the key deployment risk.
- A testable extension is to increase the number of in-batch negative speakers during contrastive training to see whether the mismatched-enrollment failure rate drops.
Reading between the lines
- The authors do not test whether the SQ-Former transfers to self-supervised encoders such as HuBERT; a testable extension is to attach it to such a model and check whether the query-search mechanism works on representations learned without ASR supervision.
- The mismatched-enrollment collapse suggests the model may use the prompt as a hard gate rather than a soft bias; if true, training with deliberately mismatched enrollments could improve robustness.
- Because the contrastive loss currently uses only in-batch negatives, scaling to much larger batches or explicit speaker banks could make the prompts more discriminative and reduce the enrollment sensitivity.
- The prompt that separates speakers for ASR could also be probed for speaker verification or diarization, although the paper does not evaluate those tasks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SQ-Whisper, an adaptation of the Whisper speech foundation model to target-speaker ASR (TS-ASR). The method inserts a Speaker-Querying Transformer (SQ-Former) between Whisper's convolutional front-end and transformer encoder; a set of trainable queries cross-attends to the mixture representation conditioned on a target-speaker enrollment, producing speaker prompts that are appended to both encoder and decoder inputs. A speaker contrastive loss is added to make the prompts discriminative. Experiments on Libri2Mix, WSJ0-2Mix, and AMI compare SQ-Whisper with TSE-Whisper, TS-HuBERT, WavLM, and other baselines, and the paper reports state-of-the-art WERs of 14.6% on Libri2Mix and 4.4% on WSJ0-2Mix when using data augmentation. The paper also integrates and ablates four TSE modules and evaluates LoRA-based parameter-efficient fine-tuning.
Significance. If the results are reproducible, the contribution is valuable: it offers a principled way to extend a supervised speech foundation model to overlapped-speech recognition, shows consistent gains over a Whisper-based TSE baseline across three datasets, and releases code. The internal comparisons between TSE-Whisper and SQ-Whisper provide solid evidence for the adaptation method itself. The headline comparisons to TS-HuBERT are partially confounded by Whisper's pretraining data overlap, which tempers the strength of the state-of-the-art claims, but the underlying technical idea is sound and the paper is well structured.
major comments (4)
- [Abstract / Section VI-B] The abstract states that SQ-Whisper yields 'up to 15% relative reduction in WER' compared with TS-HuBERT, but Table II shows TS-HuBERT at 24.8 and SQ-Whisper (Full) at 20.1, which is a 19% relative reduction; the 15% figure in Section VI-B corresponds to the comparison against PIT-Transformer (23.5 to 20.1) and should be attributed to that baseline, not to TS-HuBERT.
- [Section II-A / Tables II and VI] Whisper medium is a 764M-parameter supervised ASR model trained on 680k hours of labeled web audio, which very likely includes LibriSpeech and possibly WSJ0, the source corpora of Libri2Mix and WSJ0-2Mix. The TS-HuBERT and WavLM baselines are self-supervised and were not trained on transcripts of these corpora, so the reported gains over TS-HuBERT (19% on Libri2Mix, 10% on WSJ0-2Mix) may partly reflect the base model's prior exposure to the test data rather than the SQ-Former. The internal TSE-Whisper comparisons are controlled and reassuring, but the abstract's headline claims should either be qualified or supported by a control experiment using a Whisper-style model that has no LibriSpeech/WSJ0 exposure (e.g., an OWSM model).
- [Section VI-F / Table V] The robustness analysis shows that with mismatched enrollment, WER degrades from 20.1% to 71.8% on the Libri2Mix Test set, which is worse than the vanilla Whisper baseline (54.3%). This failure mode is a central limitation of the method and should be clearly stated in the abstract and conclusion; the current framing of the method as robustly 'eliminating interfering speakers' is misleading without this qualifier.
- [Tables I-VII] All reported WERs come from a single training run without error bars or significance tests, yet the text repeatedly uses the word 'significant' (e.g., Sections VI-B and VII). Since some of the headline differences are small (e.g., 23.2 vs. 24.8 on Libri2Mix, 22.3 vs. 21.2 on AMI), the authors should report variance across seeds or perform a statistical test to substantiate these claims.
minor comments (6)
- [Section V-B] The temperature kappa in Eq. (14) is never given a value in the hyperparameter description; please report it.
- [Tables I and II] The parameter count for TSE-Whisper (Full) Add is 762.98M in Table I but 762.85M in Table II; please make these consistent.
- [Section VI-A] The statement that LoRA fine-tuned models outperform fully fine-tuned models 'across all adaptation methods' is contradicted by the Cat row in Table I (30.2 full vs. 31.2 LoRA on Test).
- [Figures 2 and 3] The captions and diagrams contain typos: 'Conv1D + GLUE' should be 'Conv1D + GELU', and 'Framwork' should be 'Framework'.
- [Section IV-D] 'liner projection' should be 'linear projection'.
- [Section V-A] Please clarify whether the 'Separation + Whisper' baseline in Tables I-II uses a separation model trained on the same Libri2Mix training data; if so, the comparison should be contextualized in the text because the separation model may be trained on the same test conditions.
Circularity Check
No circularity: the paper is an empirical adaptation study measured against external benchmarks, with no prediction that reduces to a fitted parameter or self-citation chain.
full rationale
SQ-Whisper is an empirical adaptation study. The claimed results (Tables I, II, VI, VII) are WERs on external benchmarks (Libri2Mix, WSJ0-2Mix, AMI) compared against baselines from independent groups (TS-HuBERT [28], WavLM [19], PT-Whisper [29]) plus the paper's own TSE-Whisper controls. No quantity stated as a prediction is derived from a fitted parameter or from an assumption that already contains the result. The SQ-Former's query count and loss weight are selected by dev-set ablations (Section VI-C and the alpha discussion in Section IV-B), which is standard model selection, and the contrastive loss is an auxiliary training regularizer, not the evaluation metric. Self-citations (ESPnet [54], OWSM [23,24], prior multi-speaker ASR papers [22,36,37,42,57]) are toolkit and background references and do not carry the TS-ASR comparison. The abstract's '15%' Libri2Mix figure is internally inconsistent (Table II gives 19% relative improvement over TS-HuBERT; 15% matches the PIT-Transformer comparison), but that is a reporting error, not circular reasoning. Likewise, the possible overlap between Whisper's pretraining data and LibriSpeech/WSJ0 is a fairness concern about the external comparison, not a circularity of the derivation chain.
Assumptions & free parameters
free parameters (6)
- Number of trainable queries Lq =
16
- Contrastive loss weight alpha =
20
- Number of negative samples K in contrastive loss =
10
- Temperature kappa in contrastive loss =
not reported
- SQ-Former hidden dimension Dq =
768
- Number of SQ-Former blocks =
2
assumptions (4)
- domain assumption Whisper encoder representations contain enough information to extract target-speaker attributes from mixtures.
- domain assumption Enrollment speech is clean, single-speaker, and comes from a speaker present in the mixture.
- domain assumption Speech segments in the AMI evaluation are available with human-annotated boundaries.
- standard math Standard backpropagation and Transformer attention converge with the combined cross-entropy and contrastive losses.
invented entities (2)
-
SQ-Former module
-
Speaker prompts (P-tilde)
Cite this review
Pith. "Pith review of SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR." pith.science (2026). https://pith.science/paper/33VLJ6PY
@misc{pith2026241205589,
author = {Pith},
title = {Pith review of: SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR},
year = {2026},
howpublished = {\url{https://pith.science/paper/33VLJ6PY}},
note = {Machine review of arXiv:2412.05589}
}
read the original abstract
Benefiting from massive and diverse data sources, speech foundation models exhibit strong generalization and knowledge transfer capabilities to a wide range of downstream tasks. However, a limitation arises from their exclusive handling of single-speaker speech input, making them ineffective in recognizing multi-speaker overlapped speech, a common occurrence in real-world scenarios. In this study, we delve into the adaptation of speech foundation models to eliminate interfering speakers from overlapping speech and perform target-speaker automatic speech recognition (TS-ASR). Initially, we utilize the Whisper model as the foundation for adaptation and conduct a thorough comparison of its integration with existing target-speaker adaptation techniques. We then propose an innovative model termed Speaker-Querying Whisper (SQ-Whisper), which employs a set number of trainable queries to capture speaker prompts from overlapping speech based on target-speaker enrollment. These prompts serve to steer the model in extracting speaker-specific features and accurately recognizing target-speaker transcriptions. Experimental results demonstrate that our approach effectively adapts the pre-trained speech foundation model to TS-ASR. Compared with the robust TS-HuBERT model, the proposed SQ-Whisper significantly improves performance, yielding up to 15% and 10% relative reductions in word error rates (WERs) on the Libri2Mix and WSJ0-2Mix datasets, respectively. With data augmentation, we establish new state-of-the-art WERs of 14.6% on the Libri2Mix Test set and 4.4% on the WSJ0-2Mix Test set. Furthermore, we evaluate our model on the real-world AMI meeting dataset, which shows consistent improvement over other adaptation methods.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
Diarization-conditioned Whisper with frame-level transforms and query-key biasing achieves strong target-speaker ASR on AMI, NOTSOFAR-1, Libri2Mix, and LibriCSS.
Reference graph
Works this paper leans on
-
[1]
BERT: Pre-training of deep bidirectional Transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional Transformers for language understanding,” in Proc. NAACL, 2019, pp. 4171–4186
work page 2019
-
[2]
Language models are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language models are few-shot learners,” in Proc. NeurIPS, vol. 33, 2020, pp. 1877–1901
work page 2020
-
[3]
LLaMA 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al. , “LLaMA 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023
arXiv 2023
-
[4]
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “GPT-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
arXiv 2023
-
[5]
Gemini: a family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, Y . Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth et al., “Gemini: a family of highly capable multimodal models,” arXiv preprint arXiv:2312.11805 , 2023
arXiv 2023
-
[6]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. ICLR, 2021
2021
-
[7]
Masked autoencoders are scalable vision learners,
K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proc. IEEE/CVF CVPR , 2022, pp. 16 000–16 009
work page 2022
-
[8]
Flamingo: a visual language model for few-shot learning,
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” in Proc. NeurIPS, vol. 35, 2022, pp. 23 716–23 736
work page 2022
Show all 57 references
-
[9]
BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” in Proc. ICML, 2023, pp. 19 730–19 742
2023
-
[10]
Self-supervised speech representation learning: A review,
A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe et al. , “Self-supervised speech representation learning: A review,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1179–1210, 2022
2022
-
[11]
Robust speech recognition via large-scale weak super- vision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak super- vision,” in Proc. ICML, 2023, pp. 28 492–28 518
2023
-
[12]
Google USM: Scaling automatic speech recognition beyond 100 languages,
Y . Zhang, W. Han, J. Qin, Y . Wang, A. Bapna, Z. Chen, N. Chen, B. Li, V . Axelrod, G. Wang et al. , “Google USM: Scaling automatic speech recognition beyond 100 languages,” arXiv preprint arXiv:2303.01037 , 2023
2023 arXiv
-
[13]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS, 2020
2020
-
[14]
Un- supervised cross-lingual representation learning for speech recognition,
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Un- supervised cross-lingual representation learning for speech recognition,” in Proc. Interspeech, 2020, pp. 2426–2430
2020
-
[15]
w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,
Y .-A. Chung, Y . Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y . Wu, “w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” in Proc. IEEE ASRU. IEEE, 2021, pp. 244–250
2021
-
[16]
Non-autoregressive predictive coding for learning speech representations from local dependencies,
A. H. Liu, Y .-A. Chung, and J. Glass, “Non-autoregressive predictive coding for learning speech representations from local dependencies,” arXiv preprint arXiv:2011.00406 , 2020
2011 arXiv
-
[17]
TERA: Self-supervised learning of transformer encoder representation for speech,
A. T. Liu, S.-W. Li, and H.-y. Lee, “TERA: Self-supervised learning of transformer encoder representation for speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2351–2366, 2021
2021
-
[18]
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3451–3460, 2021
2021
-
[19]
WavLM: Large-scale self-supervised pre- training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al., “WavLM: Large-scale self-supervised pre- training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
-
[20]
data2vec: A general framework for self-supervised learning in speech, vision and language,
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “data2vec: A general framework for self-supervised learning in speech, vision and language,” in Porc. ICML, 2022, pp. 1298–1312
2022
-
[21]
SUPERB: Speech processing universal performance benchmark,
S.-w. Yang, P.-H. Chi, Y .-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y . Y . Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin et al. , “SUPERB: Speech processing universal performance benchmark,” in Proc. Interspeech , 2021, pp. 1194–1198
2021
-
[22]
An exploration of self- supervised pretrained representations for end-to-end speech recognition,
X. Chang, T. Maekaku, P. Guo, J. Shi, Y .-J. Lu, A. S. Subramanian, T. Wang, S.-w. Yang, Y . Tsao, H.-y. Leeet al., “An exploration of self- supervised pretrained representations for end-to-end speech recognition,” in Proc. IEEE ASRU, 2021, pp. 228–235
2021
-
[23]
Reproducing Whisper-style training using an open-source toolkit and publicly available data,
Y . Peng, J. Tian, B. Yan, D. Berrebbi, X. Chang, X. Li, J. Shi, S. Arora, W. Chen, R. Sharma et al., “Reproducing Whisper-style training using an open-source toolkit and publicly available data,” inProc. IEEE ASRU, 2023, pp. 1–8
2023
-
[24]
OWSM v3. 1: Better and faster open Whisper-style speech models based on E-Branchformer,
Y . Peng, J. Tian, W. Chen, S. Arora, B. Yan, Y . Sudo, M. Shakeel, K. Choi, J. Shi, X. Chang et al., “OWSM v3. 1: Better and faster open Whisper-style speech models based on E-Branchformer,” arXiv preprint arXiv:2401.16658, 2024
2024 arXiv
-
[25]
Large-scale pre-training of end-to-end multi-talker ASR for meeting transcription with single distant microphone,
N. Kanda, G. Ye, Y . Wu, Y . Gaur, X. Wang, Z. Meng, Z. Chen, and T. Yoshioka, “Large-scale pre-training of end-to-end multi-talker ASR for meeting transcription with single distant microphone,” in Proc. Interspeech, 2021, pp. 3430–3434
2021
-
[26]
Adapting multi-lingual ASR models for handling multiple talkers,
C. Li, Y . Qian, Z. Chen, N. Kanda, D. Wang, T. Yoshioka, Y . Qian, and M. Zeng, “Adapting multi-lingual ASR models for handling multiple talkers,” in Proc. Interspeech, 2023, pp. 1314–1318
2023
-
[27]
Adapting self- supervised models to multi-talker speech recognition using speaker embeddings,
Z. Huang, D. Raj, P. Garc ´ıa, and S. Khudanpur, “Adapting self- supervised models to multi-talker speech recognition using speaker embeddings,” in Proc. IEEE ICASSP , 2023, pp. 1–5
2023
-
[28]
Weakly-supervised speech pre-training: A case study on target speech recognition,
W. Zhang and Y . Qian, “Weakly-supervised speech pre-training: A case study on target speech recognition,” in Proc. Interspeech , 2023, pp. 3517–3521
2023
-
[29]
Extending Whisper with prompt tuning to target-speaker ASR,
H. Ma, Z. Peng, M. Shao, J. Li, and J. Liu, “Extending Whisper with prompt tuning to target-speaker ASR,” in Proc. IEEE ICASSP , 2024, pp. 12 516–12 520
2024
-
[30]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones et al. , “Attention is all you need,” in Proc. NeurIPS, 2017, pp. 5998–6008
2017
-
[31]
Whisper-AT: Noise- robust automatic speech recognizers are also strong general audio event taggers,
Y . Gong, S. Khurana, L. Karlinsky, and J. Glass, “Whisper-AT: Noise- robust automatic speech recognizers are also strong general audio event taggers,” in Proc. Interspeech, 2023, pp. 2798–2802
2023
-
[32]
WhiSLU: End-to-end spoken language under- standing with whisper,
M. Wang, Y . Li, J. Guo, X. Qiao, Z. Li, H. Shang, D. Wei, S. Tao, M. Zhang, and H. Yang, “WhiSLU: End-to-end spoken language under- standing with whisper,” in Proc. Interspeech, 2023, pp. 770–774
2023
-
[33]
Distil-Whisper: Robust knowledge distillation via large-scale pseudo labelling,
S. Gandhi, P. von Platen, and A. M. Rush, “Distil-Whisper: Robust knowledge distillation via large-scale pseudo labelling,” arXiv preprint arXiv:2311.00430, 2023
2023 arXiv
-
[34]
Whisper- KDQ: A lightweight Whisper via guided knowledge distillation and quantization for efficient ASR,
H. Shao, W. Wang, B. Liu, X. Gong, H. Wang, and Y . Qian, “Whisper- KDQ: A lightweight Whisper via guided knowledge distillation and quantization for efficient ASR,” arXiv preprint arXiv:2305.10788, 2023
2023 arXiv
-
[35]
Recognizing multi-talker speech with permutation invariant training,
D. Yu, X. Chang, and Y . Qian, “Recognizing multi-talker speech with permutation invariant training,” in Proc. Interspeech, 2017, pp. 2456– 2460
2017
-
[36]
End-to- end multi-speaker speech recognition with Transformer,
X. Chang, W. Zhang, Y . Qian, J. Le Roux, and S. Watanabe, “End-to- end multi-speaker speech recognition with Transformer,” in Proc. IEEE ICASSP, 2020, pp. 6134–6138. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, MAY 2024 11
2020
-
[37]
Multi-speaker ASR combining non-autoregressive conformer CTC and conditional speaker chain,
P. Guo, X. Chang, S. Watanabe, and L. Xie, “Multi-speaker ASR combining non-autoregressive conformer CTC and conditional speaker chain,” in Proc. Interspeech, 2021, pp. 3720–3724
2021
-
[38]
Serialized output training for end-to-end overlapped speech recognition,
N. Kanda, Y . Gaur, X. Wang, Z. Meng, and T. Yoshioka, “Serialized output training for end-to-end overlapped speech recognition,” in Proc. Interspeech, 2020, pp. 2797–2801
2020
-
[39]
BA- SOT: Boundary-aware serialized output training for multi-talker ASR,
Y . Liang, F. Yu, Y . Li, P. Guo, S. Zhang, Q. Chen, and L. Xie, “BA- SOT: Boundary-aware serialized output training for multi-talker ASR,” in Proc. Interspeech, 2023, pp. 3487–3491
2023
-
[40]
Joint speaker counting, speech recognition, and speaker identification for overlapped speech of any number of speakers,
N. Kanda, Y . Gaur, X. Wang, Z. Meng, Z. Chen, T. Zhou, and T. Yoshioka, “Joint speaker counting, speech recognition, and speaker identification for overlapped speech of any number of speakers,” inProc. Interspeech, 2020, pp. 36–40
2020
-
[41]
A comparative study of modular and joint approaches for speaker-attributed ASR on monaural long-form audio,
N. Kanda, X. Xiao, J. Wu, T. Zhou, Y . Gaur, X. Wang, Z. Meng, Z. Chen, and T. Yoshioka, “A comparative study of modular and joint approaches for speaker-attributed ASR on monaural long-form audio,” inProc. IEEE ASRU, 2021, pp. 296–303
2021
-
[42]
Sa- Paraformer: Non-autoregressive end-to-end speaker-attributed ASR,
Y . Li, F. Yu, Y . Liang, P. Guo, M. Shi, Z. Du, S. Zhang, and L. Xie, “Sa- Paraformer: Non-autoregressive end-to-end speaker-attributed ASR,” in Proc. IEEE ASRU. IEEE, 2023, pp. 1–7
2023
-
[43]
Conformer- based target-speaker automatic speech recognition for single-channel audio,
Y . Zhang, K. C. Puvvada, V . Lavrukhin, and B. Ginsburg, “Conformer- based target-speaker automatic speech recognition for single-channel audio,” in Proc. IEEE ICASSP , 2023, pp. 1–5
2023
-
[44]
FiLM: Visual reasoning with a general conditioning layer,
E. Perez, F. Strub, H. De Vries, V . Dumoulin, and A. Courville, “FiLM: Visual reasoning with a general conditioning layer,” in Proc. AAAI , vol. 32, no. 1, 2018
2018
-
[45]
Conditionally adaptive multi-task learning: Improving transfer learning in NLP using fewer parameters & less data,
J. Pilault, A. Elhattami, and C. Pal, “Conditionally adaptive multi-task learning: Improving transfer learning in NLP using fewer parameters & less data,” in Proc. ICLR, 2021
2021
-
[46]
LoRA: Low-rank adaptation of large language models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. ICLR, 2022, pp. 1–13
2022
-
[47]
LibriMix: An open-source dataset for generalizable speech separation,
J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vincent, “LibriMix: An open-source dataset for generalizable speech separation,” arXiv preprint arXiv:2005.11262 , 2020
2005 arXiv
-
[48]
Deep clustering: Discriminative embeddings for segmentation and separation,
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Proc. IEEE ICASSP. IEEE, 2016, pp. 31–35
2016
-
[49]
The AMI meeting corpus: A pre-announcement,
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V . Karaiskos, W. Kraaij, M. Kronenthal et al. , “The AMI meeting corpus: A pre-announcement,” in Proc. MLMI, 2005, pp. 28– 39
2005
-
[50]
Librispeech: An ASR corpus based on public domain audio books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. IEEE ICASSP, 2015, pp. 5206–5210
2015
-
[51]
WHAM!: Extending speech separation to noisy environments,
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. L. Roux, “WHAM!: Extending speech separation to noisy environments,” in Proc. Interspeech, 2019, pp. 1368–1372
2019
-
[52]
The design for the Wall Street Journal-based CSR corpus,
D. B. Paul and J. Baker, “The design for the Wall Street Journal-based CSR corpus,” in Workshop on Speech and Natural Language , 1992
1992
-
[53]
V oxCeleb: A large-scale speaker identification dataset,
A. Nagrani, J. S. Chung, and A. Zisserman, “V oxCeleb: A large-scale speaker identification dataset,” in Proc. Interspeech , 2017, pp. 2616– 2620
2017
-
[54]
ESPnet: End-to-end speech processing toolkit,
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y . Unno, N. E. Y . Soplin, J. Heymann, M. Wiesner, N. Chen et al. , “ESPnet: End-to-end speech processing toolkit,” in Proc. Interspeech, 2018, pp. 2207–2211
2018
-
[55]
Front- end factor analysis for speaker verification,
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front- end factor analysis for speaker verification,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 19, no. 4, pp. 788–798, 2010
2010
-
[56]
X-Vectors: Robust DNN embeddings for speaker recognition,
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-Vectors: Robust DNN embeddings for speaker recognition,” in Proc. IEEE ICASSP, 2018, pp. 5329–5333
2018
-
[57]
Train from scratch: Single- stage joint training of speech separation and recognition,
J. Shi, X. Chang, S. Watanabe, and B. Xu, “Train from scratch: Single- stage joint training of speech separation and recognition,” Computer Speech & Language , vol. 76, p. 101387, 2022
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.