REVIEW 3 major objections 5 minor 39 references
REAL-TSE challenge shows real-data adaptation consistently beats synthetic-only training for extracting a target speaker from real overlapping conversations, and no leading system dominates all quality metrics.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
The REAL-TSE challenge benchmarks target-speaker extraction on real bilingual conversational recordings with online and offline tracks, reporting that top systems beat baselines but no single system wins all quality metrics.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection A genuinely useful bilingual real-conversation TSE benchmark with an honest analysis of metric gaming, but the post-hoc metric switch and self-reported adaptation claims need careful framing. the 3 major comments →
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that REAL-TSE is the first benchmark to assess target speaker extraction on real, naturally overlapping conversational recordings in two languages under streaming and offline constraints, and that the submitted systems demonstrate two lessons: real-data adaptation is the largest driver of improvement over synthetic-only training, and no single system wins on all four metrics (TER, speaker similarity, DNSMOS-P808, and target-activity F1). The challenge also reveals that DNSMOS OVRL, a public reference-free quality metric, could be inflated by adversarial waveform perturbations and score-based selection without genuine perceptual improvement, leading the organizers
What carries the argument
The central object is the REAL-TSE evaluation corpus: 6,991 mixture–enrollment trials over 2,309 real conversational mixtures (11.3 h) in Mandarin and English, with a seen in-domain set (EVAL-1) and a newly recorded unseen set (EVAL-2) covering meetings, cafés, homes, and vehicles with matched and mismatched microphones. The evaluative mechanism is a four-metric protocol (Token Error Rate for intelligibility, speaker-embedding cosine similarity for consistency, DNSMOS-P808 for perceptual quality, and a target-activity F1) combined with a dense ranking that averages ranks across metrics, plus a perturbation-based latency verification for the online track.
Load-bearing premise
The claim that real-data adaptation is what made top systems succeed rests on teams' self-reported training practices without a controlled comparison; if those reports are incomplete or confounded by differences in architecture, compute, or data scale, the adaptation lesson does not follow.
What would settle it
Run the same BSRNN extractor twice, once trained only on synthetic Libri2Mix-style data and once with an added pseudo-labeled real conversational corpus, controlling for architecture, compute, and data size; if the synthetic-only system matches the adapted one on EVAL-1 and EVAL-2 TER, SpkSim, P808, and F1, the paper's central adaptation lesson is falsified. Additionally, a listening test comparing the top OVRL-spiked outputs against the top P808 outputs should show that OVRL spiking does not correspond to higher perceived quality.
If this is right
- If real-data adaptation is the dominant factor, system development should invest in pseudo-labeling and iterative fine-tuning on real far-field recordings rather than only in simulated mixtures.
- The multi-dimensional scoring means leaderboards should be reported by use case (intelligibility, quality, identity, activity) rather than a single aggregate rank.
- Public non-intrusive metrics must be treated as optimizable targets; future challenges should forbid using them during development or require reporting of metric-aware post-processing.
- The online track's measured-latency verification shows that analytic latency estimates are unreliable; standardized perturbation-based latency measurement should be standard for streaming TSE.
- Unseen scenario labels (meeting, restaurant, car) do not reliably predict difficulty; sample-level characteristics like target activity ratio are better predictors.
Where Pith is reading between the lines
- A controlled ablation with identical architecture and training recipe, differing only in the inclusion of real-data adaptation, would directly test the paper's main lesson; the current evidence is from teams' self-reports and is likely confounded.
- The OVRL over-optimization finding suggests that other public reference-free metrics (e.g., speaker similarity) could similarly be gamed; future defenses may need to use hidden or randomized metric variants.
- The finding that H2 (farthest) microphones hurt all metrics more than enrollment–mixture mismatch suggests that far-field array placement, not enrollment channel diversity, is the bottleneck for robustness.
- If scenario labels are poor difficulty predictors, a natural extension is to build difficulty-aware sampling for training sets, weighting by measured target activity and overlap rather than scene tags.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge for target speaker extraction (TSE) from real Mandarin and English conversational recordings. The task provides a multi-speaker mixture and enrollment utterance(s), and systems must output the target speaker's speech. The challenge defines online (≤100 ms latency) and offline tracks and evaluates submissions with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. The paper describes the datasets (6,991 trials, including a newly recorded EVAL-2 set), four BSRNN baselines, the 24 submitted systems, condition-wise analyses, and a post-submission reliability analysis that led to replacing DNSMOS OVRL with DNSMOS-P808 as the official perceptual-quality metric. The main empirical lessons are that real-data adaptation appears to outperform synthetic-only training, no leading system uniformly dominates across all dimensions, and public reference-free metrics such as DNSMOS OVRL can be over-optimized until they decouple from human perception.
Significance. If the results hold, REAL-TSE provides a valuable new evaluation resource for TSE: it is among the first benchmarks to use real conversational recordings, with both online and offline tracks, and it releases baselines, toolkits, and a large evaluation set. The paper's cautionary analysis of DNSMOS OVRL over-optimization, supported by human-MOS correlations, is an important contribution to benchmark design. The authors are also to be credited for releasing the WESEP baselines, submission-checking and metric-computation toolkits, and for collecting 24 independent system submissions. However, the empirical conclusions are weakened by two load-bearing issues: the official perceptual-quality metric was changed after submissions, and the claim that real-data adaptation is the cause of top-system performance is inferred from self-reports rather than controlled comparisons.
major comments (3)
- [§II-C, §VII, Figs. 1–4] The official perceptual-quality metric was changed from DNSMOS OVRL to DNSMOS P808 after all submissions were received. Participants developed, validated, and selected systems under OVRL, yet the final rankings and all aggregate analyses use P808. The paper never reports what the OVRL-based rankings would have been, nor does it qualify the findings as conditional on this metric switch. This confound could systematically alter which systems are top-ranked and thus the 'no leading system uniformly dominates' pattern and the condition-wise trends. Please provide a comparison of rankings under OVRL versus P808 and discuss how the conclusions in Sections VI and VIII change under each metric.
- [§V-A] The statement that 'real-data adaptation consistently outperformed training based solely on synthetic data' is a causal claim derived from self-reported system descriptions and aggregate scores, without a controlled comparison. The top teams differ in architecture, compute, training data scale, and other factors, so the observed superiority cannot be attributed to real-data adaptation alone. Please weaken this claim to a reported trend, or provide supporting evidence (e.g., ablations from at least a subset of systems) that isolates the effect of real-data adaptation.
- [§VII, Table VI] The decision to replace OVRL with P808 rests on the correlation with human MOS in Table VI, but the paper provides no methodology for the human listening test: no number of listeners, number of stimuli per condition, rating scale, or screening procedure. Without these details, the reliability analysis cannot be evaluated. Please include the listening-test protocol or a supplementary description.
minor comments (5)
- [§VI] Figures are referenced out of order: Fig. 3 is discussed before Fig. 2. Please renumber or reorder the references.
- [Table V] The caption says 'GRAY NUMBERS DENOTE THE PREVIOUSLY REPORTED DNSMOS OVRL SCORES,' but the table appears to have a dedicated OVRL column. Clarify what the gray text refers to.
- [§V-C] The mention of 'adversarial waveform perturbations' is intriguing but lacks detail. Please provide at least a one-sentence description or a reference to a future report.
- [Throughout] The text inconsistently uses 'EV AL' and 'EVAL' (e.g., 'EV AL-1' vs 'EVAL-1'). Please standardize.
- [References] Reference [17] appears as 'Fireredasr2s'; the official name is likely 'FireRedASR' or similar. Please verify capitalization.
Circularity Check
No circularity: the paper's empirical claims rest on externally submitted systems and independently collected evaluation data, not on a derivation that reduces to its own inputs.
full rationale
The paper is a challenge overview rather than a derivation. Its central empirical claims are aggregate scores of 24 externally submitted systems on newly recorded EVAL-2 data plus REAL-T-derived development data. Metrics (TER, SpkSim, DNSMOS, F1) are defined with external components (Zipformer, WeSpeaker, DNSMOS, FireRedV AD) and are not fitted from the submissions. The conclusion that real-data adaptation helps is based on teams' self-reported practices and observed score deltas, not on a parameter fitted to the same scores and then renamed as a prediction; it is an observational summary, so it may be confounded but is not circular. The post-submission switch from DNSMOS OVRL to P808, based on correlation with human MOS on the same final submissions, is a selection/overfitting risk and a protocol confound, but the paper explicitly reports OVRL as reference and does not present P808 as an independent prediction derived from OVRL; hence it is not a circular reduction under the criteria here. Self-citations to REAL-T [11] and WESEP [12] provide the dataset and baseline toolkit; the evaluation data include an independently collected unseen subset, and the central leaderboard results come from outside teams, so these citations are not load-bearing in a way that forces the conclusions. No quoted equation or fitted parameter reduces to its own input.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption The official ASR backbone (Zipformer) transcribes extracted speech accurately enough that TER is a valid measure of target-speech intelligibility.
- domain assumption WeSpeaker ResNet-34 speaker embeddings and cosine similarity reflect target-speaker consistency.
- domain assumption DNSMOS-P808 correlates with human perceptual quality better than OVRL in this setting.
- domain assumption FireRedV AD activity detection correctly identifies target-speaker regions for the F1 metric.
- domain assumption Enrollment utterances are non-overlapping, single-speaker, and belong to the target speaker in every trial.
- domain assumption EVAL-2 recordings are natural conversations and the target-speaker annotations are correct.
Cite this review
Pith. "Pith review of SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings." pith.science (2026). https://pith.science/paper/JETX3NRS
@misc{pith2026260715198,
author = {Pith},
title = {Pith review of: SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings},
year = {2026},
howpublished = {\url{https://pith.science/paper/JETX3NRS}},
note = {Machine review of arXiv:2607.15198}
}
read the original abstract
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.
Figures
Reference graph
Works this paper leans on
-
[1]
V oiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,
Q. Wang, H. Muckenhirn, K. Wilson, P. Sridhar, Z. Wu, J. Hershey, R. A. Saurous, R. J. Weiss, Y . Jia, and I. L. Moreno, “V oiceFilter: Targeted voice separation by speaker-conditioned spectrogram masking,” inProc. Interspeech, 2019, pp. 2728–2732
2019
-
[2]
SpEx: Multi-scale time domain speaker extraction network,
C. Xu, W. Rao, E. S. Chng, and H. Li, “SpEx: Multi-scale time domain speaker extraction network,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 1370–1384, 2020
2020
-
[3]
Neural target speech extraction: An overview,
K. Zmolikova, M. Delcroix, T. Ochiai, K. Kinoshita, J. ˇCernock`y, and D. Yu, “Neural target speech extraction: An overview,”IEEE Signal Processing Magazine, vol. 40, no. 3, pp. 8–29, 2023
2023
-
[4]
LibriMix: An open-source dataset for generalizable speech separation,
J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vincent, “LibriMix: An open-source dataset for generalizable speech separation,” arXiv preprint arXiv:2005.11262, 2020
Pith/arXiv arXiv 2005
-
[5]
Deep clustering: Discriminative embeddings for segmentation and separation,
J. R. Hershey, Z. Chen, J. L. Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” inProc. ICASSP, 2016, pp. 31–35
2016
-
[6]
The Fifth ’CHiME’ Speech Separation and Recognition Challenge: Dataset, Task and Baselines,
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, “The Fifth ’CHiME’ Speech Separation and Recognition Challenge: Dataset, Task and Baselines,” inProc. Interspeech, 2018, pp. 1561–1565
2018
-
[7]
CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings,
S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V . Manohar, D. Povey, D. Raj, D. Snyder, A. S. Subramanian, J. Trmal, B. B. Yair, C. Boeddeker, Z. Ni, Y . Fujita, S. Horiguchi, N. Kanda, T. Yoshioka, and N. Ryant, “CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings,” in6th International...
2020
-
[8]
Descriptor: Enhancing conversations for the hearing impaired in the 9th computational hearing in multisource environments challenge (CHiME9 ECHI),
R. Sutherland, J. Clarke, H. Elghazaly, T. Kuebert, M. Lugger, S. Pe- trausch, J. A. Ortiz, B. Xu, S. Goetze, and J. Barker, “Descriptor: Enhancing conversations for the hearing impaired in the 9th computational hearing in multisource environments challenge (CHiME9 ECHI),”IEEE Data Descriptions, vol. 3, pp. 73–81, 2026
2026
-
[9]
M2MeT: The ICASSP 2022 multi- channel multi-party meeting transcription challenge,
F. Yu, S. Zhang, Y . Fu, L. Xie, S. Zheng, Z. Du, W. Huang, P. Guo, Z. Yan, B. Ma, X. Xu, and H. Bu, “M2MeT: The ICASSP 2022 multi- channel multi-party meeting transcription challenge,” inProc. ICASSP. IEEE, 2022, pp. 6167–6171
2022
-
[10]
AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,
Y . Fu, L. Cheng, S. Lv, Y . Jv, Y . Kong, Z. Chen, Y . Hu, L. Xie, J. Wu, H. Bu, X. Xu, J. Du, and J. Chen, “AISHELL-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario,” inProc. Interspeech, 2021, pp. 3665–3669
2021
-
[11]
REAL-T: Real Conversational Mixtures for Target Speaker Extraction,
S. Li, S. Wang, J. Han, K. Zhang, W. Wang, and H. Li, “REAL-T: Real Conversational Mixtures for Target Speaker Extraction,” inProc. Interspeech, 2025, pp. 1923–1927
2025
-
[12]
WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction,
S. Wang, K. Zhang, S. Lin, J. Li, X. Wang, M. Ge, J. Yu, Y . Qian, and H. Li, “WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction,” inProc. Interspeech, 2024, pp. 4273–4277
2024
-
[13]
Zipformer: A faster and better encoder for automatic speech recognition,
Z. Yao, L. Guo, X. Yang, W. Kang, F. Kuang, Y . Yang, Z. Jin, L. Lin, and D. Povey, “Zipformer: A faster and better encoder for automatic speech recognition,” inProc. ICLR, vol. 2024, 2024, pp. 44 440–44 455
2024
-
[14]
Wespeaker: A research and production oriented speaker embedding learning toolkit,
H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y . Deng, and Y . Qian, “Wespeaker: A research and production oriented speaker embedding learning toolkit,” inProc. ICASSP. IEEE, 2023, pp. 1–5
2023
-
[15]
DNSMOS P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,
C. K. A. Reddy, V . Gopal, and R. Cutler, “DNSMOS P.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,” inProc. ICASSP, 2022, pp. 886–890
2022
-
[16]
DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,
C. K. Reddy, V . Gopal, and R. Cutler, “DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” inProc. ICASSP. IEEE, 2021, pp. 6493–6497
2021
-
[17]
Fireredasr2s: A state-of-the-art industrial-grade all-in-one automatic speech recognition system,
K. Xu, Y . Jia, K. Huang, J. Chen, W. Li, K. Liu, F.-L. Xie, X. Tang, and Y . Hu, “Fireredasr2s: A state-of-the-art industrial-grade all-in-one automatic speech recognition system,”arXiv preprint arXiv:2603.10420, 2026
arXiv 2026
-
[18]
Image method for efficiently simulating small-room acoustics,
J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small-room acoustics,”The Journal of the Acoustical Society of America, vol. 65, no. 4, pp. 943–950, 1979
1979
-
[19]
The AMI meeting corpus: A pre-announcement,
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V . Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lin- coln, A. Lisowska, I. McCowan, W. Post, D. Reidsma, and P. Wellner, “The AMI meeting corpus: A pre-announcement,” inMachine Learning for Multimodal Interaction (MLMI 2005), ser. Lecture Notes in Computer Science, S. R...
2005
-
[20]
DiPCo – Dinner Party Corpus,
M. Van Segbroeck, A. Zaid, K. Kutsenko, C. Huerta, T. Nguyen, X. Luo, B. Hoffmeister, J. Trmal, M. Omologo, and R. Maas, “DiPCo – Dinner Party Corpus,” inProc. Interspeech, 2020, pp. 434–436
2020
-
[21]
Music source separation with band-split RNN,
Y . Luo and J. Yu, “Music source separation with band-split RNN,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 1215–1225, 2023
2023
-
[22]
Multi-level speaker representation for target speaker extraction,
K. Zhang, J. Li, S. Wang, Y . Wei, Y . Wang, Y . Wang, and H. Li, “Multi-level speaker representation for target speaker extraction,” inProc. ICASSP, 2025, pp. 1–5
2025
-
[23]
ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification,” inProc. Interspeech, 2020, pp. 3830–3834
2020
-
[24]
Librispeech: an asr corpus based on public domain audio books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” inProc. ICASSP. IEEE, 2015, pp. 5206–5210
2015
-
[25]
V oxceleb: A large-scale speaker identification dataset,
A. Nagrani, J. S. Chung, and A. Zisserman, “V oxceleb: A large-scale speaker identification dataset,” inProc. Interspeech, 2017, pp. 2616–2620
2017
-
[26]
CN-Celeb: A challenging chinese speaker recognition dataset,
Y . Fan, J. Kang, L. Li, K. Li, H. Chen, S. Cheng, P. Zhang, Z. Zhou, Y . Cai, and D. Wang, “CN-Celeb: A challenging chinese speaker recognition dataset,” inProc. ICASSP, 2020, pp. 7604–7608
2020
-
[27]
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment (O-COCOSDA), 2017, pp. 1–5
2017
-
[28]
CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit,
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2019, sound dataset. [Online]. Available: https://doi.org/10.7488/ds/2645
doi:10.7488/ds/2645 2019
-
[29]
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,
J. Richter, Y .-C. Wu, S. Krenn, S. Welker, B. Lay, S. Watanabe, A. Richard, and T. Gerkmann, “EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,” in Proc. Interspeech, 2024, pp. 4873–4877
2024
-
[30]
WHAM!: Extending Speech Separation to Noisy Environments,
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. L. Roux, “WHAM!: Extending Speech Separation to Noisy Environments,” inInterspeech 2019, 2019, pp. 1368–1372
2019
-
[31]
Musan: A music, speech, and noise corpus,
D. Snyder, G. Chen, and D. Povey, “Musan: A music, speech, and noise corpus,” 2015. [Online]. Available: https://arxiv.org/abs/1510.08484
Pith/arXiv arXiv 2015
-
[32]
DEMAND: A collection of multi- channel recordings of acoustic noise in diverse environments,
J. Thiemann, N. Ito, and E. Vincent, “DEMAND: A collection of multi- channel recordings of acoustic noise in diverse environments,” 2013, version 1.0. [Online]. Available: https://doi.org/10.5281/zenodo.1227121
-
[33]
The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results,
C. K. Reddy, V . Gopal, R. Cutler, E. Beyrami, R. Cheng, H. Dubey, S. Matusevych, R. Aichner, A. Aazami, S. Braun, P. Rana, S. Srinivasan, and J. Gehrke, “The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results,” inProc. Interspeech, 2020, pp. 2492–2496
2020
-
[34]
Building and evaluation of a real room impulse response dataset,
I. Szöke, M. Skácel, L. Mošner, J. Paliesek, and J. ˇCernocký, “Building and evaluation of a real room impulse response dataset,”IEEE Journal of Selected Topics in Signal Processing, vol. 13, no. 4, pp. 863–876, 2019
2019
-
[35]
Fast random approximation of multi-channel room impulse response,
Y . Luo and R. Gu, “Fast random approximation of multi-channel room impulse response,” in2024 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), 2024, pp. 449– 454
2024
-
[36]
TF- GridNet: Integrating full-and sub-band modeling for speech separation,
Z.-Q. Wang, S. Cornell, S. Choi, Y . Lee, B.-Y . Kim, and S. Watanabe, “TF- GridNet: Integrating full-and sub-band modeling for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 3221–3236, 2023
2023
-
[37]
Notes on the history of correlation,
K. Pearson, “Notes on the history of correlation,”Biometrika, vol. 13, no. 1, pp. 25–45, 1920
1920
-
[38]
The proof and measurement of association between two things
C. Spearman, “The proof and measurement of association between two things.” 1961
1961
-
[39]
DNSMOS Pro: A reduced-size DNN for probabilistic MOS of speech
F. Cumlin, X. Liang, V . Ungureanu, C. K. Reddy, C. Schüldt, and S. Chatterjee, “DNSMOS Pro: A reduced-size DNN for probabilistic MOS of speech.” inProc. Interspeech, 2024, pp. 4818–4822
2024
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.