REVIEW 3 major objections 6 minor 32 references
Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Crowdsourced MUSHRA tests can stand in for expert lab tests when screening and quality-control steps are added.
desk verdict A useful, narrow empirical study showing crowdsourced MUSHRA can reproduce expert codec rankings for generative speech codecs, with Prolific closer in absolute terms than MTurk—but the expert ground truth is only 4–6 votes per file and the objective-metric analysis leans on substituted Prolific scores. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a screening-and-renormalization pipeline built around the MUSHRA protocol. It includes platform quality filters (97% success rate and 100 completed tasks), a digits-in-noise hearing test, a training MUSHRA question with three attempts, real-time score screening that discards responses deviating beyond thresholds, automatic test partitioning into smaller blocks, listener-level disqualification when the anchor is rated above the hidden reference or when all non-anchor scores are identical, inter-quartile-range outlier removal, and renormalization of results across sub-experiments so the hidden reference sits at 100 and the anchor matches across tests. The anchor is Opus 6 kbps, chosen because its coding artifacts resemble those of the test codecs, which the authors found easier for non-experts to judge than a low-pass-filtered anchor. This pipeline is what lets non-expert listeners produce expert-like rankings.
What would settle it
Run the same protocol on a held-out set of at least six generative and DSP codecs with both expert and Prolific listeners; if per-condition expert-to-Prolific correlation drops below about 0.9, or if SCOREQ's correlation with subjective scores no longer stays consistent between DNN and DSP codecs when computed per file, the central claims fail.
Extended reading notes
Core claim
The authors' central claim is that a crowdsourced MUSHRA protocol, once adapted with participant screening, a hearing test, training, real-time score screening, test partitioning, and post-screening outlier rules, yields results that are both repeatable and aligned with an internal expert MUSHRA test, even for generative speech codecs. With 40 English speech files and four codecs plus an Opus 6 kbps anchor, the per-file mean scores from Prolific correlate at 0.95 (Pearson) with the expert scores in two independent runs, MTurk at 0.89-0.90, and repeated crowdsourced runs correlate at 0.98-0.99 with each other. Prolific's absolute scores stay close to the expert curve, while MTurk reproduces rankings but drifts in absolute value. On the objective side, the paper finds that PESQ, POLQA, and ViSQOL correlate better for DSP codecs and systematically undervalue DNN codecs, while SCOREQ's correlation with subjective scores is stable across both architectures (-0.80 overall, -0.79 for both groups, with lower values meaning better quality), and NISQA and DNSMOS align poorly overall.
Load-bearing premise
All conclusions rest on treating the internal expert MUSHRA scores, just four to six votes per audio file, as a stable ground truth, and on accepting the later substitution of Prolific scores for those expert scores in the objective-metric analysis, with two added codec conditions validated only by an informal listening test.
Editorial extensions
If this is right
- Crowdsourced MUSHRA can substitute for expert lab tests during codec development, making frequent perceptual checks practical.
- Prolific is the better platform when absolute score alignment matters; MTurk still yields valid rankings, but its ceiling effect must be accounted for.
- PESQ, POLQA, and ViSQOL should not be relied on to rank neural codecs, because they systematically undervalue DNN-based systems.
- SCOREQ, as a reference-based contrastive metric, is a more architecture-consistent choice among the six metrics tested.
- MUSHRA test data can serve as reference ground truth for evaluating new objective metrics.
Reading between the lines
- A logical testable extension is to apply the same screening pipeline to other relative-judgment audio tasks, such as preference or intelligibility tests; the paper does not claim this, but nothing in its mechanism is task-specific.
- The objective-metric comparison is computed at condition level over a small codec set; per-file analysis with a wider set of generative codecs would be needed to confirm that SCOREQ's advantage generalizes.
- The platform difference may stem from different participant pools as much as from platform mechanics; a crossover design with the same listeners on both platforms would separate those effects.
- Because the anchor impairment was matched to coding artifacts, the protocol's conclusions may be specific to codec-like degradation; other artifact types may need their own anchor validation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes adaptations to the MUSHRA protocol for crowdsourced, non-expert listeners, focusing on generative speech codecs. It compares results from MTurk and Prolific against a small internal expert listening test, reports test–retest correlations, and evaluates six objective metrics against subjective scores. The main claims are that crowdsourced MUSHRA tests can reliably replace expert lab tests, that Prolific shows closer absolute alignment with expert scores while MTurk shows a ceiling effect, and that traditional objective metrics (PESQ, POLQA, ViSQOL) undervalue DNN-based codecs while SCOREQ is more architecture-consistent.
Significance. If the findings are robust, this is a practical contribution: it provides an open-source tool, a direct platform comparison, and evidence that architecture-aware objective metrics are needed for generative codecs. The test–retest correlations and the large per-file samples on the crowdsourcing platforms are concrete assets. The main caveat is that the validity claims hinge on a very small expert reference (4–6 votes per file) and on a post-hoc substitution of Prolific scores for expert scores in the objective-metric analysis; these weaken the force of the conclusions as currently presented.
major comments (3)
- [Section 3.1 / Table 1] The internal expert test used only 4–6 votes per file, yet it is treated as the ground truth for validating both platforms. Table 1 reports Pearson/Spearman correlations with this reference but gives no confidence intervals, significance tests, or per-file variance of the expert scores. With such a small rater pool, per-file means are highly sensitive to individual scale use, and the reported correlations could be inflated by shared response patterns rather than true alignment. The authors should quantify the uncertainty in the expert reference (e.g., bootstrap CIs, inter-rater agreement) and show that the Table 1 correlations are robust; otherwise the central validity claim is not fully supported.
- [Section 4.2 / Table 2] The objective-metric comparison uses Prolific scores as the subjective reference instead of the expert scores, and it adds Opus 9 kbps and Webex AI Codec 1 kbps conditions that were validated only by an informal listening test. Because the conclusion that SCOREQ outperforms PESQ/POLQA/ViSQOL across architectures rests on this merged dataset, the switch of reference and the undocumented validation are load-bearing. The authors should repeat the metric analysis on the internal expert scores for the original four codecs, and/or provide the details and criteria of the informal validation, to show that the metric rankings are not an artifact of the chosen subjective reference or of the added conditions.
- [Section 2.7] The renormalization rule for aggregating sub-experiments sets the reference to 100 and the anchor to the average anchor across sub-experiments, but the paper does not justify why this preserves inter-condition comparability or how sensitive the final scores are to this choice. Since all subsequent correlations and rankings use these renormalized values, the authors should provide a sensitivity analysis (e.g., alternative anchor settings, or separate analyses per sub-experiment) to demonstrate that the reported validity and metric results do not depend on this ad-hoc step.
minor comments (6)
- [Section 3.2] The text in Section 3.2 says correlations were calculated using "per-condition mean scores," but the surrounding text and Table 1 refer to "per-file mean scores"; please clarify the unit of analysis (per file, per condition, or per file-within-condition).
- [Table 2] The text says the table lists Pearson and Spearman correlations, but the table shows only one set of numbers; either add Spearman values or revise the text.
- [Section 4.2 / Table 2] The negative sign of SCOREQ correlations should be explicitly explained in the table caption or text, since a reader may otherwise misinterpret a negative correlation as disagreement; the reversal of the y-axis in Figure 2 is helpful but the table alone is ambiguous.
- [Figure 2] The figure would benefit from a legend or a caption statement clarifying what the blue and red lines and dots represent, and whether the regression line is computed on per-file or per-condition scores.
- [References] The text in Section 2.2 cites ITU-T P.800.1, but reference [17] is listed as P.800; please correct the reference or the citation.
- [Section 1] The claim of being "the first, to our knowledge, dedicated crowdsourced MUSHRA evaluation design" is strong given that references [10]–[13] already describe crowdsourced MUSHRA adaptations; the novelty should be qualified with respect to generative codecs specifically.
Circularity Check
No significant circularity: the study's claims rest on external subjective ratings and standard correlation analysis, not on self-citation or fitted inputs.
full rationale
The paper is an empirical measurement study rather than a derivation. The central claim that crowdsourced MUSHRA tests can serve as reliable and repeatable alternatives to expert lab tests is supported by comparing crowdsourced ratings with an internal expert MUSHRA test (Section 3.1, Table 1); the expert scores are an external benchmark, not a function of the crowdsourced scores. The repeatability claim is tested by correlating two independent crowdsourced runs (Prolific 1 vs 2, MTurk 1 vs 2). The objective-metric evaluation (Section 4.2) regresses six metrics against Prolific subjective scores; this is a standard correlation analysis, and the decision to use Prolific rather than internal scores is an explicit, reported data-selection choice ('Based on the strong correlation of Prolific tests with our internal results... we decided to use the Prolific scores'), not an equation that defines the outcome in terms of the inputs. The informal listening test used to add two codec conditions is a supplementary validation check, not a fitted parameter renamed as a prediction. The only same-author citation ([8], Lechler and Wojcicki) is used as an example of a prior crowdsourced speech-test adaptation and is not load-bearing for any conclusion. No quantity is defined in terms of the result it is used to predict, and no self-citation is invoked to forbid alternatives. Therefore the derivation chain is self-contained with respect to circularity.
Assumptions & free parameters
free parameters (5)
- Listener disqualification threshold =
20% failures per block
- Participant quality filters =
97% success rate and 100+ completed tasks on both platforms; Prolific additionally required English as first language…
- Maximum conditions per crowdsourced test =
6
- Renormalization of merged experiments =
Reference set to 100 and anchor set to the average of anchors across tests
- Hearing test pass threshold =
Not reported exactly
assumptions (4)
- domain assumption Expert internal MUSHRA scores with 4-6 votes per file are a stable ground truth for validity and metric comparisons.
- domain assumption Per-condition mean scores justify Pearson and Spearman correlations without modeling unequal per-condition listener counts or variance.
- ad hoc to paper Renormalizing scores from different sub-experiments by fixing reference to 100 and anchor to the average anchor preserves inter-condition comparability.
- ad hoc to paper The merged subjective scores including Opus 9 kbps and Webex AI Codec 1 kbps remain valid after informal listening validation.
Cite this review
Pith. "Pith review of Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods." pith.science (2026). https://pith.science/paper/Q5VXIRNW
@misc{pith2026250600950,
author = {Pith},
title = {Pith review of: Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q5VXIRNW}},
note = {Machine review of arXiv:2506.00950}
}
read the original abstract
The MUSHRA framework is widely used for detecting subtle audio quality differences but traditionally relies on expert listeners in controlled environments, making it costly and impractical for model development. As a result, objective metrics are often used during development, with expert evaluations conducted later. While effective for traditional DSP codecs, these metrics often fail to reliably evaluate generative models. This paper proposes adaptations for conducting MUSHRA tests with non-expert, crowdsourced listeners, focusing on generative speech codecs. We validate our approach by comparing results from MTurk and Prolific crowdsourcing platforms with expert listener data, assessing test-retest reliability and alignment. Additionally, we evaluate six objective metrics, showing that traditional metrics undervalue generative models. Our findings reveal platform-specific biases and emphasize codec-aware metrics, offering guidance for scalable perceptual testing of speech codecs.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction Subjective perceptual studies are the gold standard for evalu- ating human perception of audio quality and related research tasks. Conducting these studies requires significant time and cost, often relying on specialized labs or resorting to small- sample internal listening tests with expert listeners. As a result, evaluation using objective ...
-
[2]
Crowdsourced MUSHRA testing Traditional MUSHRA testing has always faced difficulties and high costs. Setting up controlled listening environments, re- cruiting trained participants, and managing equipment can be expensive and time-consuming. Moreover, ensuring consistent and reliable results can be challenging due to variations in lis- tener focus and fat...
work page Pith review arXiv 2025
- [3]
-
[4]
All ratings for the non-anchor systems are identical, i.e. Var(r1, r2, . . . , rK ) = 0, where r1, . . . , rK are the scores for the K conditions under test (excluding the anchor but in- cluding the hidden reference), and Var(·) denotes variance. If 20% of the questions in a test block is less than 1, we take 1 as the threshold. Score-level outlier remova...
work page 2025
-
[5]
Experimental setup We set two main goals for evaluating our crowdsourced MUSHRA listening test
Experiments 3.1. Experimental setup We set two main goals for evaluating our crowdsourced MUSHRA listening test. First, it should correlate well with an internal listening test conducted by expert listeners under con- trolled conditions (i.e., validity). Second, repeating the same test should yield the same results (i.e., repeatability). In this sec- tion...
-
[6]
Results and discussion 4.1. Repeatability and Validity From Figure 1, we see that both Prolific and MTurk listeners produce the same overall ranking of codecs as Internal listeners (i.e., “ground truth”). However, Prolific’s scores more closely trace the Internal curve, with closer absolute means and nar- rower confidence intervals, implying better alignm...
work page 2025
-
[7]
Conclusion Our study demonstrates that crowdsourced MUSHRA tests—when designed with appropriate screening and pro- cedural adjustments—can serve as reliable and repeatable alternatives to expert lab tests, even for generative speech codecs. We show that Prolific yields closer absolute alignment with expert ratings than MTurk, though both maintain reliable...
work page 2025
-
[8]
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221) , vol. 2. Salt Lake City, UT, USA: IEEE, 2001, pp. 749–752. [Online...
work page 2001
Show all 32 references
-
[9]
Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part I–Temporal Alignment,
J. G. Beerends, M. Obermann, R. Ullmann, J. Pomy, and M. Keyhl, “Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part I–Temporal Alignment,”J. Au- dio Eng. Soc., vol. 61, no. 6, 2013
2013
-
[10]
Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part II–Perceptual Model,
——, “Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part II–Perceptual Model,”J. Audio Eng. Soc., vol. 61, no. 6, 2013
2013
-
[11]
Non-intrusive Speech Quality Assess- ment for Super-wideband Speech Communication Networks,
G. Mittag and S. Moller, “Non-intrusive Speech Quality Assess- ment for Super-wideband Speech Communication Networks,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Brighton, United Kingdom: IEEE, May 2019, pp. 7125–7...
2019
-
[12]
NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,
G. Mittag, B. Naderi, A. Chehadi, and S. M ¨oller, “NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,” in Interspeech 2021 . ISCA, Aug. 2021, pp. 2127–2131. [On- line]. Available: https://www.isca-archive.org/inte...
2021
-
[13]
CROWD- MOS: An approach for crowdsourcing mean opinion score studies,
F. Ribeiro, D. Florencio, C. Zhang, and M. Seltzer, “CROWD- MOS: An approach for crowdsourcing mean opinion score studies,” in 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Prague, Czech Re- public: IEEE, May 2011, pp. 2416–2419. [Onl...
2011
-
[14]
An Open Source Implementa- tion of ITU-T Recommendation P.808 with Validation,
B. Naderi and R. Cutler, “An Open Source Implementa- tion of ITU-T Recommendation P.808 with Validation,” in Interspeech 2020 . ISCA, Oct. 2020, pp. 2862–2866. [On- line]. Available: https://www.isca-archive.org/interspeech 2020/ naderi20 interspeech.html
2020
-
[15]
Crowdsourced Multilingual Speech Intelligibility Testing,
L. Lechler and K. Wojcicki, “Crowdsourced Multilingual Speech Intelligibility Testing,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Seoul, Korea, Republic of: IEEE, Apr. 2024, pp. 1441–1445. [Online]. Available: htt...
2024
-
[16]
Method for the subjective assessment of intermediate quality level of coding systems,
ITU, “Method for the subjective assessment of intermediate quality level of coding systems,” Geneva, Switzerland, 2015. [Online]. Available: https://www.itu.int/rec/R-REC-BS.1534/en
2015
-
[17]
Audio quality evaluation by experienced and inexperienced listeners,
N. Schinkel-Bielefeld, N. Lotze, and F. Nagel, “Audio quality evaluation by experienced and inexperienced listeners,” Montreal, Canada, 2013, pp. 060 016–060 016. [Online]. Available: https: //pubs.aip.org/asa/poma/article/808581
2013
-
[18]
Fast and easy crowdsourced perceptual audio evaluation,
M. Cartwright, B. Pardo, G. J. Mysore, and M. Hoffman, “Fast and easy crowdsourced perceptual audio evaluation,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Shanghai: IEEE, Mar. 2016, pp. 619–
2016
-
[19]
High fidelity neural audio compression,
A. D ´efossez, J. Copet, G. Synnaeve, and Y . Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research, 2023, featured Certification, Reproducibility Certification. [Online]. Available: https://openreview.net/forum? id=ivCd8z8zR2
2023
-
[20]
webMUSHRA — A Comprehensive Framework for Web-based Listening Tests,
M. Schoeffler, S. Bartoschek, F.-R. St ¨oter, M. Roess, S. Westphal, B. Edler, and J. Herre, “webMUSHRA — A Comprehensive Framework for Web-based Listening Tests,” Journal of Open Research Software , vol. 6, no. 1, p. 8, Feb. 2018. [Online]. Available: https://openresearchsoft...
2018
-
[21]
Reproducible sub- jective evaluation,
M. Morrison, B. Tang, G. Tan, and B. Pardo, “Reproducible sub- jective evaluation,”arXiv preprint arXiv:2203.04444, 2022
2022 arXiv
-
[22]
Speech quality evaluation of neural audio codecs,
T. Muller, S. Ragot, L. Gros, P. Philippe, and P. Scalart, “Speech quality evaluation of neural audio codecs,” in Interspeech 2024 . ISCA, Sep. 2024, pp. 1760–1764. [On- line]. Available: https://www.isca-archive.org/interspeech 2024/ muller24c interspeech.html
2024
-
[23]
P.808 : Subjective evaluation of speech quality with a crowdsourcing approach,
ITU, “P.808 : Subjective evaluation of speech quality with a crowdsourcing approach,” 2021. [Online]. Available: https: //www.itu.int/rec/T-REC-P.808/en
2021
-
[24]
Is it harder to perceive coding artifact in foreign language items? – A study with Mandarin Chinese and German speaking listeners,
N. Schinkel-Bielefeld, Z. Jiandong, Q. Yili, A. K. Leschanowsky, and F. Shanshan, “Is it harder to perceive coding artifact in foreign language items? – A study with Mandarin Chinese and German speaking listeners,” Journal of the Audio Engineering Society, no. 9739, May 2017
2017
-
[25]
P.800 : Methods for subjective determination of transmission quality,
ITU, “P.800 : Methods for subjective determination of transmission quality,” 1998. [Online]. Available: https://www.itu. int/rec/T-REC-P.800-199608-I/en
1998
-
[26]
Webex AI Codec: Delivering Next-level Audio Experiences with AI/ML,
A. Dhingra, “Webex AI Codec: Delivering Next-level Audio Experiences with AI/ML,” Mar. 2024. [Online]. Available: https://blog.webex.com/collaboration/hybrid-work/ next-level-audio-with-webex-ai-codec/
2024
-
[28]
V oice coding with opus,
K. V os, K. V . Sørensen, S. S. Jensen, and J.-M. Valin, “V oice coding with opus,” inAudio Engineering Society Convention 135. Audio Engineering Society, 2013
2013
-
[29]
Overview of the EVS codec architecture,
M. Dietz, M. Multrus, V . Eksler, V . Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache, Y . Kamamoto, K. Kikuiri, S. Ragot, J. Faure, H. Ehara, V . Rajendran, V . Atti, H. Sung, E. Oh, H. Yuan, and C. Zhu, “Overview of the EVS codec architecture...
2015
-
[30]
ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric,
M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric,” Apr. 2020, arXiv:2004.09584 [eess]. [Online]. Available: http://arxiv.org/ abs/2004.09584
2020 arXiv
-
[31]
Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,
C. K. Reddy, V . Gopal, and R. Cutler, “Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6493–6497
2021
-
[32]
SCOREQ: Speech quality assessment with contrastive regression,
A. Ragano, J. Skoglund, and A. Hines, “SCOREQ: Speech quality assessment with contrastive regression,” in Advances in Neural Information Processing Systems , A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, In...
2024
-
[623]
Available: http://ieeexplore.ieee.org/document/ 7471749/
[Online]. Available: http://ieeexplore.ieee.org/document/ 7471749/
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.