{"id":"3eeb30ab-74fc-49f2-b0d9-5f2c197da5d3","arxiv_id":"2506.00950","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Crowdsourced MUSHRA tests on Prolific and MTurk reproduce expert codec rankings for generative speech codecs, and SCOREQ tracks subjective quality more consistently than PESQ, POLQA, or ViSQOL.","lead":"This paper tests whether MUSHRA audio quality tests can be run with non-expert crowd workers instead of expert listeners in a lab. It finds that both MTurk and Prolific reproduce expert codec rankings, with Prolific closer in absolute scores, and that the SCOREQ metric tracks subjective quality better than traditional metrics for neural codecs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validity rests on a 4–6-vote expert MUSHRA reference; if that reference is unstable, the platform-alignment and metric conclusions lose their anchor.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the small internal expert sample is the sole external validity anchor for the crowdsourced protocol, and the later substitution of Prolific scores for expert scores in the objective-metric analysis compounds the risk. My independent reading confirms that the paper's headline claim of a reliable and repeatable alternative to expert lab tests depends on the expert MUSHRA scores being a stable ground truth. The paper does provide supporting evidence: two independent crowdsourced runs on two platforms, high correlations, repeated tests months apart, and open-source tooling. However, the correlations are computed against a 4–6-vote-per-file expert reference, and no uncertainty quantification is reported. The additional conditions added for the objective-metric comparison are validated only informally, so the metric conclusions rest on Prolific subjective scores rather than the expert reference. This does not warrant rejection, because the observed effects are large and internally consistent, but it does warrant a conditional verdict: the central claim is credible yet not fully established without a stronger expert reference or an explicit reliability analysis of the existing expert scores. Therefore no change to the reader's CONDITIONAL verdict is needed.","tokens_in":7926,"tokens_out":4903,"duration_ms":52991,"concrete_test":"Re-run the internal expert MUSHRA on the same 40 files and original four codec conditions with a panel of at least 20 listeners following BS.1534-3, then recompute Table 1 and Table 2 using these expert scores for the shared conditions. If the crowd-expert correlations drop below about 0.9, or if SCOREQ's advantage over PESQ/POLQA disappears or reverses, the central claim and the metric ranking are not established. As a cheaper first step, compute split-half reliability of the existing 4–6-vote expert means and report bootstrap confidence intervals for the Table 1 correlations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the internal expert MUSHRA scores are a trustworthy, stable ground truth. Section 3.1 reports only 4–6 votes per file for the expert test; ITU-R BS.1534-3 recommends a much larger expert panel, and per-file means from 4–6 raters are highly sensitive to individual scale use. Table 1 then reports Pearson/Spearman correlations of these per-file means as validity evidence, with no confidence intervals or significance tests. The same expert scores are also the basis for declaring Prolific 'closer absolute alignment,' and Section 4.2 goes further by substituting Prolific scores for the expert scores in the objective-metric comparison, adding two conditions validated only by an informal listening test. If the small expert reference is noisy or biased, the crowd-to-expert correlations could be inflated or artifacts of shared response patterns, and the SCOREQ/PESQ/POLQA conclusions would not transfer to genuine expert-lab judgments. The claim is plausible, but its evidential foundation is a single small-sample comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes adaptations to the MUSHRA protocol for crowdsourced, non-expert listeners, focusing on generative speech codecs. It compares results from MTurk and Prolific against a small internal expert listening test, reports test–retest correlations, and evaluates six objective metrics against subjective scores. The main claims are that crowdsourced MUSHRA tests can reliably replace expert lab tests, that Prolific shows closer absolute alignment with expert scores while MTurk shows a ceiling effect, and that traditional objective metrics (PESQ, POLQA, ViSQOL) undervalue DNN-based codecs while SCOREQ is more architecture-consistent.","tokens_in":8168,"tokens_out":3603,"duration_ms":37110,"significance":"If the findings are robust, this is a practical contribution: it provides an open-source tool, a direct platform comparison, and evidence that architecture-aware objective metrics are needed for generative codecs. The test–retest correlations and the large per-file samples on the crowdsourcing platforms are concrete assets. The main caveat is that the validity claims hinge on a very small expert reference (4–6 votes per file) and on a post-hoc substitution of Prolific scores for expert scores in the objective-metric analysis; these weaken the force of the conclusions as currently presented.","major_comments":[{"comment":"The internal expert test used only 4–6 votes per file, yet it is treated as the ground truth for validating both platforms. Table 1 reports Pearson/Spearman correlations with this reference but gives no confidence intervals, significance tests, or per-file variance of the expert scores. With such a small rater pool, per-file means are highly sensitive to individual scale use, and the reported correlations could be inflated by shared response patterns rather than true alignment. The authors should quantify the uncertainty in the expert reference (e.g., bootstrap CIs, inter-rater agreement) and show that the Table 1 correlations are robust; otherwise the central validity claim is not fully supported.","section":"Section 3.1 / Table 1"},{"comment":"The objective-metric comparison uses Prolific scores as the subjective reference instead of the expert scores, and it adds Opus 9 kbps and Webex AI Codec 1 kbps conditions that were validated only by an informal listening test. Because the conclusion that SCOREQ outperforms PESQ/POLQA/ViSQOL across architectures rests on this merged dataset, the switch of reference and the undocumented validation are load-bearing. The authors should repeat the metric analysis on the internal expert scores for the original four codecs, and/or provide the details and criteria of the informal validation, to show that the metric rankings are not an artifact of the chosen subjective reference or of the added conditions.","section":"Section 4.2 / Table 2"},{"comment":"The renormalization rule for aggregating sub-experiments sets the reference to 100 and the anchor to the average anchor across sub-experiments, but the paper does not justify why this preserves inter-condition comparability or how sensitive the final scores are to this choice. Since all subsequent correlations and rankings use these renormalized values, the authors should provide a sensitivity analysis (e.g., alternative anchor settings, or separate analyses per sub-experiment) to demonstrate that the reported validity and metric results do not depend on this ad-hoc step.","section":"Section 2.7"}],"minor_comments":[{"comment":"The text in Section 3.2 says correlations were calculated using \"per-condition mean scores,\" but the surrounding text and Table 1 refer to \"per-file mean scores\"; please clarify the unit of analysis (per file, per condition, or per file-within-condition).","section":"Section 3.2"},{"comment":"The text says the table lists Pearson and Spearman correlations, but the table shows only one set of numbers; either add Spearman values or revise the text.","section":"Table 2"},{"comment":"The negative sign of SCOREQ correlations should be explicitly explained in the table caption or text, since a reader may otherwise misinterpret a negative correlation as disagreement; the reversal of the y-axis in Figure 2 is helpful but the table alone is ambiguous.","section":"Section 4.2 / Table 2"},{"comment":"The figure would benefit from a legend or a caption statement clarifying what the blue and red lines and dots represent, and whether the regression line is computed on per-file or per-condition scores.","section":"Figure 2"},{"comment":"The text in Section 2.2 cites ITU-T P.800.1, but reference [17] is listed as P.800; please correct the reference or the citation.","section":"References"},{"comment":"The claim of being \"the first, to our knowledge, dedicated crowdsourced MUSHRA evaluation design\" is strong given that references [10]–[13] already describe crowdsourced MUSHRA adaptations; the novelty should be qualified with respect to generative codecs specifically.","section":"Section 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically plausible but the evidential foundation is thinner than the conclusions suggest. The small expert reference and the post-hoc use of Prolific scores are fixable with additional analyses, so I am not recommending rejection. The fit with the journal's scope is fine. One editorial concern: the paper states it is already accepted for INTERSPEECH 2025; if this manuscript is being considered for a journal version, the authors should be asked to extend the analysis beyond the conference paper, not just resubmit the same text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, narrow empirical paper on adapting MUSHRA for crowdsourced testing of generative speech codecs. The new bits are the direct Prolific-versus-MTurk comparison with test-retest data and the codec-aware objective metric correlations, including SCOREQ.\n\nWhat it does well: the protocol adaptations are sensible and reported in enough detail to be useful—screening, qualification, post-screening, anchor choice, renormalization. The platform comparison is a real data point: both platforms reproduce expert rankings, Prolific absolute scores track better, MTurk shows a ceiling effect. Test-retest correlations are high. That is credible and worth having.\n\nSoft spots: the expert reference is the load-bearing wall, and it is 4–6 votes per file with no confidence intervals or significance tests. That in itself doesn't sink the ranking claim—the crowd-to-expert correlations are strong across 40 files, and pure noise on the expert side would attenuate correlations rather than inflate them—but it does limit how strongly you can claim 'absolute alignment.' The objective metric comparison is a step further from the ground truth: they substitute Prolific scores for the expert scores and add two codec conditions validated only informally. The ranking conclusions are robust; the metric conclusions (SCOREQ good, PESQ/POLQA undervalue DNN codecs) are suggestive but provisional. Raw data isn't released, which is a shame for a paper whose main asset is its empirical comparison.\n\nWho for: anyone doing perceptual evaluation of neural codecs, or deciding between MTurk and Prolific for MUSHRA-like tasks. It deserves a serious referee, not a desk reject. My recommendation: accept with requests for confidence intervals, raw per-file scores as supplementary material, and softer wording around the absolute-alignment and SCOREQ claims.","headline":"A useful, narrow empirical study showing crowdsourced MUSHRA can reproduce expert codec rankings for generative speech codecs, with Prolific closer in absolute terms than MTurk—but the expert ground truth is only 4–6 votes per file and the objective-metric analysis leans on substituted Prolific scores.","tokens_in":8639,"tokens_out":2496,"would_cite":true,"duration_ms":25975,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Crowdsourced MUSHRA tests can stand in for expert lab tests when screening and quality-control steps are added.","keywords":["crowdsourced MUSHRA","speech codec evaluation","generative speech codecs","subjective audio quality testing","objective speech quality metrics","SCOREQ","Prolific","MTurk"],"falsifier":"Run the same protocol on a held-out set of at least six generative and DSP codecs with both expert and Prolific listeners; if per-condition expert-to-Prolific correlation drops below about 0.9, or if SCOREQ's correlation with subjective scores no longer stays consistent between DNN and DSP codecs when computed per file, the central claims fail.","tokens_in":7711,"feed_emoji":"🎧","tokens_out":9112,"duration_ms":80564,"temperature":0.7,"pith_summary":"The paper tries to establish that the MUSHRA (multiple stimuli with hidden reference and anchor) listening test, normally run with trained experts in controlled labs, can be moved to crowdsourced, non-expert listeners without losing its ability to rank modern generative speech codecs. To do this, the authors add a screening and quality-control pipeline to the standard protocol and validate it by running the same test internally with experts and externally on two crowdsourcing platforms. They report that both platforms reproduce the expert codec ranking with high test-retest repeatability, that Prolific's absolute scores closely match the expert scores, and that MTurk shows a ceiling effect in absolute terms. The paper also argues that conventional objective metrics such as PESQ, POLQA, and ViSQOL underestimate DNN-based codecs, while SCOREQ tracks subjective scores consistently across codec architectures. If the claim holds, perceptual evaluation of speech codecs becomes cheaper and more frequent during model development.","feed_headline":"Crowdsourced MUSHRA passes the expert-listener test for neural codecs","feed_subtitle":"Prolific listeners reproduce expert scores; MTurk reproduces rankings with a ceiling effect.","key_machinery":"The carrying mechanism is a screening-and-renormalization pipeline built around the MUSHRA protocol. It includes platform quality filters (97% success rate and 100 completed tasks), a digits-in-noise hearing test, a training MUSHRA question with three attempts, real-time score screening that discards responses deviating beyond thresholds, automatic test partitioning into smaller blocks, listener-level disqualification when the anchor is rated above the hidden reference or when all non-anchor scores are identical, inter-quartile-range outlier removal, and renormalization of results across sub-experiments so the hidden reference sits at 100 and the anchor matches across tests. The anchor is Opus 6 kbps, chosen because its coding artifacts resemble those of the test codecs, which the authors found easier for non-experts to judge than a low-pass-filtered anchor. This pipeline is what lets non-expert listeners produce expert-like rankings.","core_discovery":"The authors' central claim is that a crowdsourced MUSHRA protocol, once adapted with participant screening, a hearing test, training, real-time score screening, test partitioning, and post-screening outlier rules, yields results that are both repeatable and aligned with an internal expert MUSHRA test, even for generative speech codecs. With 40 English speech files and four codecs plus an Opus 6 kbps anchor, the per-file mean scores from Prolific correlate at 0.95 (Pearson) with the expert scores in two independent runs, MTurk at 0.89-0.90, and repeated crowdsourced runs correlate at 0.98-0.99 with each other. Prolific's absolute scores stay close to the expert curve, while MTurk reproduces rankings but drifts in absolute value. On the objective side, the paper finds that PESQ, POLQA, and ViSQOL correlate better for DSP codecs and systematically undervalue DNN codecs, while SCOREQ's correlation with subjective scores is stable across both architectures (-0.80 overall, -0.79 for both groups, with lower values meaning better quality), and NISQA and DNSMOS align poorly overall.","pith_inferences":["A logical testable extension is to apply the same screening pipeline to other relative-judgment audio tasks, such as preference or intelligibility tests; the paper does not claim this, but nothing in its mechanism is task-specific.","The objective-metric comparison is computed at condition level over a small codec set; per-file analysis with a wider set of generative codecs would be needed to confirm that SCOREQ's advantage generalizes.","The platform difference may stem from different participant pools as much as from platform mechanics; a crossover design with the same listeners on both platforms would separate those effects.","Because the anchor impairment was matched to coding artifacts, the protocol's conclusions may be specific to codec-like degradation; other artifact types may need their own anchor validation."],"forward_implications":["Crowdsourced MUSHRA can substitute for expert lab tests during codec development, making frequent perceptual checks practical.","Prolific is the better platform when absolute score alignment matters; MTurk still yields valid rankings, but its ceiling effect must be accounted for.","PESQ, POLQA, and ViSQOL should not be relied on to rank neural codecs, because they systematically undervalue DNN-based systems.","SCOREQ, as a reference-based contrastive metric, is a more architecture-consistent choice among the six metrics tested.","MUSHRA test data can serve as reference ground truth for evaluating new objective metrics."],"supporting_citations":[{"why":"Supplies the standard MUSHRA testing procedure that the authors adapt for crowdsourced, non-expert listeners.","marker":"[9]"},{"why":"Provides the crowdsourcing guidelines for speech quality testing that the qualification and screening flow builds on.","marker":"[15]"},{"why":"Supplies the open-source implementation and hearing-test toolkit used for the digits-in-noise qualification step.","marker":"[7]"},{"why":"Supplies the base web-based listening-test framework that the authors extend with real-time score screening and automatic test partitioning.","marker":"[12]"},{"why":"Provides prior evidence that expert and non-expert listeners produce largely consistent relative rankings, which motivates the crowdsourcing adaptation.","marker":"[10]"},{"why":"Defines EnCodec, one of the two generative speech codecs used as a test condition in the MUSHRA experiments.","marker":"[19]"},{"why":"Defines SCOREQ, the objective metric the paper finds most consistent across both DNN- and DSP-based codecs.","marker":"[24]"},{"why":"Defines PESQ, one of the traditional intrusive metrics whose underestimation of DNN codecs is a central finding.","marker":"[1]"},{"why":"Defines POLQA, another traditional intrusive metric compared against subjective scores and found to undervalue generative codecs.","marker":"[2]"},{"why":"Defines ViSQOL, the third traditional objective metric whose correlation with subjective scores is weaker for DNN codecs.","marker":"[22]"}],"fun_headline_variants":["Crowdsourced MUSHRA passes for neural codecs","Prolific beats MTurk in crowd MUSHRA for speech","Objective metrics misfire on generative speech codecs","Crowd MUSHRA: Prolific aligns, MTurk ranks only","SCOREQ stays stable on generative codecs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All conclusions rest on treating the internal expert MUSHRA scores, just four to six votes per audio file, as a stable ground truth, and on accepting the later substitution of Prolific scores for those expert scores in the objective-metric analysis, with two added codec conditions validated only by an informal listening test.","fun_headline_variants_meta":{"raw":{"variants":["Crowdsourced MUSHRA passes for neural codecs","Prolific beats MTurk in crowd MUSHRA for speech","Objective metrics misfire on generative speech codecs","Crowd MUSHRA: Prolific aligns, MTurk ranks only","SCOREQ stays stable on generative codecs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1263,"prompt_tokens":946,"completion_tokens":317,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":232}},"tokens_in":562,"tokens_out":317,"duration_ms":3591,"temperature":1.0,"reasoning_tokens":232,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:53:43.190692+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same protocol on a held-out set of at least six generative and DSP codecs with both expert and Prolific listeners; if per-condition expert-to-Prolific correlation drops below about 0.9, or if SCOREQ's correlation with subjective scores no longer stays consistent between DNN and DSP codecs when computed per file, the central claims fail.","supporting_citations":[{"cited_title":"Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part I–Temporal Alignment,","cited_arxiv_id":null,"evidence_quote":"Supplies the standard MUSHRA testing procedure that the authors adapt for crowdsourced, non-expert listeners."},{"cited_title":"Crowdsourced Multilingual Speech Intelligibility Testing,","cited_arxiv_id":null,"evidence_quote":"Provides the crowdsourcing guidelines for speech quality testing that the qualification and screening flow builds on."},{"cited_title":"We show that Prolific yields closer absolute alignment with expert ratings than MTurk, though both maintain reliable relative rankings","cited_arxiv_id":null,"evidence_quote":"Supplies the open-source implementation and hearing-test toolkit used for the digits-in-noise qualification step."},{"cited_title":"NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,","cited_arxiv_id":null,"evidence_quote":"Supplies the base web-based listening-test framework that the authors extend with real-time score screening and automatic test partitioning."},{"cited_title":"Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part II–Perceptual Model,","cited_arxiv_id":null,"evidence_quote":"Provides prior evidence that expert and non-expert listeners produce largely consistent relative rankings, which motivates the crowdsourcing adaptation."},{"cited_title":"High fidelity neural audio compression,","cited_arxiv_id":null,"evidence_quote":"Defines EnCodec, one of the two generative speech codecs used as a test condition in the MUSHRA experiments."},{"cited_title":"Is it harder to perceive coding artifact in foreign language items? – A study with Mandarin Chinese and German speaking listeners,","cited_arxiv_id":null,"evidence_quote":"Defines SCOREQ, the objective metric the paper finds most consistent across both DNN- and DSP-based codecs."},{"cited_title":"Conducting these studies requires significant time and cost, often relying on specialized labs or resorting to small- sample internal listening tests with expert listeners","cited_arxiv_id":null,"evidence_quote":"Defines PESQ, one of the traditional intrusive metrics whose underestimation of DNN codecs is a central finding."},{"cited_title":"Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods","cited_arxiv_id":"2506.00950","evidence_quote":"Defines POLQA, another traditional intrusive metric compared against subjective scores and found to undervalue generative codecs."},{"cited_title":"Speech quality evaluation of neural audio codecs,","cited_arxiv_id":null,"evidence_quote":"Defines ViSQOL, the third traditional objective metric whose correlation with subjective scores is weaker for DNN codecs."}],"review_version":1}