Pith. sign in

REVIEW 3 major objections 6 minor 32 references

Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Crowdsourced MUSHRA tests can stand in for expert lab tests when screening and quality-control steps are added.

desk verdict A useful, narrow empirical study showing crowdsourced MUSHRA can reproduce expert codec rankings for generative speech codecs, with Prolific closer in absolute terms than MTurk—but the expert ground truth is only 4–6 votes per file and the objective-metric analysis leans on substituted Prolific scores. read the letter →

arxiv 2506.00950 v1 pith:Q5VXIRNW submitted 2025-06-01 eess.AS cs.SD

classification eess.AScs.SD
keywords crowdsourcedMUSHRAspeechcodecevaluationgenerativecodecssubjectiveaudioqualitytestingobjectivemetricsSCOREQProlificMTurk
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the MUSHRA (multiple stimuli with hidden reference and anchor) listening test, normally run with trained experts in controlled labs, can be moved to crowdsourced, non-expert listeners without losing its ability to rank modern generative speech codecs. To do this, the authors add a screening and quality-control pipeline to the standard protocol and validate it by running the same test internally with experts and externally on two crowdsourcing platforms. They report that both platforms reproduce the expert codec ranking with high test-retest repeatability, that Prolific's absolute scores closely match the expert scores, and that MTurk shows a ceiling effect in absolute terms. The paper also argues that conventional objective metrics such as PESQ, POLQA, and ViSQOL underestimate DNN-based codecs, while SCOREQ tracks subjective scores consistently across codec architectures. If the claim holds, perceptual evaluation of speech codecs becomes cheaper and more frequent during model development.

What carries the argument

The carrying mechanism is a screening-and-renormalization pipeline built around the MUSHRA protocol. It includes platform quality filters (97% success rate and 100 completed tasks), a digits-in-noise hearing test, a training MUSHRA question with three attempts, real-time score screening that discards responses deviating beyond thresholds, automatic test partitioning into smaller blocks, listener-level disqualification when the anchor is rated above the hidden reference or when all non-anchor scores are identical, inter-quartile-range outlier removal, and renormalization of results across sub-experiments so the hidden reference sits at 100 and the anchor matches across tests. The anchor is Opus 6 kbps, chosen because its coding artifacts resemble those of the test codecs, which the authors found easier for non-experts to judge than a low-pass-filtered anchor. This pipeline is what lets non-expert listeners produce expert-like rankings.

What would settle it

Run the same protocol on a held-out set of at least six generative and DSP codecs with both expert and Prolific listeners; if per-condition expert-to-Prolific correlation drops below about 0.9, or if SCOREQ's correlation with subjective scores no longer stays consistent between DNN and DSP codecs when computed per file, the central claims fail.

Watch

Extended reading notes

Core claim

The authors' central claim is that a crowdsourced MUSHRA protocol, once adapted with participant screening, a hearing test, training, real-time score screening, test partitioning, and post-screening outlier rules, yields results that are both repeatable and aligned with an internal expert MUSHRA test, even for generative speech codecs. With 40 English speech files and four codecs plus an Opus 6 kbps anchor, the per-file mean scores from Prolific correlate at 0.95 (Pearson) with the expert scores in two independent runs, MTurk at 0.89-0.90, and repeated crowdsourced runs correlate at 0.98-0.99 with each other. Prolific's absolute scores stay close to the expert curve, while MTurk reproduces rankings but drifts in absolute value. On the objective side, the paper finds that PESQ, POLQA, and ViSQOL correlate better for DSP codecs and systematically undervalue DNN codecs, while SCOREQ's correlation with subjective scores is stable across both architectures (-0.80 overall, -0.79 for both groups, with lower values meaning better quality), and NISQA and DNSMOS align poorly overall.

Load-bearing premise

All conclusions rest on treating the internal expert MUSHRA scores, just four to six votes per audio file, as a stable ground truth, and on accepting the later substitution of Prolific scores for those expert scores in the objective-metric analysis, with two added codec conditions validated only by an informal listening test.

Editorial extensions

If this is right

  • Crowdsourced MUSHRA can substitute for expert lab tests during codec development, making frequent perceptual checks practical.
  • Prolific is the better platform when absolute score alignment matters; MTurk still yields valid rankings, but its ceiling effect must be accounted for.
  • PESQ, POLQA, and ViSQOL should not be relied on to rank neural codecs, because they systematically undervalue DNN-based systems.
  • SCOREQ, as a reference-based contrastive metric, is a more architecture-consistent choice among the six metrics tested.
  • MUSHRA test data can serve as reference ground truth for evaluating new objective metrics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A logical testable extension is to apply the same screening pipeline to other relative-judgment audio tasks, such as preference or intelligibility tests; the paper does not claim this, but nothing in its mechanism is task-specific.
  • The objective-metric comparison is computed at condition level over a small codec set; per-file analysis with a wider set of generative codecs would be needed to confirm that SCOREQ's advantage generalizes.
  • The platform difference may stem from different participant pools as much as from platform mechanics; a crossover design with the same listeners on both platforms would separate those effects.
  • Because the anchor impairment was matched to coding artifacts, the protocol's conclusions may be specific to codec-like degradation; other artifact types may need their own anchor validation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes adaptations to the MUSHRA protocol for crowdsourced, non-expert listeners, focusing on generative speech codecs. It compares results from MTurk and Prolific against a small internal expert listening test, reports test–retest correlations, and evaluates six objective metrics against subjective scores. The main claims are that crowdsourced MUSHRA tests can reliably replace expert lab tests, that Prolific shows closer absolute alignment with expert scores while MTurk shows a ceiling effect, and that traditional objective metrics (PESQ, POLQA, ViSQOL) undervalue DNN-based codecs while SCOREQ is more architecture-consistent.

Significance. If the findings are robust, this is a practical contribution: it provides an open-source tool, a direct platform comparison, and evidence that architecture-aware objective metrics are needed for generative codecs. The test–retest correlations and the large per-file samples on the crowdsourcing platforms are concrete assets. The main caveat is that the validity claims hinge on a very small expert reference (4–6 votes per file) and on a post-hoc substitution of Prolific scores for expert scores in the objective-metric analysis; these weaken the force of the conclusions as currently presented.

major comments (3)
  1. [Section 3.1 / Table 1] The internal expert test used only 4–6 votes per file, yet it is treated as the ground truth for validating both platforms. Table 1 reports Pearson/Spearman correlations with this reference but gives no confidence intervals, significance tests, or per-file variance of the expert scores. With such a small rater pool, per-file means are highly sensitive to individual scale use, and the reported correlations could be inflated by shared response patterns rather than true alignment. The authors should quantify the uncertainty in the expert reference (e.g., bootstrap CIs, inter-rater agreement) and show that the Table 1 correlations are robust; otherwise the central validity claim is not fully supported.
  2. [Section 4.2 / Table 2] The objective-metric comparison uses Prolific scores as the subjective reference instead of the expert scores, and it adds Opus 9 kbps and Webex AI Codec 1 kbps conditions that were validated only by an informal listening test. Because the conclusion that SCOREQ outperforms PESQ/POLQA/ViSQOL across architectures rests on this merged dataset, the switch of reference and the undocumented validation are load-bearing. The authors should repeat the metric analysis on the internal expert scores for the original four codecs, and/or provide the details and criteria of the informal validation, to show that the metric rankings are not an artifact of the chosen subjective reference or of the added conditions.
  3. [Section 2.7] The renormalization rule for aggregating sub-experiments sets the reference to 100 and the anchor to the average anchor across sub-experiments, but the paper does not justify why this preserves inter-condition comparability or how sensitive the final scores are to this choice. Since all subsequent correlations and rankings use these renormalized values, the authors should provide a sensitivity analysis (e.g., alternative anchor settings, or separate analyses per sub-experiment) to demonstrate that the reported validity and metric results do not depend on this ad-hoc step.
minor comments (6)
  1. [Section 3.2] The text in Section 3.2 says correlations were calculated using "per-condition mean scores," but the surrounding text and Table 1 refer to "per-file mean scores"; please clarify the unit of analysis (per file, per condition, or per file-within-condition).
  2. [Table 2] The text says the table lists Pearson and Spearman correlations, but the table shows only one set of numbers; either add Spearman values or revise the text.
  3. [Section 4.2 / Table 2] The negative sign of SCOREQ correlations should be explicitly explained in the table caption or text, since a reader may otherwise misinterpret a negative correlation as disagreement; the reversal of the y-axis in Figure 2 is helpful but the table alone is ambiguous.
  4. [Figure 2] The figure would benefit from a legend or a caption statement clarifying what the blue and red lines and dots represent, and whether the regression line is computed on per-file or per-condition scores.
  5. [References] The text in Section 2.2 cites ITU-T P.800.1, but reference [17] is listed as P.800; please correct the reference or the citation.
  6. [Section 1] The claim of being "the first, to our knowledge, dedicated crowdsourced MUSHRA evaluation design" is strong given that references [10]–[13] already describe crowdsourced MUSHRA adaptations; the novelty should be qualified with respect to generative codecs specifically.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the study's claims rest on external subjective ratings and standard correlation analysis, not on self-citation or fitted inputs.

full rationale

The paper is an empirical measurement study rather than a derivation. The central claim that crowdsourced MUSHRA tests can serve as reliable and repeatable alternatives to expert lab tests is supported by comparing crowdsourced ratings with an internal expert MUSHRA test (Section 3.1, Table 1); the expert scores are an external benchmark, not a function of the crowdsourced scores. The repeatability claim is tested by correlating two independent crowdsourced runs (Prolific 1 vs 2, MTurk 1 vs 2). The objective-metric evaluation (Section 4.2) regresses six metrics against Prolific subjective scores; this is a standard correlation analysis, and the decision to use Prolific rather than internal scores is an explicit, reported data-selection choice ('Based on the strong correlation of Prolific tests with our internal results... we decided to use the Prolific scores'), not an equation that defines the outcome in terms of the inputs. The informal listening test used to add two codec conditions is a supplementary validation check, not a fitted parameter renamed as a prediction. The only same-author citation ([8], Lechler and Wojcicki) is used as an example of a prior crowdsourced speech-test adaptation and is not load-bearing for any conclusion. No quantity is defined in terms of the result it is used to predict, and no self-citation is invoked to forbid alternatives. Therefore the derivation chain is self-contained with respect to circularity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claims rest on hand-selected protocol thresholds and on the assumption that small expert samples and renormalized crowd scores are valid ground truth. No invented entities are introduced.

free parameters (5)
  • Listener disqualification threshold = 20% failures per block
    Chosen by hand in Section 2.4; it determines which listeners are excluded and therefore affects all reported score distributions.
  • Participant quality filters = 97% success rate and 100+ completed tasks on both platforms; Prolific additionally required English as first language…
    Set in Section 2.1; different filters on the two platforms may contribute to the observed platform differences.
  • Maximum conditions per crowdsourced test = 6
    Adopted from experience in Section 2.7; the paper states this limit is not yet scientifically established.
  • Renormalization of merged experiments = Reference set to 100 and anchor set to the average of anchors across tests
    Defined in Section 2.7; this normalization is required to pool results and directly shapes the compared score distributions.
  • Hearing test pass threshold = Not reported exactly
    Mentioned in Section 2.2 as a threshold of correct digits, but the numeric value is not given, leaving replication incomplete.
assumptions (4)
  • domain assumption Expert internal MUSHRA scores with 4-6 votes per file are a stable ground truth for validity and metric comparisons.
    Invoked in Sections 3.1 and 4.1; if the expert scores are noisy, the correlation claims weaken.
  • domain assumption Per-condition mean scores justify Pearson and Spearman correlations without modeling unequal per-condition listener counts or variance.
    Used for Table 1 and Table 2; Prolific had roughly 15 responses per item and MTurk roughly 25, but this imbalance is not modeled.
  • ad hoc to paper Renormalizing scores from different sub-experiments by fixing reference to 100 and anchor to the average anchor preserves inter-condition comparability.
    Defined in Section 2.7; this is introduced for this study and is not an established ITU MUSHRA practice.
  • ad hoc to paper The merged subjective scores including Opus 9 kbps and Webex AI Codec 1 kbps remain valid after informal listening validation.
    Stated in Section 4.2; no formal listening test or published evidence supports this merge.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods." pith.science (2026). https://pith.science/paper/Q5VXIRNW

@misc{pith2026250600950,
  author       = {Pith},
  title        = {Pith review of: Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q5VXIRNW}},
  note         = {Machine review of arXiv:2506.00950}
}
read the original abstract

The MUSHRA framework is widely used for detecting subtle audio quality differences but traditionally relies on expert listeners in controlled environments, making it costly and impractical for model development. As a result, objective metrics are often used during development, with expert evaluations conducted later. While effective for traditional DSP codecs, these metrics often fail to reliably evaluate generative models. This paper proposes adaptations for conducting MUSHRA tests with non-expert, crowdsourced listeners, focusing on generative speech codecs. We validate our approach by comparing results from MTurk and Prolific crowdsourcing platforms with expert listener data, assessing test-retest reliability and alignment. Additionally, we evaluate six objective metrics, showing that traditional metrics undervalue generative models. Our findings reveal platform-specific biases and emphasize codec-aware metrics, offering guidance for scalable perceptual testing of speech codecs.

Figures

Figures reproduced from arXiv: 2506.00950 by the authors.

Figure 1
Figure 1. Repeatability and validity of crowdsourced test on two different crowdsourcing platforms. This makes comparisons easier and reduces confusion, espe￾cially for non-expert listeners. Since we evaluated generative speech codecs, which typically introduce coding artifacts, we used Opus at 6 kbps as the anchor. In a pilot experiment, we tried using a low-pass filtered ver￾sion of the original signal—an anchor sometimes u… view at source ↗
Figure 2
Figure 2. Linear regression plots for each objective metric against subjective scores (Prolific 1). Blue and red lines/dots correspond to 3 DNN- and 3 DSP-based codecs, respectively, across 40 test items [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

32 extracted references · 30 canonical work pages

  1. [1]

    Conducting these studies requires significant time and cost, often relying on specialized labs or resorting to small- sample internal listening tests with expert listeners

    Introduction Subjective perceptual studies are the gold standard for evalu- ating human perception of audio quality and related research tasks. Conducting these studies requires significant time and cost, often relying on specialized labs or resorting to small- sample internal listening tests with expert listeners. As a result, evaluation using objective ...

  2. [2]

    Crowdsourcing MUSHRA Tests in the Age of Generative Speech Technologies: A Comparative Analysis of Subjective and Objective Testing Methods

    Crowdsourced MUSHRA testing Traditional MUSHRA testing has always faced difficulties and high costs. Setting up controlled listening environments, re- cruiting trained participants, and managing equipment can be expensive and time-consuming. Moreover, ensuring consistent and reliable results can be challenging due to variations in lis- tener focus and fat...

  3. [3]

    ranchor > rref

    The anchor is rated higher than the hidden reference, i.e. ranchor > rref

  4. [4]

    Var(r1, r2,

    All ratings for the non-anchor systems are identical, i.e. Var(r1, r2, . . . , rK ) = 0, where r1, . . . , rK are the scores for the K conditions under test (excluding the anchor but in- cluding the hidden reference), and Var(·) denotes variance. If 20% of the questions in a test block is less than 1, we take 1 as the threshold. Score-level outlier remova...

  5. [5]

    Experimental setup We set two main goals for evaluating our crowdsourced MUSHRA listening test

    Experiments 3.1. Experimental setup We set two main goals for evaluating our crowdsourced MUSHRA listening test. First, it should correlate well with an internal listening test conducted by expert listeners under con- trolled conditions (i.e., validity). Second, repeating the same test should yield the same results (i.e., repeatability). In this sec- tion...

  6. [6]

    ground truth

    Results and discussion 4.1. Repeatability and Validity From Figure 1, we see that both Prolific and MTurk listeners produce the same overall ranking of codecs as Internal listeners (i.e., “ground truth”). However, Prolific’s scores more closely trace the Internal curve, with closer absolute means and nar- rower confidence intervals, implying better alignm...

  7. [7]

    We show that Prolific yields closer absolute alignment with expert ratings than MTurk, though both maintain reliable relative rankings

    Conclusion Our study demonstrates that crowdsourced MUSHRA tests—when designed with appropriate screening and pro- cedural adjustments—can serve as reliable and repeatable alternatives to expert lab tests, even for generative speech codecs. We show that Prolific yields closer absolute alignment with expert ratings than MTurk, though both maintain reliable...

  8. [8]

    Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

    A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221) , vol. 2. Salt Lake City, UT, USA: IEEE, 2001, pp. 749–752. [Online...

Show all 32 references
  1. [9]

    Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part I–Temporal Alignment,

    J. G. Beerends, M. Obermann, R. Ullmann, J. Pomy, and M. Keyhl, “Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part I–Temporal Alignment,”J. Au- dio Eng. Soc., vol. 61, no. 6, 2013

  2. [10]

    Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part II–Perceptual Model,

    ——, “Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part II–Perceptual Model,”J. Audio Eng. Soc., vol. 61, no. 6, 2013

  3. [11]

    Non-intrusive Speech Quality Assess- ment for Super-wideband Speech Communication Networks,

    G. Mittag and S. Moller, “Non-intrusive Speech Quality Assess- ment for Super-wideband Speech Communication Networks,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Brighton, United Kingdom: IEEE, May 2019, pp. 7125–7...

  4. [12]

    NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,

    G. Mittag, B. Naderi, A. Chehadi, and S. M ¨oller, “NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,” in Interspeech 2021 . ISCA, Aug. 2021, pp. 2127–2131. [On- line]. Available: https://www.isca-archive.org/inte...

  5. [13]

    CROWD- MOS: An approach for crowdsourcing mean opinion score studies,

    F. Ribeiro, D. Florencio, C. Zhang, and M. Seltzer, “CROWD- MOS: An approach for crowdsourcing mean opinion score studies,” in 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Prague, Czech Re- public: IEEE, May 2011, pp. 2416–2419. [Onl...

  6. [14]

    An Open Source Implementa- tion of ITU-T Recommendation P.808 with Validation,

    B. Naderi and R. Cutler, “An Open Source Implementa- tion of ITU-T Recommendation P.808 with Validation,” in Interspeech 2020 . ISCA, Oct. 2020, pp. 2862–2866. [On- line]. Available: https://www.isca-archive.org/interspeech 2020/ naderi20 interspeech.html

  7. [15]

    Crowdsourced Multilingual Speech Intelligibility Testing,

    L. Lechler and K. Wojcicki, “Crowdsourced Multilingual Speech Intelligibility Testing,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Seoul, Korea, Republic of: IEEE, Apr. 2024, pp. 1441–1445. [Online]. Available: htt...

  8. [16]

    Method for the subjective assessment of intermediate quality level of coding systems,

    ITU, “Method for the subjective assessment of intermediate quality level of coding systems,” Geneva, Switzerland, 2015. [Online]. Available: https://www.itu.int/rec/R-REC-BS.1534/en

  9. [17]

    Audio quality evaluation by experienced and inexperienced listeners,

    N. Schinkel-Bielefeld, N. Lotze, and F. Nagel, “Audio quality evaluation by experienced and inexperienced listeners,” Montreal, Canada, 2013, pp. 060 016–060 016. [Online]. Available: https: //pubs.aip.org/asa/poma/article/808581

  10. [18]

    Fast and easy crowdsourced perceptual audio evaluation,

    M. Cartwright, B. Pardo, G. J. Mysore, and M. Hoffman, “Fast and easy crowdsourced perceptual audio evaluation,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Shanghai: IEEE, Mar. 2016, pp. 619–

  11. [19]

    High fidelity neural audio compression,

    A. D ´efossez, J. Copet, G. Synnaeve, and Y . Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research, 2023, featured Certification, Reproducibility Certification. [Online]. Available: https://openreview.net/forum? id=ivCd8z8zR2

  12. [20]

    webMUSHRA — A Comprehensive Framework for Web-based Listening Tests,

    M. Schoeffler, S. Bartoschek, F.-R. St ¨oter, M. Roess, S. Westphal, B. Edler, and J. Herre, “webMUSHRA — A Comprehensive Framework for Web-based Listening Tests,” Journal of Open Research Software , vol. 6, no. 1, p. 8, Feb. 2018. [Online]. Available: https://openresearchsoft...

  13. [21]

    Reproducible sub- jective evaluation,

    M. Morrison, B. Tang, G. Tan, and B. Pardo, “Reproducible sub- jective evaluation,”arXiv preprint arXiv:2203.04444, 2022

  14. [22]

    Speech quality evaluation of neural audio codecs,

    T. Muller, S. Ragot, L. Gros, P. Philippe, and P. Scalart, “Speech quality evaluation of neural audio codecs,” in Interspeech 2024 . ISCA, Sep. 2024, pp. 1760–1764. [On- line]. Available: https://www.isca-archive.org/interspeech 2024/ muller24c interspeech.html

  15. [23]

    P.808 : Subjective evaluation of speech quality with a crowdsourcing approach,

    ITU, “P.808 : Subjective evaluation of speech quality with a crowdsourcing approach,” 2021. [Online]. Available: https: //www.itu.int/rec/T-REC-P.808/en

  16. [24]

    Is it harder to perceive coding artifact in foreign language items? – A study with Mandarin Chinese and German speaking listeners,

    N. Schinkel-Bielefeld, Z. Jiandong, Q. Yili, A. K. Leschanowsky, and F. Shanshan, “Is it harder to perceive coding artifact in foreign language items? – A study with Mandarin Chinese and German speaking listeners,” Journal of the Audio Engineering Society, no. 9739, May 2017

  17. [25]

    P.800 : Methods for subjective determination of transmission quality,

    ITU, “P.800 : Methods for subjective determination of transmission quality,” 1998. [Online]. Available: https://www.itu. int/rec/T-REC-P.800-199608-I/en

  18. [26]

    Webex AI Codec: Delivering Next-level Audio Experiences with AI/ML,

    A. Dhingra, “Webex AI Codec: Delivering Next-level Audio Experiences with AI/ML,” Mar. 2024. [Online]. Available: https://blog.webex.com/collaboration/hybrid-work/ next-level-audio-with-webex-ai-codec/

  19. [28]

    V oice coding with opus,

    K. V os, K. V . Sørensen, S. S. Jensen, and J.-M. Valin, “V oice coding with opus,” inAudio Engineering Society Convention 135. Audio Engineering Society, 2013

  20. [29]

    Overview of the EVS codec architecture,

    M. Dietz, M. Multrus, V . Eksler, V . Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache, Y . Kamamoto, K. Kikuiri, S. Ragot, J. Faure, H. Ehara, V . Rajendran, V . Atti, H. Sung, E. Oh, H. Yuan, and C. Zhu, “Overview of the EVS codec architecture...

  21. [30]

    ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric,

    M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric,” Apr. 2020, arXiv:2004.09584 [eess]. [Online]. Available: http://arxiv.org/ abs/2004.09584

  22. [31]

    Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,

    C. K. Reddy, V . Gopal, and R. Cutler, “Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise sup- pressors,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6493–6497

  23. [32]

    SCOREQ: Speech quality assessment with contrastive regression,

    A. Ragano, J. Skoglund, and A. Hines, “SCOREQ: Speech quality assessment with contrastive regression,” in Advances in Neural Information Processing Systems , A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, Eds., vol. 37. Curran Associates, In...

  24. [623]

    Available: http://ieeexplore.ieee.org/document/ 7471749/

    [Online]. Available: http://ieeexplore.ieee.org/document/ 7471749/

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.