{"id":"a30eb37b-3c70-436a-96c6-e25a50d4db72","arxiv_id":"2506.01731","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper introduces SITool, a Flask-based DRT/MRT testing application, and uses it to show that STOI and ESTOI, but not WER, correlate with subjective intelligibility across 13 speech codecs.","lead":"SITool is a new open-source web toolkit for running diagnostic and modified rhyme intelligibility tests on speech codecs, used here to score 13 neural and traditional codecs with crowdsourced listeners. The benchmark finds that STOI and ESTOI track human intelligibility of these codecs while Whisper word error rate does not, which matters for anyone designing or evaluating low-bitrate neural audio systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'WER does not correlate' conclusion rests on a mismatched ASR measure: Whisper already errs on 19-25% of clean reference words, so the near-zero WER correlation may be an artifact of isolated-word ASR rather than evidence against WER.","rationale":"The reader's verdict is CONDITIONAL, and I agree with that overall assessment. My focus differs from the reader's weakest_assumption: the reader emphasizes the female-reference confound as the load-bearing concern, whereas I see the WER measurement as the more direct threat to the headline claim. The reader does mention 'the WER result is tied to a questionable Whisper-on-isolated-words setup' in the rationale, so there is partial overlap, but the reader does not make it the central weakest assumption. My concern is load-bearing because the abstract makes a comparative claim about which objective metrics correlate with subjective intelligibility; if WER is measured with a tool that fails on clean speech, then the non-correlation is an artifact of the tool, not a property of WER as an intelligibility metric. The paper's own admission that reference WER is high (Section 4.2) is the key piece of in-scope evidence. No ad hominem is intended; the issue is the validity of a proxy measure. The proposed test is feasible with publicly available ASR and alignment tools and would settle whether the WER conclusion survives a better-matched measurement. Since the paper remains a useful, clearly reported empirical contribution when the WER claim is appropriately qualified, my recommendation is unchanged from the reader's CONDITIONAL verdict.","tokens_in":10095,"tokens_out":3502,"duration_ms":38787,"concrete_test":"Recompute the Section 4.3 correlations after replacing Whisper-large WER with either (a) Whisper large fine-tuned on isolated words or prompted with the DRT closed-set alternatives, or (b) a phoneme-error rate from a forced aligner (e.g., Montreal Forced Aligner) on the same stimuli. Also report 95% bootstrap confidence intervals for all Pearson r values. If the WER-based correlation becomes significantly negative (e.g., 95% CI excluding 0) or the STOI/ESTOI correlations lose significance, the central 'not WER' claim fails; if the null result persists under a matched ASR, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central comparative claim is that 'only STOI and ESTOI, not WER, significantly correlate with subjective results.' This claim depends entirely on the WER values in Table 1, which are produced by Whisper 'large' applied to single-word utterances. Section 4.2 itself notes that WER is unexpectedly high even for the unprocessed reference signals (Original row: 0.25 female, 0.19 male). A clean reference should be almost perfectly recognized; error rates of 19-25% indicate that Whisper is not a reliable word recognizer for isolated DRT words, which are short, rhyming, and differ by a single phoneme. Consequently, the reported correlations in Section 4.3 (WER r = -0.11 to -0.15, versus STOI/ESTOI r = 0.595-0.958) conflate codec-induced intelligibility loss with ASR's task-level failure. The paper provides no significance tests or confidence intervals for any correlation, despite the word 'significantly' in the abstract. If WER is measured with a system better matched to isolated-word recognition, or with human word recognition on the same stimuli, the WER-subjective correlation could become substantial, directly undermining the headline. This is a measurement-validity concern, not a question of author diligence; the authors themselves flag the anomaly but do not treat it as disqualifying for the WER conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents SITool, an open Flask-based web application for conducting DRT and MRT intelligibility tests, and uses it to benchmark 14 codec conditions (13 codecs plus the original reference signal, at various bitrates) with crowdsourced and in-house listeners. It reports phoneme-level feature patterns, compares subjective DRT scores with STOI, ESTOI, and Whisper-based WER, and concludes that STOI/ESTOI correlate with subjective scores only after aggregating over gender and wordlist, while WER does not correlate. The paper also documents a lower female reference intelligibility and wordlist-specific phoneme effects, and discusses the mismatch between objective metrics and subjective gender/wordlist variation.","tokens_in":10309,"tokens_out":6087,"duration_ms":62170,"significance":"At the level of a systems and benchmark paper, this is a useful contribution: SITool fills a real gap as an open, deployable tool for standardized DRT/MRT tests, and the study design is careful, with trap questions, a clean gold-standard reference, a synthetic lower anchor, language screening, and LMM/repeated-measures ANOVA analyses. The DRT feature heatmaps and the gender/wordlist interaction analyses are informative for codec developers. However, the central statistical claim about WER is not yet supported, because the WER probe is not valid for the isolated-word DRT stimuli and because the correlation analysis lacks formal inference. With the WER claim repaired or carefully qualified, the toolkit's value stands.","major_comments":[{"comment":"The conclusion that 'only STOI and ESTOI, not WER, significantly correlate with subjective results' is not supported by the WER measurements as reported. Whisper-large is applied to isolated, single-word DRT utterances, and Table 1 shows a WER of 0.25 (female) and 0.19 (male) on the unprocessed reference signal; the authors themselves note in §4.2 that 'single-word utterances pose a challenge to Whisper.' Under those conditions, the near-zero WER correlation (§4.3: r = -0.11/-0.15) is an artifact of an ASR system failing on the task in general, not evidence that WER is unrelated to intelligibility for these codecs. The authors should re-compute WER with an ASR system validated on isolated-word or closed-set stimuli, or with a human word-recognition baseline on the same stimuli, or explicitly restrict the conclusion to the fact that their chosen ASR configuration did not track the subjective scores. As written, the headline comparative claim needs a load-bearing revision.","section":"§4.2, Table 1, §4.3"},{"comment":"Pearson correlations are reported without significance tests, confidence intervals, or any correction for the non-independence of observations repeated over codec conditions, wordlists, and genders, even though the abstract uses the word 'significantly.' The increases from r=0.595 to r=0.958 for STOI and from r=0.499 to r=0.890 for ESTOI after averaging are described descriptively, but the number of independent points after averaging is only the number of codec conditions (approximately 14), and with no interval estimate the strength of the claim is unclear. Please add bootstrap or nested correlation confidence intervals, test whether the averaged correlations differ from zero and from each other, and report the N underlying each correlation.","section":"§4.3"},{"comment":"The gender comparison is built on reference materials that already differ in intelligibility (female M=85, SD=9.66 vs male M=91.87, SD=8.23), and the authors acknowledge in §4.1 that 'these gender-specific differences already exist at the reference level.' Since the paper uses gender and wordlist interactions to argue that objective metrics fail to capture subjective variation, the analysis would be more convincing with a per-talker or per-word-pair control for the reference-level gender difference. The limitation is stated, but it should be treated as a constraint on the reported gender-related rankings rather than only as a future-work comment.","section":"§4.1"}],"minor_comments":[{"comment":"The first sentence contains a typo: 'We show ojective results' should read 'We show objective results.'","section":"§4.2"},{"comment":"The 'kHz' column lists values of 8, 16, and 24 for different codecs; please clarify whether this denotes the internal codec sampling rate before resampling to the 16 kHz test rate, so that the column is not misread as the test condition.","section":"Table 1"},{"comment":"The heatmap text for feature accuracies is very small, and the six distinctive features are only implicit from the DRT literature; adding a sentence that defines graveness, compactness, and the other features would help readers interpret the diagnostic analysis.","section":"Figure 3"},{"comment":"The comparison with existing DRT/MRT tools is brief; a short capabilities table contrasting SITool with the Qualtrics scripts of [15] and the NIST MRT tool would make the contribution claim more concrete.","section":"§1 and §3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's core toolkit and subjective benchmark are sound and likely publishable, but the WER comparison in the abstract and conclusions should not appear in its current form. If the authors can re-run with a validated ASR or human baseline, or reframe the WER finding as a tooling limitation, I would be happy to see a revision. The correlation analysis also needs formal inference before the word 'significantly' is used. The paper is more of a systems/benchmark contribution than a methods contribution, but it fits the journal's applied speech/audio evaluation scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core artifact here is legitimate: a public, Flask-based DRT/MRT toolkit that fills a real gap, plus a benchmark of 13 codecs with subjective intelligibility scores, phoneme-level breakdowns, and a correlation analysis against STOI/ESTOI/WER. The subjective design is careful — trap questions, gold standard, lower anchor, LMM/ANOVA, and a reasonable post-screening rule. The authors earn credit for releasing SITool and for flagging the female-reference anomaly rather than hiding it.\n\nThe soft spots are mostly in the objective correlation story. First, the abstract says STOI and ESTOI \"significantly correlate\" with subjective results and WER does not, but there are no significance tests or confidence intervals anywhere in Section 4.3. The word \"significantly\" is doing work without a p-value behind it. Second, the aggregation effect is real: r jumps from 0.595/0.499 to 0.958/0.890 when averaging over gender and wordlist. That is not necessarily a flaw — averaging across nuisance factors is a legitimate analysis choice — but it should be presented as such, not as one rising correlation.\n\nThe WER criticism from the stress test lands. The paper itself notes that Whisper large produces WER of 0.25 on clean female reference and 0.19 on clean male reference. For single isolated DRT words that are short, rhyming, and differ by one phoneme, a 19–25% error on the reference means the ASR is not measuring codec degradation; it is struggling with the task itself. The near-zero WER correlation is therefore likely an artifact of putting a general ASR on isolated-word recognition. The authors mention the anomaly but still draw a strong \"WER is useless\" conclusion. They should either use an ASR better matched to the task, or restrict the claim to say that this particular Whisper-based WER does not track DRT.\n\nThe female-reference deficit is a genuine confound for gender comparisons, though the authors acknowledge it and recommend controlling talker-specific effects in future work. That is an acceptable limitation for a benchmark paper, but some of the gender-interaction results should be read as exploratory.\n\nOverall: the tool and the subjective benchmark are solid and useful to codec developers and speech quality researchers. The correlation analysis is under-supported as currently written. This deserves a serious referee, but I would ask for revision rather than acceptance. The paper would be stronger with significance tests, confidence intervals on the correlations, a reworked or strongly qualified WER section, and ideally a second ASR comparison on the same stimuli.","headline":"SITool is a genuinely useful open-source DRT/MRT toolkit with a careful subjective benchmark, but the WER conclusion rests on an ASR mismatch and the correlation claims need statistical backing.","tokens_in":10962,"tokens_out":2201,"would_cite":true,"duration_ms":22180,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"STOI and ESTOI track listener-rated codec intelligibility; WER does not.","keywords":["SITool","speech intelligibility","Diagnostic Rhyme Test","Modified Rhyme Test","neural speech codecs","STOI","ESTOI","word error rate"],"falsifier":"Run a DRT study with several male and female talkers per wordlist and recompute per-condition correlations; if WER correlates strongly with subjective scores (for example, r > 0.5), or if STOI and ESTOI correlations fall well below 0.9 after gender and wordlist averaging, the paper's central claim is contradicted.","tokens_in":9843,"feed_emoji":"🎧","tokens_out":4725,"duration_ms":44685,"temperature":0.7,"pith_summary":"The paper introduces SITool, a publicly available web application for running Diagnostic Rhyme Tests (DRT) and Modified Rhyme Tests (MRT), and uses it to measure how well 13 neural and traditional speech codecs preserve intelligibility. Its central finding is that, judged by listeners, neural codecs can be more intelligible than traditional ones at lower bitrates, but the objective metrics that predict intelligibility are only STOI and ESTOI, and only when talker gender and wordlist are averaged away. Word error rate (WER) from an automatic speech recognizer shows no correlation with listener DRT scores, so WER should not be used as a proxy for intelligibility in this setting. The paper also reports that female reference recordings scored lower than male ones, suggesting talker-specific material effects that objective metrics miss.","feed_headline":"STOI and ESTOI track codec intelligibility; WER does not","feed_subtitle":"In a 13-codec DRT benchmark, only STOI and ESTOI match listeners, and only after averaging over gender and wordlist.","key_machinery":"The load-bearing instrument is SITool, a Flask-based web application that administers closed-set rhyme tests: listeners choose the word they heard from two options (DRT) or six (MRT). DRT pairs are organized by six distinctive acoustic features (voicing, nasality, sustension, sibilation, graveness, compactness), and scores are chance-corrected with P(c) = (R - W)/(R + W) * 100. Analysis rests on a Linear Mixed-Effects Model over codec condition, talker gender, and wordlist, plus Pearson correlations between mean subjective scores and the objective metrics STOI, ESTOI, and WER.","core_discovery":"On the paper's own terms, the discovery is a measurement result: across 13 codecs evaluated with DRT, subjective intelligibility correlates strongly with STOI (Pearson r = 0.958) and ESTOI (r = 0.890) once results are averaged over talker gender and wordlist, while WER is uncorrelated (r = -0.15 averaged). Because the correlation is much weaker before averaging (STOI r = 0.595, ESTOI r = 0.499), the objective metrics fail to capture the gender and wordlist variation that listeners show. The paper further finds that the female reference signal was less intelligible than the male one (85 vs 91.87), and that this gender gap propagates through codec comparisons, so future tests should control talker-specific effects.","pith_inferences":["Because DRT is closed-set and shows ceiling effects, an open-set SUS test might reveal intelligibility differences that STOI and ESTOI averaging hides; the paper itself points toward SUS as future work.","The gender gap in the reference signal suggests codec rankings could shift if re-tested with multiple talkers per gender, so a natural extension is to build a SITool benchmark with balanced multi-talker stimuli.","The weak pre-averaging correlation implies that a per-condition intelligibility heatmap from STOI or ESTOI would miss exactly the phoneme-level failures (such as /f/ and /θ/ in graveness items) that subjective DRT surfaces; combining objective metrics with phoneme-level error analysis may be more informative than any single score."],"forward_implications":["STOI and ESTOI can serve as screening metrics for codec intelligibility when talker gender and wordlist are averaged, but they should not be trusted for per-condition comparisons.","WER computed with an ASR system is not a valid proxy for DRT intelligibility, at least for single-word closed-set tests.","Below roughly 1.1 kbps, neural codec intelligibility drops relative to the uncoded signal, so ultra-low-bitrate codecs still have room for improvement in intelligibility.","Future SITool deployments should include reference signals and pre-test their intelligibility, since talker-specific and wordlist-specific effects can masquerade as codec performance.","DAC-S, despite a lower bitrate than EVS, showed higher subjective intelligibility for both genders, indicating neural codecs can beat traditional ones on intelligibility."],"supporting_citations":[{"why":"Supplies the DRT word lists, six distinctive features, and the chance-corrected scoring formula that define the subjective test.","marker":"[13]"},{"why":"Supplies the crowdsourced English audio stimuli and the AMT-based test procedure that SITool builds on.","marker":"[15]"},{"why":"Defines the ESTOI objective intelligibility metric whose correlation with subjective scores is a central result.","marker":"[5]"},{"why":"Defines the STOI objective intelligibility metric whose correlation with subjective scores is a central result.","marker":"[18]"},{"why":"Provides the ASR-based intelligibility prediction context that motivates comparing WER against subjective DRT scores.","marker":"[19]"},{"why":"Defines the EVS traditional 3GPP codec used as a high-bitrate baseline in the benchmark.","marker":"[33]"},{"why":"Introduces the DAC RVQGAN codec family from which the DAC-S neural codec is fine-tuned.","marker":"[7]"},{"why":"Introduces SemantiCodec, the diffusion-based ultra-low-bitrate neural codec included in the benchmark.","marker":"[27]"}],"fun_headline_variants":["STOI and ESTOI match subjective codec scores only after averaging","WER fails as a codec intelligibility metric in 13-codec DRT test","SITool: STOI and ESTOI correlate with DRT scores, WER doesn't","SITool: open-source DRT/MRT toolkit for codec intelligibility"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The crowdsourced reference recordings are assumed to be equally intelligible across talker genders, yet the female reference scored notably lower (85 vs 91.87), so gender comparisons and some codec rankings could reflect the recordings rather than the codecs.","fun_headline_variants_meta":{"raw":{"variants":["STOI and ESTOI match subjective codec scores only after averaging","WER fails as a codec intelligibility metric in 13-codec DRT test","SITool: STOI and ESTOI correlate with DRT scores, WER doesn't","SITool: open-source DRT/MRT toolkit for codec intelligibility"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000744,"raw_usage":{"total_tokens":3283,"prompt_tokens":872,"completion_tokens":2411,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":2322}},"tokens_in":488,"tokens_out":2411,"duration_ms":17740,"temperature":1.0,"reasoning_tokens":2322,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:34:30.820740+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a DRT study with several male and female talkers per wordlist and recompute per-condition correlations; if WER correlates strongly with subjective scores (for example, r > 0.5), or if STOI and ESTOI correlations fall well below 0.9 after gender and wordlist averaging, the paper's central claim is contradicted.","supporting_citations":[{"cited_title":"Speech quality evaluation of neural audio codecs,","cited_arxiv_id":null,"evidence_quote":"Supplies the DRT word lists, six distinctive features, and the chance-corrected scoring formula that define the subjective test."},{"cited_title":"The SUS test: A method for the assessment of text-to-speech synthesis intelligibility us- ing semantically unpredictable sentences,","cited_arxiv_id":null,"evidence_quote":"Supplies the crowdsourced English audio stimuli and the AMT-based test procedure that SITool builds on."},{"cited_title":"Unlike speech quality evaluation, DRT offered a more detailed analysis of phoneme-specific degradations and distinctive acoustic features","cited_arxiv_id":null,"evidence_quote":"Defines the ESTOI objective intelligibility metric whose correlation with subjective scores is a central result."},{"cited_title":"Subjective test methodology for assessing speech intelligibility,","cited_arxiv_id":null,"evidence_quote":"Defines the STOI objective intelligibility metric whose correlation with subjective scores is a central result."},{"cited_title":"Subjective evaluation of speech quality with a crowdsourcing approach,","cited_arxiv_id":null,"evidence_quote":"Provides the ASR-based intelligibility prediction context that motivates comparing WER against subjective DRT scores."},{"cited_title":"Methods for subjective determination of transmission quality,","cited_arxiv_id":null,"evidence_quote":"Introduces the DAC RVQGAN codec family from which the DAC-S neural codec is fine-tuned."},{"cited_title":"Neural dis- crete representation learning,","cited_arxiv_id":null,"evidence_quote":"Introduces SemantiCodec, the diffusion-based ultra-low-bitrate neural codec included in the benchmark."}],"review_version":1}