{"id":"53de18dc-b1c1-4c45-a7bc-0bed71773a07","arxiv_id":"1908.04751","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In a small study, replacing hard-to-hear consonant-vowel sounds with more robust talkers or with equally intelligible different vowels improved consonant recognition for most hearing-impaired ears, though results varied.","lead":"Speech scientists tested whether replacing hard-to-hear consonant-vowel sounds with easier versions or with different vowels can improve consonant recognition in hearing-impaired listeners. On average, most listeners did better, but the effects varied by person, vowel, and consonant, suggesting hearing aids could be tuned to each ear.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The frequency fine-tuning results depend on NH SNR90 transferring to HI ears as a cue-salience metric; the paper's own HI6 case shows tokens matched on SNR90 are not matched for that ear, so the vowel-change improvements may reflect audiogram audibility rather than fine-tuning.","rationale":"I agree with the reader's weakest assumption: the load-bearing step is the transfer of NH-measured SNR90 to HI ears as a cue-salience metric. The entire experimental design, especially the frequency fine-tuning comparison, relies on this transfer to control the intensity of the primary cue. Without it, the vowel-change improvements could simply reflect that the replacement token happens to place acoustic energy in a region of the audiogram that is more audible for a given HI ear, not that 'frequency fine-tuning' per se is beneficial. The manuscript contains direct evidence for this concern: the discussion of subject HI6 (Figs. 9-10) shows large recognition differences among tokens that were matched on SNR90, explained by audiogram-specific energy placement. The abstract/text inconsistency (92% vs 85% improved) and the small sample without inferential statistics are real reporting issues, but they are addressable by re-reporting and do not threaten the qualitative direction of the cue-enhancement result. The SNR90-transfer concern, by contrast, undermines the interpretation of the fine-tuning experiment and the proposed prescriptive use of SNR90. I therefore recommend retaining the reader's CONDITIONAL verdict, now with an explicit condition that the transfer assumption be tested via audiogram-based audibility analysis.","tokens_in":14527,"tokens_out":7708,"duration_ms":79940,"concrete_test":"Reanalyze the vowel-change data for each HI ear using an audiogram-weighted audibility metric (e.g., Speech Intelligibility Index or a band-importance-weighted SNR) computed for the consonant cue region of each token. Split the matched token pairs into those with similar versus different audiogram-based audibility in the tested ear. If the improvement percentages in Table 1 are concentrated in pairs where the replacement token has higher audiogram-based audibility, then SNR90 matching did not control cue intensity and the frequency fine-tuning claim is an audibility confound. As a second check, test whether the NH-based SNR90 ordering of T1/T2 tokens is preserved in each HI ear's error rates; if not, the transfer assumption fails directly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim presumes that SNR90, estimated from 30 NH listeners, is a valid measure of the perceptual salience of the primary cue in each HI ear (Introduction; Methods II.A, 'noise-robust... error as measured by 30 NH ears'). The frequency fine-tuning experiment compares CVs matched on NH SNR90 and attributes error changes to vowel context. If SNR90 does not predict the audibility of the consonant cue in a given HI ear, then 'similar SNR90' tokens are not matched in the variable being controlled, and the reported 63-75% improvement rates for vowel changes (Table 1) could reflect uncontrolled token difficulty. The paper itself supplies evidence that this transfer fails in at least one case: subject HI6 recognizes /bae/ much better than /bA/ and /bE/ despite comparable SNR90, and the authors explain the difference by the placement of burst energy and vowel formants relative to HI6's audiogram (Discussion, Figs. 9-10). That is precisely an audiogram-dependent cue salience effect that SNR90 matching did not control. Hence the proposed SNR90-based prescription is not supported unless the transfer assumption is directly tested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a speech-based hearing test that uses a normal-hearing (NH) perceptual measure, SNR90, to select noise-robust tokens and to identify error-prone phones for individual hearing-impaired (HI) ears. In two experiments, the authors study (1) cue enhancement, where a CV token is replaced by the same CV spoken by a talker with a lower (more salient) SNR90, and (2) frequency fine-tuning, where the consonant is kept fixed and the vowel is changed to another vowel matched in SNR90. The reported results are that cue enhancement improves recognition on average (92% in the abstract, 85% in the text), while vowel changes improve 63–75% of tokens depending on the vowel. The authors propose using confusion patterns, cue-enhancement responses, and fine-tuning responses to tailor hearing aid amplification to individual ears.","tokens_in":14759,"tokens_out":2504,"duration_ms":27212,"significance":"If the SNR90-based approach works, it would provide a principled, token-specific alternative to broad audibility-based hearing aid fitting, with the potential to explain idiosyncratic HI consonant confusions. The paper's strengths include its fine-grained token-level analysis rather than class-averaged errors, the use of an external NH benchmark (SNR90) as a control variable, and a detailed case study (HI6) linking spectrotemporal token properties to audiometric configuration. However, the significance is substantially limited by the small sample (8 ears from 6 subjects), the lack of any statistical inference, and an internal inconsistency in the headline improvement rate. The central idea is promising, but the current evidence is preliminary and the quantitative claims are not yet reliable.","major_comments":[{"comment":"The abstract states that 'On average, 92% of tokens are improved when we replaced the CV with the same CV but with a more robust talker,' but the Results section reports that '85% of tokens are improved and 10% of tokens are degraded' and no 92% figure appears anywhere in the text. This is a direct internal inconsistency in the paper's primary quantitative claim. The reader cannot tell which number is correct, and the discrepancy between an 85% improvement rate and a 92% improvement rate (with an unexplained 10% degradation) is large enough to affect the interpretation of the cue-enhancement result.","section":"Abstract and §3.A.ii (Fig. 6)"},{"comment":"The frequency fine-tuning experiment matches tokens on SNR90 values estimated from 30 NH listeners (Methods II.A, Table 1), but the paper's own Discussion shows that this matching does not control cue salience in HI ears. For subject HI6, /bae/ is recognized far better than /bA/ and /bE/ despite comparable SNR90, and the authors attribute this to the placement of burst energy and vowel formants relative to HI6's audiogram (Figs. 9–10). That is precisely an audiogram-dependent, HI-specific effect that SNR90 matching does not capture. Therefore the improvements attributed to vowel change in Table 1 may reflect uncontrolled token audibility or difficulty in the HI ear rather than a frequency fine-tuning effect. To support the central claim, the authors need to directly test whether NH-derived SNR90 transfers to HI ears (e.g., by measuring HI recognition at SNRs around the putative SNR90 or by demonstrating that matched tokens produce matched errors in HI ears) or, failing that, substantially weaken the claims about frequency fine-tuning.","section":"Methods II.A and Discussion (Figs. 9–10)"},{"comment":"All quantitative claims in §3 (the 85%/10% improvement/degradation rates, the 63–75% vowel-change improvements, and the subject/consonant differences in Fig. 8) are reported as raw token counts with no confidence intervals, standard errors, or significance tests. With only 8 ears, a response-repeat option, and an adaptive selection procedure that determines which tokens reach List 3, the sampling variability is considerable. For example, a claim that '75% of tokens are improved' based on a small number of errorful tokens per ear could easily arise from chance. The manuscript should include per-subject and per-token error bars, a statement of how many tokens contributed to each percentage, and a statistical test (e.g., sign test or mixed-effects model) for the improvement rates. Without this, the numerical improvement percentages are not interpretable.","section":"Methods II.D and Table 1"}],"minor_comments":[{"comment":"The caption of Fig. 6 is incomplete: it reads only 'Figure 6: .' and should describe the left and right panels and the meaning of the bars.","section":"Fig. 6"},{"comment":"There is a typo: 'Conversly' should be 'Conversely'.","section":"Introduction, after Fig. 2"},{"comment":"There is a typo: 'when the nosie decreases' should be 'when the noise decreases.'","section":"§3.A.i"},{"comment":"The header 'Changed V owel' has a spacing typo; it should read 'Changed Vowel.'","section":"Table 1"},{"comment":"The figure caption states that cases where the vowel change vanished the error are not shown, but Table 1 appears to include such cases in the improvement percentages. Please clarify whether the Table 1 percentages count only error-reducing changes or also error-vanishing changes, and make the definitions consistent between text, table, and figure.","section":"Fig. 8 and Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper presents an interesting clinical direction, but the main quantitative claim is internally inconsistent, and the key assumption that NH-derived SNR90 transfers to HI ears is directly contradicted in one of the paper's own case studies. The revision would need to correct the abstract/text discrepancy, add statistical support, and either validate or explicitly disclaim the transfer assumption. If the transfer assumption cannot be tested within the scope of this dataset, the frequency fine-tuning conclusions should be reframed as exploratory. The scope of the journal (q-bio.QM) is appropriate, but the paper may also benefit from a more careful discussion of how the adaptive list transitions affect the comparability of tokens across ears."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth a look, but with a clear head. The genuinely new thing is the experimental design: by matching tokens on SNR90 measured in NH listeners, the authors try to separate talker-based cue enhancement from vowel-based frequency fine-tuning in HI consonant recognition. That is a real step beyond earlier work that averaged over tokens, and the cue enhancement result — most tokens improve when the talker is swapped for one with lower SNR90 — is plausible and probably holds up in direction.\n\nThe paper also does something good that is easy to miss: it reports individual token-level confusion patterns and goes to the spectrogram to explain why one HI subject hears /bæ/ much better than /bA/ or /bE/. That case analysis is honest and useful, even if it undercuts the paper's own control assumption.\n\nNow the soft spots, in proportion. The abstract says 92% of tokens improved with talker change; the body says 85% improved and 10% degraded. That is a real internal inconsistency in the central quantitative claim, and it needs fixing before this can be cited for numbers. Second, no statistical analysis. Eight ears from six subjects, no error bars, no test of whether the improvement rates differ from chance. For a diagnostic proposal, that is thin. Third, the transfer assumption: SNR90 is measured in NH listeners, and the paper assumes it works as a cue-salience metric in HI ears. The HI6 case shows this can fail — tokens matched on SNR90 were not matched for that ear, and the difference is explained by the audiogram. That means the \"frequency fine-tuning\" improvements may be partly uncontrolled token difficulty rather than a clean effect of vowel context. The authors are aware of this; they qualify their conclusions, but the qualification doesn't rescue the average percentages.\n\nAlso note the selective adaptive design: only tokens that reach List 3 get full testing, so the improvement/degradation counts are conditional on being errorful in the first place. That is not fatal, but it limits generalization to the full token set.\n\nWhere does this land? The core idea — using SNR90-matched tokens to test cue-level interventions per ear — is worth pursuing, and the confusion-pattern analysis is a good example for the field. But the evidence in this paper is preliminary. The 92 vs 85 discrepancy alone should be resolved before I'd trust the numbers.\n\nThat said, this deserves peer review, not a desk reject. A serious referee can push for the re-analysis and a direct test of the SNR90 transfer assumption. I'd bring it to a reading group if the group works on speech perception or hearing aids, but I wouldn't cite it for the quantitative claims.","headline":"Small-sample but genuinely novel SNR90-based test separates cue enhancement from vowel fine-tuning in HI consonant recognition; the qualitative direction is credible, but the abstract/text inconsistency and the NH-to-HI transfer assumption keep the quantitative claims provisional.","tokens_in":15278,"tokens_out":2023,"would_cite":false,"duration_ms":21330,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A token-specific noise-robustness measure, SNR90, can identify which consonant cues to amplify or frequency-shift for each hearing-impaired ear, improving recognition for 92% of tokens when the cue is enhanced.","keywords":["SNR90","hearing-impaired speech recognition","consonant recognition","perceptual cue","hearing aid amplification","confusion pattern","frequency fine-tuning","phone recognition"],"falsifier":"Directly measure SNR90 in hearing-impaired ears for the same tokens, or compare HI error rates across tokens matched by NH SNR90: if tokens with equal NH SNR90 produce significantly different HI recognition when presented at the same level above threshold, the transfer assumption fails and the reported vowel-change improvements would be attributable to uncontrolled token difficulty rather than frequency fine-tuning.","tokens_in":14314,"feed_emoji":"🦻","tokens_out":5403,"duration_ms":50065,"temperature":0.7,"pith_summary":"This paper tries to establish that a single token-specific noise-robustness measure, SNR90, can guide hearing-aid-style interventions for individual hearing-impaired ears. It reports that replacing a difficult consonant-vowel token with the same CV spoken by a more salient talker improves recognition in 92% of tokens, and that changing the vowel while holding SNR90 similar improves a majority of tokens (63 to 75 percent depending on the vowel). If true, speech-based hearing tests could prescribe which consonant cues to amplify, and how to fine-tune their frequency content, for each ear rather than applying the same gain curve to everyone. The authors propose using per-token confusion patterns to decide when cue enhancement or frequency fine-tuning helps.","feed_headline":"Better talker fixes 92% of hard consonant tokens","feed_subtitle":"The SNR90 score from normal hearing points to the consonant cue each impaired ear needs amplified.","key_machinery":"The central object is the SNR90 perceptual measure, defined as the signal-to-speech-weighted-noise ratio at which a normal-hearing listener recognizes a token with at least 90% accuracy on average. It does two jobs: it selects noise-robust tokens (error below 10% at SNR = -2 dB) for the hearing test, and it defines equivalence classes of tokens so that swapping talkers or vowels holds the approximate intensity of the primary cue constant. The adaptive List 1 to List 2 to List 3 procedure then identifies which tokens are error-prone for a particular HI ear and tests alternatives against them.","core_discovery":"The paper's claim is that SNR90, the signal-to-speech-weighted-noise ratio at which normal-hearing listeners recognize a token with 90% accuracy on average, is a valid perceptual measure of the salience of the primary cue in a consonant-vowel token, and that it can be used to control token difficulty while testing hearing-impaired ears. In an adaptive experiment, error-prone tokens identified for each HI ear were replaced either by the same CV from a talker with a lower SNR90 (more salient) or by a token with the same consonant, a different vowel, and a similar SNR90. On average, the more salient talker improved recognition for 92% of tokens, including cases where the error vanished; among tokens that remained errorful in both conditions, 85% improved and 10% degraded. Vowel changes matched in SNR90 improved recognition for 63 to 75 percent of tokens depending on the vowel. The authors conclude that HI ears use the same primary cues as NH ears, and that individual confusion patterns, interpreted alongside the audiogram, can identify when amplification should be applied and whether the frequency content of the cue should be shifted.","pith_inferences":["If NH SNR90 transfers to HI ears, the same token bank could be reused as a standardized perceptual probe, mapping each ear's confusion patterns to a personalized gain or filter recommendation without collecting HI-specific SNR90 curves.","A natural extension not performed in this paper would be to measure SNR90 directly in hearing-impaired ears and compare it to the NH values; if they correlate, SNR90 could be estimated from the audiogram, making the test practical for clinicians.","The frequency fine-tuning results suggest a testable prediction: for a consonant whose primary cue lies in a high-frequency band, shifting the coarticulated vowel to move burst and formant energy into a region of residual hearing should improve recognition more than flat amplification.","The 92% improvement rate under talker replacement might partly reflect talker clarity differences beyond primary-cue salience, such as speaking rate or vocal effort; an experiment controlling for duration and vocal effort would separate those mechanisms."],"forward_implications":["Hearing aid fitting could move from a universal gain prescription to a token-specific test that first finds the consonants each ear mishears.","Enhancing the primary cue by selecting a more salient talker should reduce consonant errors for most tokens in most HI ears, with degradation concentrated in ears that also have low-frequency hearing loss.","Vowel changes matched in SNR90 offer a way to shift the frequency content of the primary cue without changing its overall salience, improving most tokens though less consistently than talker enhancement.","Per-token confusion patterns can point to the acoustic feature, such as burst strength, voice onset time, or formant location, that needs amplification for a given ear.","The minority of tokens that degrade under cue enhancement or vowel change implies that these interventions should be evaluated per ear rather than assumed beneficial for everyone."],"supporting_citations":[{"why":"Supplies the accumulated error-difference method and the evidence that some HI ears are hurt by frequency-dependent gain, motivating token-specific interventions.","marker":"Abavisani and Allen (2017)"},{"why":"Source of the CV tokens and the 30-NH-ear SNR90 measurements from which the T1 and T2 sets were drawn.","marker":"Li et al. (2010)"},{"why":"Provides the master error curve showing the score drop over about 6 dB and the statistical basis for testing noise levels above SNR90 plus 6 dB.","marker":"Singh and Allen (2012)"},{"why":"Defines the noise-robust token criterion (less than 10% error at -2 dB, average error under 1 in 32) used to select the test tokens.","marker":"Phatak and Allen (2007)"},{"why":"Supports the claim that HI errors are token-specific rather than class-specific and that HI ears use cues similar to NH ears.","marker":"Trevino and Allen (2013b)"},{"why":"Documents token-specific confusion groups and the steep SNR range of the psychometric function, justifying SNR90 as a perceptual measure.","marker":"Toscano and Allen (2014)"},{"why":"Shows that SNR90 correlates with the relative intensity of the primary cue, the bridge that lets the authors match tokens by SNR90 when changing vowels.","marker":"Li and Allen (2011)"},{"why":"Used to interpret voice onset time as a perceptual cue in the /b/ case study that explains why one vowel context helped a particular HI ear.","marker":"Lisker (1975)"}],"fun_headline_variants":["SNR90-guided talker choice fixes 92% of consonant errors","Robust talker boosts impaired-ear consonant recognition 92%","Frequency fine-tuning of cue yields 63-75% improvement per vowel","Personalized hearing aid amplification guided by SNR90 cue salience","Cue enhancement via talker replacement improves 92% of tokens"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that a token's SNR90 measured in normal-hearing listeners carries over to hearing-impaired ears as a measure of the salience of the primary cue, so that matching tokens on NH SNR90 actually controls cue intensity for HI listeners.","fun_headline_variants_meta":{"raw":{"variants":["SNR90-guided talker choice fixes 92% of consonant errors","Robust talker boosts impaired-ear consonant recognition 92%","Frequency fine-tuning of cue yields 63-75% improvement per vowel","Personalized hearing aid amplification guided by SNR90 cue salience","Cue enhancement via talker replacement improves 92% of tokens"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001418,"raw_usage":{"total_tokens":5780,"prompt_tokens":1052,"completion_tokens":4728,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":668,"completion_tokens_details":{"reasoning_tokens":4637}},"tokens_in":668,"tokens_out":4728,"duration_ms":28576,"temperature":1.0,"reasoning_tokens":4637,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:07:21.059600+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Directly measure SNR90 in hearing-impaired ears for the same tokens, or compare HI error rates across tokens matched by NH SNR90: if tokens with equal NH SNR90 produce significantly different HI recognition when presented at the same level above threshold, the transfer assumption fails and the reported vowel-change improvements would be attributable to uncontrolled token difficulty rather than frequency fine-tuning.","supporting_citations":[{"cited_title":"and Allen, J","cited_arxiv_id":null,"evidence_quote":"Supplies the accumulated error-difference method and the evidence that some HI ears are hurt by frequency-dependent gain, motivating token-specific interventions."},{"cited_title":"and Allen, J","cited_arxiv_id":null,"evidence_quote":"Provides the master error curve showing the score drop over about 6 dB and the statistical basis for testing noise levels above SNR90 plus 6 dB."},{"cited_title":"and Allen, J","cited_arxiv_id":null,"evidence_quote":"Defines the noise-robust token criterion (less than 10% error at -2 dB, average error under 1 in 32) used to select the test tokens."},{"cited_title":"and Allen, J","cited_arxiv_id":null,"evidence_quote":"Documents token-specific confusion groups and the steep SNR range of the psychometric function, justifying SNR90 as a perceptual measure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Used to interpret voice onset time as a perceptual cue in the /b/ case study that explains why one vowel context helped a particular HI ear."}],"review_version":1}