{"id":"53975a0e-abde-4b74-86d7-94073bcffc15","arxiv_id":"2506.16969","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A Mamba-based state-space ASR model and four fine-tuned self-supervised models achieve claimed state-of-the-art word error rates on whispered and normal speech across three English dialects, including near-perfect results on wTIMIT and CHAINS.","lead":"This paper reports very low word error rates for whispered and normal English speech across Singaporean, US, and Irish dialects using a Mamba state-space model and fine-tuned self-supervised systems. It claims the best published results on the wTIMIT and CHAINS whispered-speech benchmarks, with an efficient model trained on limited whispered data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported WERs are internally inconsistent: Table 2 and Section 4 assign different values to US whispered/normal (0.40 vs 0.92) and SG normal/whispered; the SOTA claim rests on these exact numbers, so they must be reconciled before the central claim can be evaluated.","rationale":"The reader's conditional verdict rests on possible partition leakage. That is a genuine risk, but it cannot be adjudicated from the manuscript alone, and the authors do provide code, which reduces the concern. The internal inconsistency, by contrast, is directly visible in the text and affects the exact numbers that constitute the abstract's SOTA claim. The prose and Table 2 disagree on whether US whispered WER is 0.40 or 0.92 and whether SG normal WER is 0.12 or 0.51. The pattern across models and the baseline table suggests the US columns are swapped, and the SG description in Section 4 is also inconsistent. If the table is as printed, the reader's strongest_claim misstates the US numbers; if the prose is correct, the table is wrong. Either way, the central empirical result is not reliably established until the discrepancy is resolved. Because the code is public, regeneration is feasible and should be the acceptance condition. This does not change the overall CONDITIONAL posture, but it changes the reason: not merely a hypothesized leakage, but a concrete misreporting that must be fixed before the SOTA claim can be accepted.","tokens_in":8014,"tokens_out":10134,"duration_ms":90408,"concrete_test":"Regenerate Table 2 from the public repositories (github.com/areffarhadi/Whisper_fine_tuning_ASR and github.com/areffarhadi/Mamba-ASR) using the released checkpoints and the exact wTIMIT/CHAINS splits described in Section 3.1. Verify, for the fine-tuned Whisper row, whether the audio condition for WER=0.40 is whispered US or normal US, and whether the WER=0.92 entry is normal US or whispered US; check the same for SG (0.51 vs 0.12). If the column labels are swapped, correct the tables and recompute the comparison against prior wTIMIT/CHAINS results; the abstract's best-performance claim should be re-evaluated against the corrected numbers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is not primarily partition leakage but an internal numerical inconsistency in the results. Section 4 states: \"For SG normal speech, a WER of 0.51% was obtained. For US dialect, the WER was 0.92% for whispered speech and 0.40% for normal speech.\" Table 2, however, lists for the fine-tuned Whisper row exactly the opposite pairing: Whisper-SG=0.51, Normal-SG=0.12, Whisper-US=0.40, Normal-US=0.92, Whisper-IRI=2.11. The same swap appears in baseline Table 1: the text says the US whispered WER is 5.97%, but the table shows Whisper-US=3.4 and Normal-US=5.97; the statement that SG normal WER is \"nearly twice\" the US value only works if Normal-US=3.4. Moreover, for HuBERT, Wav2Vec2, and WavLM, the labeled Normal-US columns are consistently worse than the labeled Whisper-US columns, which is the reverse of the expected difficulty ordering and is fixed by swapping the US column labels. Because the abstract's SOTA claim is quantified by these exact WERs, the headline numbers are not self-consistent as printed. Either Table 2 or the prose must be corrected, and the SOTA comparison re-derived.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Mamba-based state-space model (ConMamba encoder with Mamba decoder, versions Ver1 and Ver2) and fine-tuned self-supervised models (Wav2Vec2, WavLM, HuBERT, Whisper) for whispered speech recognition across Singaporean, US, and Irish dialects. Training uses wTIMIT Singaporean whispered and normal speech, CHAINS Irish normal speech, and (for the Mamba models) 1000 hours of LibriSpeech; evaluation covers SG whispered/normal, US whispered/normal, and Irish whispered speech. The paper reports very low WERs, e.g., 0.12% for Whisper on SG normal, 0.92% on US normal, and 2.11% on Irish whispered speech, and claims state-of-the-art performance on the wTIMIT and CHAINS datasets. The central technical claim is that a comparatively small Mamba model trained from scratch approaches or beats fine-tuned Whisper on several conditions while using far less whispered data.","tokens_in":8302,"tokens_out":6297,"duration_ms":61496,"significance":"If the reported results hold, they would represent a striking advance in whispered and multi-dialect ASR: current published WERs on whispered CHAINS are around 9%, whereas this paper reports 1.19% for Mamba-Ver2 and 2.11% for Whisper on the Irish whispered condition. The paper also makes a useful methodological contribution by comparing a from-scratch state-space model against fine-tuned self-supervised baselines under a controlled evaluation protocol, and it releases code. The main caveats are that the numerical results contain internal label inconsistencies that must be reconciled before the claims can be evaluated, the claim that the systems were not exposed to US data is contradicted by the use of LibriSpeech in Mamba training, and the state-of-the-art claim is supported by only a single self-cited comparison. These are fixable issues, but they currently prevent acceptance of the headline claims.","major_comments":[{"comment":"The US whispered/normal column assignments are interchanged between the prose and the tables, and this affects the headline numbers. In Table 2, the Whisper row lists Whisper-US=0.40 and Normal-US=0.92, but the prose states that for the US dialect the WER was 0.92% for whispered speech and 0.40% for normal speech. The same swap appears in Table 1: the prose says the US whispered WER is 5.97%, while the table shows Whisper-US=3.4 and Normal-US=5.97, and the statement that the SG normal WER is nearly twice the US value only holds if Normal-US=3.4. Since the abstract's state-of-the-art claim is quantified by these exact WERs, the tables and the prose must be reconciled and the resulting values re-derived before the results can be assessed.","section":"Section 4, Tables 1 and 2"},{"comment":"The claim that 'the systems were not provided with data for the US dialect, neither normal nor whispered speech' is contradicted by the training setup described for the Mamba models. Section 3.1 states that the Mamba models are trained from scratch on a mixture that includes 1000 hours of LibriSpeech, and LibriSpeech consists predominantly of US English speakers. Consequently, US normal speech is present in the Mamba training mixture, which invalidates the zero-shot interpretation of the US normal results unless LibriSpeech is explicitly excluded or the claim is restricted to the fine-tuning data of the self-supervised models. This needs to be clarified and, if necessary, the evaluation must be rerun without US English material in the Mamba training set.","section":"Section 3.1 and Section 4"},{"comment":"The claim of 'best performance reported on the wTIMIT and CHAINS datasets' is supported only by a single prior WER of 9.22% from the authors' own reference [4]. No comparison table of published wTIMIT/CHAINS results is provided, and no confidence intervals or statistical significance tests accompany the reported WERs. Since the central claim is a state-of-the-art statement, the manuscript should include a systematic comparison with all relevant prior published results on these datasets and report the variance across evaluation subsets or multiple runs.","section":"Abstract and Section 4"},{"comment":"The efficiency claim in the title and abstract is not quantified. The manuscript states that Mamba-Ver2 is 'significantly smaller' and trains with 'tiny data' compared to Whisper, but it does not report parameter counts, training time, FLOPs, or the actual hours of data used by each model. This is particularly important because the Mamba models are trained on 1000 hours of LibriSpeech in addition to the whispered data, so 'low data' refers only to the whispered portion. Please provide concrete efficiency metrics so that the efficiency contribution can be evaluated.","section":"Section 3.2 and Section 4"}],"minor_comments":[{"comment":"There are several typos and grammatical issues, for example 'autoencode' should be 'autoencoder', and 'Whereby we proposed an efficient model' should be rephrased as 'We therefore propose an efficient model'.","section":"Throughout"},{"comment":"The statement that 'the Whisper model developed by OpenAI is italicized to differentiate it easily from the concept of whispered speech' is not visibly implemented in the manuscript; the model name appears in the same font as surrounding text. Either apply the italics or remove the note.","section":"Section 1"},{"comment":"The caption says results are based on greedy and beam searches, but the Whisper and Mamba rows are evaluated only with beam search. Please state this explicitly in the caption or add the missing greedy values.","section":"Table 2 caption"},{"comment":"Figure 3 shows training and evaluation losses for Mamba-Ver1 and Ver2, but the text does not discuss the loss curves in detail. Please relate the curves to the reported WERs, especially the plateau behavior and whether early stopping was used.","section":"Figure 3"},{"comment":"Reference [4] is a self-citation of prior work by the same authors; it should be clearly marked as such in the text, and the comparison against it should be presented in a dedicated table rather than only in prose.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has a potentially valuable result, but the internal inconsistency between the prose and the tables in Section 4 is exactly the kind of error that undermines a state-of-the-art claim. In addition, the LibriSpeech mixing issue is not a minor wording problem: it directly affects the validity of the claimed zero-shot US dialect evaluation. I would recommend requiring a revised manuscript that corrects the tables, re-checks the training/evaluation separation, and provides a fuller comparison to prior work before sending the paper back for review. The self-citation of [4] as the sole SOTA baseline is acceptable only if it is indeed the only comparable result, but a broader comparison is needed for a journal-level claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read on 2506.16969. The genuinely new thing is the experimental design: they train on SG whispered + normal and Irish normal, then test on US whispered/normal as a zero-shot dialect condition. That's a more realistic evaluation than prior whispered-ASR work, which usually mixes both dialects into training. The Mamba-based model trained from scratch on LibriSpeech plus a few hours of whispered data comes close to a fine-tuned Whisper Large-v2 — that is a practical efficiency result worth having.\n\nThe weak spot is the numbers. The prose in Section 4 says SG normal WER 0.51% and US whispered 0.92%, but Table 2 lists SG whisper 0.51 and US normal 0.92. For every model, the US columns are reversed in difficulty: whispered US is better than normal US, which makes no sense. This is almost certainly a column-label swap, but until corrected the abstract's 'best performance' claim is unverifiable. The SOTA baseline is also only the self-cited 9.22% from [4]; no comparison table with independent results is given.\n\nSecondary issues: no confidence intervals or significance tests, and the efficiency claim (smaller model, low data) is stated without concrete numbers. The claim that no US audio leaked into training is credible but needs code and data splits to back it up. The code links are there, so this is fixable.\n\nOverall, this is a solid empirical study with a reasonable architecture and a well-designed evaluation. It's not high theory, but it could be a useful benchmark for whispered and multi-dialect ASR. I'd send it to peer review with a request to correct the table/prose mismatch, add a real comparison to prior art, and release exact data splits and eval scripts. As printed, I wouldn't rely on the specific WERs.","headline":"Useful whispered-ASR results and an interesting zero-shot dialect split, but the printed WERs are internally inconsistent and need fixing before the SOTA claim can be trusted.","tokens_in":8889,"tokens_out":4598,"would_cite":false,"duration_ms":42242,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Mamba-based state-space model plus fine-tuned self-supervised models achieves the lowest word error rates reported on the whispered-speech benchmarks wTIMIT and CHAINS.","keywords":["whispered speech recognition","state-space models","Mamba","self-supervised speech models","multi-dialect ASR","wTIMIT","CHAINS","word error rate"],"falsifier":"Check every utterance in the wTIMIT training portion, the LibriSpeech subset, and the self-supervised pretraining corpora for US-dialect speakers or US transcripts; if any US audio or its text appears in training, a retrained system on the clean split would show sharply higher US WERs, settling the zero-shot cross-dialect claim.","tokens_in":7777,"feed_emoji":"🤫","tokens_out":7184,"duration_ms":63646,"temperature":0.7,"pith_summary":"The paper tries to show that whispered speech recognition, made harder by dialect variation, can be handled by two complementary routes: fine-tuning large self-supervised models on a small whispered-speech corpus, and training a compact state-space model from scratch on the same data mixed with normal speech. The central claim is that this combination achieves the best word error rates yet reported on the wTIMIT and CHAINS whispered-speech benchmarks. A sympathetic reader should care because the proposed Mamba-based model reaches near-Whisper accuracy while using a fraction of the pretraining data and compute, which points toward ASR systems for low-resource acoustic conditions.","feed_headline":"Small Mamba ASR posts best whispered-speech WERs yet","feed_subtitle":"Fine-tuned Whisper stays under 1% on most tests; a from-scratch Mamba model nearly matches with far less data.","key_machinery":"The load-bearing mechanism is the ConMamba encoder, which replaces self-attention with bidirectional Mamba state-space layers, selective state-space models that process a sequence in linear time, and adds depthwise-separable convolutions to capture local acoustic structure such as phoneme boundaries. A unidirectional Mamba decoder combines the encoder output with autoregressive token predictions. Around this core, the paper's recipe mixes roughly 16 hours of wTIMIT whispered speech, Singaporean dialect, and normal and whispered CHAINS data with a thousand hours of LibriSpeech to train the Mamba model from scratch, and fine-tunes Wav2Vec2, WavLM, HuBERT, and Whisper on the same small multi-dialect corpus. The efficiency claim rests on the Mamba model's linear scaling, which lets a small model train on a low-range dataset and still model long-range dependencies.","core_discovery":"The paper's central discovery, stated on its own terms, is that a system trained only on Singaporean whispered and normal speech plus normal Irish speech can transcribe unseen US whispered and normal speech and unseen Irish whispered speech at very low error rates. Fine-tuning Whisper reaches WERs of 0.51% on Singaporean whispered, 0.12% on Singaporean normal, 0.40% on US whispered, 0.92% on US normal, and 2.11% on Irish whispered speech. The from-scratch Mamba-Ver2 model reaches 0.56% on Singaporean whispered, 0.63% on Singaporean normal, 1.75% on US whispered, 0.97% on US normal, and 1.19% on Irish whispered speech, beating Whisper on the Irish whispered condition. On the strength of these numbers the paper claims the best reported performance on wTIMIT and CHAINS for whispered speech recognition, with the efficient Mamba model trained on roughly 16 hours of whispered data mixed with LibriSpeech rather than on hundreds of thousands of hours.","pith_inferences":["The paper leaves implicit that the near-perfect US results imply the model has learned dialect-invariant acoustic representations rather than dialect-specific shortcuts; probing hidden states for dialect information would test this directly.","A direct extension would be to train the same ConMamba architecture on whispered data from another language family, since the model is trained from scratch without English-specific pretraining, success there would indicate the mechanism generalizes beyond English.","A controlled next experiment would train a standard transformer or Conformer from scratch on the identical data mix, isolating the efficiency gain attributable to the Mamba layers themselves.","Reproducing the exact wTIMIT and LibriSpeech split with another whispered corpus would clarify how much of the cross-dialect transfer comes from the small whispered set versus the large normal-speech mix."],"forward_implications":["If the results hold, a from-scratch state-space model trained on about 16 hours of whispered speech plus normal audiobook speech can come within a few tenths of a percent of Whisper, a model pretrained on 680,000 hours.","Fine-tuned Whisper's near-zero WERs on unseen US dialect suggest that large self-supervised models need only a small amount of whispered target-domain data to adapt across dialects and speaking styles.","Mamba-Ver2 beating Whisper on Irish whispered speech, 1.19% versus 2.11%, suggests that the state-space model is particularly robust to the unusually high speech rate of the CHAINS corpus, not just to whispering.","The training strategy removes the need for whisper-to-normal speech conversion or pseudo-whispered data augmentation, replacing them with a simple mix of small whispered and large normal corpora."],"supporting_citations":[{"why":"Provides the wTIMIT whispered speech corpus, including the Singaporean and US dialect split the experiments are built on.","marker":"[16]"},{"why":"Supplies 1,000 hours of normal speech mixed with the small whispered dataset to train the Mamba model from scratch.","marker":"[20]"},{"why":"Provides the strong pretrained baseline and fine-tuned model that achieves the lowest overall WERs.","marker":"[15]"},{"why":"Sets the prior WER of 9.22% on CHAINS whispered speech that the paper's claimed best performance is measured against.","marker":"[4]"},{"why":"Introduces the selective state-space layers that the ConMamba encoder uses in place of self-attention.","marker":"[21]"},{"why":"Provides the Mamba encoder-decoder and cross-modality design that the proposed decoder builds on.","marker":"[22]"},{"why":"Defines the Wav2Vec2 self-supervised model that is fine-tuned on the whispered multi-dialect corpus.","marker":"[12]"},{"why":"Defines the WavLM self-supervised model that is fine-tuned and compared against.","marker":"[13]"},{"why":"Defines the HuBERT self-supervised model that is fine-tuned and compared against.","marker":"[14]"}],"fun_headline_variants":["Mamba ASR sets best whispered-speech WERs on wTIMIT, CHAINS","State-space model beats Whisper on Irish whispers, less data","Low-data Mamba matches fine-tuned Whisper on whispered ASR","Mamba model tops whispered-speech recognition with 16 hours of data","From-scratch Mamba nearly matches Whisper on cross-dialect whispers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything depends on the dataset split being exactly as described: only Singaporean speech appears in training, US speech appears only in testing, Irish whispered speech appears only in testing, and no US utterances hide inside the LibriSpeech mix or the pretraining corpora.","fun_headline_variants_meta":{"raw":{"variants":["Mamba ASR sets best whispered-speech WERs on wTIMIT, CHAINS","State-space model beats Whisper on Irish whispers, less data","Low-data Mamba matches fine-tuned Whisper on whispered ASR","Mamba model tops whispered-speech recognition with 16 hours of data","From-scratch Mamba nearly matches Whisper on cross-dialect whispers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000276,"raw_usage":{"total_tokens":1643,"prompt_tokens":938,"completion_tokens":705,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":603}},"tokens_in":554,"tokens_out":705,"duration_ms":6677,"temperature":1.0,"reasoning_tokens":603,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:14:59.316854+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check every utterance in the wTIMIT training portion, the LibriSpeech subset, and the self-supervised pretraining corpora for US-dialect speakers or US transcripts; if any US audio or its text appears in training, a retrained system on the clean split would show sharply higher US WERs, settling the zero-shot cross-dialect claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the wTIMIT whispered speech corpus, including the Singaporean and US dialect split the experiments are built on."},{"cited_title":"Whispered speech recogni- tion using deep denoising autoencoder and inverse filtering,","cited_arxiv_id":null,"evidence_quote":"Provides the strong pretrained baseline and fine-tuned model that achieves the lowest overall WERs."},{"cited_title":"As a base- line, we evaluated the performance of the pre-trained Whisper Large-v2 model on the test set to assess the need for a special- ized system for the proposed challenges","cited_arxiv_id":null,"evidence_quote":"Sets the prior WER of 9.22% on CHAINS whispered speech that the paper's claimed best performance is measured against."},{"cited_title":"wtimit whispered timit dataset,","cited_arxiv_id":null,"evidence_quote":"Provides the Mamba encoder-decoder and cross-modality design that the proposed decoder builds on."},{"cited_title":"End-to- end whispered speech recognition with frequency-weighted ap- proaches and pseudo whisper pre-training,","cited_arxiv_id":null,"evidence_quote":"Defines the WavLM self-supervised model that is fine-tuned and compared against."},{"cited_title":"Improving whispered speech recognition performance using pseudo-whispered based data augmentation,","cited_arxiv_id":null,"evidence_quote":"Defines the HuBERT self-supervised model that is fine-tuned and compared against."}],"review_version":1}