REVIEW 4 major objections 5 minor 37 references
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A Mamba-based temporal transform block matches a transformer-based block in binaural speech intelligibility prediction while using fewer learnable parameters.
desk verdict A clean but small empirical swap of transformer for Mamba in a binaural SIP model; the accuracy claim holds, but the efficiency motivation and the spatial-information claims outrun the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Mamba block, a selective state-space model that updates a hidden state $\mathbf{h}_t = \bar{A}\mathbf{h}_{t-1} + \bar{B}\mathbf{x}_t$ with input-dependent parameters $\bar{A}$, $\bar{B}$, $C$, and $\Delta$, running the sequence scan in linear time with constant memory. The paper inserts this block as the temporal transform block in place of self-attention in the monaural model and in place of self-attention plus cross-attention in the binaural model. The binaural variant uses bidirectional Mamba, defined as the sum of a forward Mamba scan and a Mamba scan on the time-reversed sequence, followed by a skip connection and Gaussian error linear units, which is intended to capture interaural contextual and spatial information without an explicit cross-attention module. The rest of the pipeline—a frozen large speech-recognition encoder's layer features, temporal pooling, layer-wise transformer, and audiogram conditioning—is kept fixed, so the Mamba block is the only replaced component.
What would settle it
Measure end-to-end latency, peak memory, and energy for the full pipeline—frozen speech encoder plus temporal transform block—on the same low-power hearing-aid-class processor for both the transformer baseline and the bidirectional Mamba model; if the Mamba system is not faster or lighter, or if a paired significance test on a larger held-out set shows no RMSE difference, the paper's motivation and comparative claim would fail.
Extended reading notes
Core claim
On the binaural speech-intelligibility data used in this study, the bidirectional Mamba model achieves an average RMSE of 27.34 against the transformer baseline's 27.49, with 5.01 million versus 5.23 million learnable parameters, and gives the lowest average RMSE on two of the three test partitions and overall. The differences are not statistically significant by the paper's own Wilcoxon signed-rank test, so the paper's claim is that Mamba performs competitively, not that it outperforms. The binaural Mamba block replaces both self-attention and cross-attention with a forward and a time-reversed Mamba scan plus a skip connection and GELU nonlinearity, and the paper interprets the binaural model's larger improvement over its monaural version as evidence that bidirectional Mamba captures contextual and spatial binaural information. In monaural experiments, Mamba also lowers average RMSE relative to the transformer baseline (28.25 versus 29.45), while LSTM-based blocks do not.
Load-bearing premise
The load-bearing premise is that Mamba's theoretical linear-time, constant-memory inference advantage survives in the full model and that its prediction accuracy equals or improves on the transformer's, but the paper reports no runtime or memory measurements and its accuracy differences are not statistically significant.
Editorial extensions
If this is right
- Lower-parameter temporal processing is achievable: the binaural Mamba model's temporal transform block has about 5.0 million parameters versus the transformer's 5.2 million.
- Binaural processing works without an explicit cross-attention module, with the bidirectional Mamba block plus skip connection and GELU appearing sufficient to leverage left-right spatial speech information.
- Temporal resolution of the input features has little effect on Mamba-based prediction, so high-resolution temporal input is not needed to reach competitive accuracy.
- Because Mamba scales linearly in sequence length, the temporal transform's cost should grow more slowly with longer inputs than self-attention's quadratic cost, which could benefit long-duration audio processing.
Reading between the lines
- The reported accuracy differences between Mamba and transformer are not statistically significant; a larger listener set or test set would be needed to establish that the two approaches are not interchangeable at the block level.
- Because the frozen speech encoder likely dominates the full model's compute and memory, the end-to-end efficiency gain from swapping the temporal block could be small; measuring full-pipeline speed on hearing-aid-class hardware would settle whether the motivation holds.
- The bidirectional Mamba uses two sequential passes over the sequence, so a unidirectional model with a larger state or a different binaural fusion rule might be a cheaper way to capture the same spatial context—a testable extension would be comparing Mamba fusion against simple inter-channel feature concatenation.
- Since temporal pooling size had little effect, an adaptive or stride-based pooling schedule could reduce compute while retaining speech intelligibility accuracy, which the paper does not pursue.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes replacing the transformer-based temporal transform blocks in the non-intrusive binaural speech intelligibility prediction (SIP) model E011 with Mamba blocks, and evaluates the resulting monaural and binaural models on the CPC2 dataset. The authors report that a binaural model using bidirectional Mamba achieves an average RMSE of 27.34, slightly lower than the transformer baseline's 27.49, while using fewer learnable parameters in the temporal block (5.01M vs. 5.23M). They also test LSTM baselines and vary the temporal pooling size of Whisper features. The central claim is that Mamba provides competitive prediction accuracy with a smaller temporal-block parameter count, and the authors further suggest that bidirectional Mamba captures contextual and spatial information from binaural signals.
Significance. If the reported accuracy holds, the paper demonstrates a credible drop-in replacement for transformer temporal blocks in a state-of-the-art SIP pipeline, using a public challenge dataset, a public baseline implementation, and official Mamba code. The work is reproducible in principle and addresses a practical question about efficient non-intrusive metrics for hearing-impaired listeners. However, the significance is limited in two ways. First, the efficiency motivation that frames the paper is never measured: no runtime, memory, or power figures are given, and the frozen Whisper feature extractor likely dominates inference cost. Second, the numerical advantage over the transformer baseline is small and, by the authors' own Wilcoxon signed-rank test, not statistically significant, so the interpretive claims about binaural information capture exceed the evidence.
major comments (4)
- [Section 1 and Eq. (7)] The efficiency motivation is not supported by any measurement. Section 1 argues that self-attention is a computational bottleneck for low-latency, power-efficient hearing aids and that Mamba's linear-time, constant-memory inference makes it suitable for such devices, but the paper reports no runtime, memory, or power comparisons. Moreover, the frozen Whisper-Large-v2 feature extractor is executed for every input and likely dominates end-to-end inference cost, while the bidirectional Mamba in Eq. (7) requires two sequential passes over the sequence. The authors should either measure end-to-end inference latency and memory for the transformer and Mamba variants, or explicitly limit the efficiency claim to the temporal transform block rather than the full SIP pipeline.
- [Section 5.2, Table 3, and footnote 6] The claim that the bidirectional Mamba model 'significantly improved' over the monaural model and 'effectively embeds contextual and spatial information' is inconsistent with the paper's own statistical test: footnote 6 states that the Wilcoxon signed-rank test found no significant differences between the models. The mean RMSE difference of 0.15 points in Table 3 is within the range of trial-to-trial variability. The authors should report confidence intervals or paired effect sizes, and should temper the interpretation to numerical trends unless a properly powered significance test supports the stronger claim.
- [Tables 2 and 3] The tables report only mean RMSE and NCC values averaged over 3 and 15 trials, respectively, with no standard deviations or confidence intervals. Without dispersion measures, the reader cannot assess whether the reported differences, such as 28.25 vs. 29.45 in monaural Table 2 or 27.34 vs. 27.49 in binaural Table 3, are meaningful. At minimum, the authors should report standard deviations or per-trial intervals, especially since the paper already performs a Wilcoxon test that suggests high variability.
- [Section 5.3, Table 4] The conclusion that 'temporal fine structure component for speech perception... is widely distributed among Whisper features' is presented as a finding, but the experiment only shows that changing the pooling size has little effect on average RMSE in one dataset. This is a negative result without statistical testing or analysis of why the features behave this way. I recommend presenting Section 5.3 as an ablation observation and softening the mechanistic interpretation.
minor comments (5)
- [Section 1] The text contains a typo: 'Manba-based SIP model' should be 'Mamba-based SIP model'.
- [Section 2] Equation (3) is notationally imprecise: it writes '\bar{A}, \bar{B} = exp(\Delta A), \Delta B', but the discretization of a state-space model is not simply an exponential of the raw matrices. Clarify that this follows the zero-order hold discretization used in Mamba.
- [Section 3.1.1] The phrase 'applying an average pooling size p = 20' should read 'applying average pooling with a pooling size p = 20' to avoid ambiguity about what is pooled.
- [Section 4.2] The sentence 'All Whisper features were statistically normalized based on each train dataset' would benefit from specifying the normalization method (e.g., z-score with statistics computed on the training partition only).
- [Section 5.2] The text refers to 'CEC2.test.1' and 'CEC2.test.3' in the narrative; ensure the naming is consistent with the dataset partition labels used in Table 1 and elsewhere.
Circularity Check
No significant circularity: the paper is an empirical model comparison on a held-out public challenge dataset with no derivation that reduces to its own inputs.
full rationale
The paper is an empirical comparison of transformer-, Mamba-, and LSTM-based temporal transform blocks for non-intrusive binaural speech intelligibility prediction on the CPC2 dataset. The central claim is that the proposed Mamba-based model achieves competitive RMSE/NCC with a slightly smaller parameter count than the transformer baseline. This claim is supported by held-out test partitions (CEC2.test.1-3) that were not used for training, and the baseline is a public challenge system (E011) with public code and hyperparameters. There is no fitted parameter that is later renamed as a prediction, no definitional equivalence between the model output and the training target, and no uniqueness theorem or load-bearing self-citation that forces the reported result. The only self-citation is reference [22], co-authored by one of the present authors, used to present the standard Mamba formulation in Eqs. (1)-(5); this citation is not load-bearing because the same equations are also attributed to the original Mamba paper [21] and are standard background. The paper also reuses the baseline's hyperparameters, which is a methodological choice for fair comparison rather than a circular step. The skeptical concerns about unmeasured runtime/memory and the lack of statistical significance in the Wilcoxon test are validity or evidence-strength issues, not circularity: they do not show that any result is equivalent to its inputs by construction. Accordingly, the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- dropout rate for Mamba/LSTM blocks =
0.3
- temporal pooling size p =
20
- embedding dimension =
384
- Mamba state dimension N =
not reported
assumptions (3)
- domain assumption CPC2 subjective intelligibility scores are reliable ground truth for hearing-impaired listeners
- domain assumption Whisper-Large-v2 encoder features contain sufficient information for intelligibility prediction
- standard math Mamba selective SSM provides linear-time inference with constant memory
Cite this review
Pith. "Pith review of Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners." pith.science (2026). https://pith.science/paper/5QHARMES
@misc{pith2026250705729,
author = {Pith},
title = {Pith review of: Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners},
year = {2026},
howpublished = {\url{https://pith.science/paper/5QHARMES}},
note = {Machine review of arXiv:2507.05729}
}
read the original abstract
Speech intelligibility prediction (SIP) models have been used as objective metrics to assess intelligibility for hearing-impaired (HI) listeners. In the Clarity Prediction Challenge 2 (CPC2), non-intrusive binaural SIP models based on transformers showed high prediction accuracy. However, the self-attention mechanism theoretically incurs high computational and memory costs, making it a bottleneck for low-latency, power-efficient devices. This may also degrade the temporal processing of binaural SIPs. Therefore, we propose Mamba-based SIP models instead of transformers for the temporal processing blocks. Experimental results show that our proposed SIP model achieves competitive performance compared to the baseline while maintaining a relatively small number of parameters. Our analysis suggests that the SIP model based on bidirectional Mamba effectively captures contextual and spatial speech information from binaural signals.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction Speech intelligibility (SI) is an index used to evaluate how spo- ken words and sentences are heard under noisy and distorted conditions. This helps assess the listening abilities of hearing- impaired (HI) individuals and the effectiveness of speech en- hancement (SE) processes in hearing aids (HAs) [1, 2, 3]. Many SI prediction (SIP) models ...
-
[2]
Mamba The Mamba architecture [21] was proposed as a sequence-to- sequence model for mapping an input sequence x ∈ RD to y ∈ RD with an element-wise hidden state hd ∈ RN , which can be formulated as follows: [22]: ht,d = ¯Aht−1,d + ¯Bxt,d, (1) yt,d = Cht,d, (2) ¯A, ¯B = exp(∆A), ∆B, (3) where A ∈ RN ×N , B ∈ RN ×1, C ∈ R1×N and ∆ ∈ R+ rep- resent continuou...
work page Pith review arXiv 2025
-
[3]
Proposed Model 3.1. Monaural SIP Model 3.1.1. Overall Architecture Figure 1(a) shows the overall architecture of the monaural ver- sion of the E011 model [11]. The model is based on the network structure of a method called Whisper-AT [27], which was proposed to perform speech recognition and environmen- tal sound classification jointly. The model inputs a...
-
[4]
Experimental Setups 4.1. Code Implementations for Models We used code from the GitHub repository for CPC2 2 to train and test the baseline and our proposed models. To implement the baseline model, we used the public code of Whisper-AT [27]3, allowing feature extraction from encoder layers of Whis- per. We employed the official Mamba codebase and parame- t...
-
[5]
In the monaural experiments, we used the left channel of input speech data
Experiments and Results We conducted comparative monaural and binaural SIP experi- ments. In the monaural experiments, we used the left channel of input speech data. We compared the root-mean-squared er- ror (RMSE) and normalized cross-correlation coefficient (NCC) between the predicted and subjective SI results. 5.1. Monaural SIP models Table 2 shows the...
-
[6]
Conclusions In this paper, we propose a Mamba-based non-intrusive binau- ral SIP model for HI listeners. Our proposed model exhibits competitive performance compared with the transformer-based model under both monaural and binaural conditions, confirm- ing the effectiveness of Mamba in temporal transform blocks. Our analysis implies that the bidirectional...
-
[7]
H. Dillon, Hearing Aids , 2nd ed. New York, United States: Thieme, 2012
work page 2012
-
[8]
P. C. Loizou, Speech Enhancement: Theory and Practice, 2nd ed. Boca Raton, Florida, United States: CRC Press, 2013
work page 2013
Show all 37 references
-
[9]
Objective Quality and Intelligibility Prediction for Users of Assistive Listening Devices: Advantages and limitations of existing tools,
T. H. Falk, V . Parsa, J. F. Santos, K. Arehart, O. Hazrati, R. Huber, J. M. Kates, and S. Scollie, “Objective Quality and Intelligibility Prediction for Users of Assistive Listening Devices: Advantages and limitations of existing tools,” IEEE Signal Processing Maga- zine, vol...
2015
-
[10]
Speech intelligibility prediction based on modulation frequency-selective processing,
H. Rela ˜no-Iborra and T. Dau, “Speech intelligibility prediction based on modulation frequency-selective processing,” Hearing Research, p. 108610, 2022
2022
-
[11]
Nonintrusive objective measurement of speech intelligibility: A review of methodology,
Y . Feng and F. Chen, “Nonintrusive objective measurement of speech intelligibility: A review of methodology,”Biomedical Sig- nal Processing and Control, vol. 71, p. 103204, 2022
2022
-
[12]
ASR-based speech intelligibility prediction: A review,
M. Karbasi and D. Kolossa, “ASR-based speech intelligibility prediction: A review,” Hearing Research, vol. 426, p. 108606, 2022
2022
-
[13]
Clarity-2021 Challenges: Machine Learning Challenges for Advancing Hearing Aid Pro- cessing,
S. Graetzer, J. Barker, T. J. Cox, M. Akeroyd, J. F. Culling, G. Naylor, E. Porter, and R. V . Mu˜noz, “Clarity-2021 Challenges: Machine Learning Challenges for Advancing Hearing Aid Pro- cessing,” in Proc. INTERSPEECH 2021 – 22 nd Annual Con- ference of the International Spee...
2021
-
[14]
The 2nd Clar- ity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and Outcomes,
M. A. Akeroyd, W. Bailey, J. Barker, T. J. Cox, J. F. Culling, S. Graetzer, G. Naylor, Z. Podwi ´nska, and Z. Tu, “The 2nd Clar- ity Enhancement Challenge for Hearing Aid Speech Intelligibility Enhancement: Overview and Outcomes,” in Proc. ICASSP 2023 – 48th International Conf...
2023
-
[15]
The 1st Clarity Prediction Challenge: A ma- chine learning challenge for hearing aid intelligibility prediction,
J. Barker, M. Akeroyd, T. J. Cox, J. F. Culling, J. Firth, S. Graet- zer, H. Griffiths, L. Harris, G. Naylor, Z. Podwinska, E. Porter, and R. V . Munoz, “The 1st Clarity Prediction Challenge: A ma- chine learning challenge for hearing aid intelligibility prediction,” in Proc. ...
2022
-
[16]
The 2nd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing Aid Intel- ligibility Prediction,
J. Barker, M. A. Akeroyd, W. Bailey, T. J. Cox, J. F. Culling, J. Firth, S. Graetzer, and G. Naylor, “The 2nd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing Aid Intel- ligibility Prediction,” in Proc. ICASSP 2024 – 49 th International Conference on Acou...
2024
-
[17]
Temporal-hierarchical features from noise-robust speech foundation models for non-intrusive intelligi- bility prediction,
S. Cuervo and R. Marxer, “Temporal-hierarchical features from noise-robust speech foundation models for non-intrusive intelligi- bility prediction,” The 4th Clarity Workshop on Machine Learn- ing Challenges for Hearing Aids (Clarity-2023), Trinity Business School, Dublin, Irel...
2023
-
[18]
Attention is All you Need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems , vol. 30, 2017, pp. 1–11
2017
-
[19]
Robust Speech Recognition via Large-Scale Weak Supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Supervision,” 2022, arXiv:2212.04356
2022 arXiv
-
[20]
WavLM: Large- Scale Self-Supervised Pre-Training for Full Stack Speech Pro- cessing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large- Scale Self-Supervised Pre-Training for Full Stack Speech Pro- cessing,” IEEE Journal of Sele...
2022
-
[21]
The Hearing-Aid Speech Percep- tion Index (HASPI) Version 2,
J. M. Kates and K. H. Arehart, “The Hearing-Aid Speech Percep- tion Index (HASPI) Version 2,”Speech Communication, vol. 131, pp. 35–46, 2021
2021
-
[22]
Unsupervised Uncertainty Mea- sures of Automatic Speech Recognition for Non-intrusive Speech Intelligibility Prediction,
Z. Tu, N. Ma, and J. Barker, “Unsupervised Uncertainty Mea- sures of Automatic Speech Recognition for Non-intrusive Speech Intelligibility Prediction,” in Proc. INTERSPEECH 2022 – 23 rd Annual Conference of the International Speech Communication Association, Incheon, South Kor...
2022
-
[23]
MBI-Net: A Non-Intrusive Multi-Branched Speech Intelligibil- ity Prediction Model for Hearing Aids,
R. E. Zezario, F. Chen, C.-S. Fuh, H.-M. Wang, and Y . Tsao, “MBI-Net: A Non-Intrusive Multi-Branched Speech Intelligibil- ity Prediction Model for Hearing Aids,” in Proc. INTERSPEECH 2022 – 23rd Annual Conference of the International Speech Com- munication Association , Inche...
2022
-
[24]
Speech Foundation Models on In- telligibility Prediction for Hearing-Impaired Listeners,
S. Cuervo and R. Marxer, “Speech Foundation Models on In- telligibility Prediction for Hearing-Impaired Listeners,” in Proc. ICASSP 2024 – 49 th International Conference on Acoustics, Speech, and Signal Processing , Seoul, South Korea, Apr. 2024, pp. 1421–1425
2024
-
[25]
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory Models,
R. Mogridge, G. Close, R. Sutherland, T. Hain, J. Barker, S. Goetze, and A. Ragni, “Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users Using Intermediate ASR Features and Human Memory Models,” in Proc. ICASSP 2024 – 49th International Conference on Acou...
2024
-
[26]
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata,
R. E. Zezario, F. Chen, C.-S. Fuh, H.-M. Wang, and Y . Tsao, “Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata,” in Proc. INTERSPEECH 2024 – 25th Annual Conference of the International Speech Communica- tion Association, Kos Island, G...
2024
-
[27]
Mamba: Linear-time sequence modeling with selective state spaces,
A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” in COLM 2024 – 1 st Conference on Lan- guage Modeling, Pennsylvania, United States, Oct. 2024
2024
-
[28]
Exploring the Ca- pability of Mamba in Speech Applications,
K. Miyazaki, Y . Masuyama, and M. Murata, “Exploring the Ca- pability of Mamba in Speech Applications,” in Proc. INTER- SPEECH 2024 – 25 th Annual Conference of the International Speech Communication Association , Kos Island, Greece, Sep. 2024, pp. 237–241
2024
-
[29]
On the Parameteriza- tion and Initialization of Diagonal State Space Models,
A. Gu, K. Goel, A. Gupta, and C. R ´e, “On the Parameteriza- tion and Initialization of Diagonal State Space Models,”Advances in Neural Information Processing Systems , vol. 35, pp. 35 971– 35 983, Dec. 2022
2022
-
[30]
Simplified State Space Layers for Sequence Modeling,
J. T. H. Smith, A. Warrington, and S. Linderman, “Simplified State Space Layers for Sequence Modeling,” in Proc. ICLR 2023 – International Conference on Learning Representations, Kigali, Rwanda, Feb. 2023
2023
-
[31]
Prefix Sums and Their Applications,
G. E. Blelloch, “Prefix Sums and Their Applications,” School of Computer Science, Carnegie Mellon University, Tech. Rep. CMU-CS-90-190, Nov. 1990
1990
-
[32]
Hungry Hungry Hippos: Towards Language Modeling with State Space Models,
D. Y . Fu, T. Dao, K. K. Saab, A. W. Thomas, A. Rudra, and C. Re, “Hungry Hungry Hippos: Towards Language Modeling with State Space Models,” in Proc. ICLR 2023 – International Conference on Learning Representations, Kigali, Rwanda, Feb. 2023
2023
-
[33]
Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,
Y . Gong, S. Khurana, L. Karlinsky, and J. Glass, “Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,” in Proc. INTERSPEECH 2023 – 24th Annual Conference of the International Speech Communica- tion Association, Dublin, Ireland, 2...
2023
-
[34]
Dropout: A Simple Way to Prevent Neural Networks from Overfitting,
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Re- search, vol. 15, no. 56, pp. 1929–1958, 2014
1929
-
[35]
Long short-term memory,
S. Hochreiter, “Long short-term memory,” Neural Computation MIT-Press, 1997
1997
-
[36]
Adam: A Method for Stochastic Op- timization,
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Op- timization,” in Proc. ICLR 2015 – International Conference on Learning Representations , San Diego, USA, 2015, pp. 1930– 1958
2015
-
[37]
The roles of temporal envelope and fine struc- ture information in auditory perception,
B. C. J. Moore, “The roles of temporal envelope and fine struc- ture information in auditory perception,” Acoustical Science and Technology, vol. 40, no. 2, pp. 61–83, 2019
2019
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.