REVIEW 4 major objections 4 minor 54 references
Uncovering the role of semantic and acoustic cues in normal and dichotic listening
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper shows that in dichotic listening, EEG matches the attended story's text (82%) better than its sound envelope (77.45%), and the joint model reaches 83.76%.
desk verdict The unified MM framework is promising and the natural-listening results are solid, but the dichotic semantic-dominance claim is compromised by a subject-identity confound the authors do not control. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is STEM3, a three-branch sequence model. One branch encodes the speech envelope (a common proxy for acoustic neural tracking), one encodes the text of the speech using word2vec embeddings, and one encodes EEG; all three are pooled to word level using forced-alignment word boundaries, passed through LSTM or transformer layers, and compared via the negative L1 distance between stimulus and EEG embeddings. A single joint model is trained with a modality weight lambda that is randomly set to text-only, speech-only, or both, allowing the same model to produce all three accuracy numbers. The word-boundary average pooling is what lets the model align the acoustic and neural time courses at the rate of words rather than raw samples.
What would settle it
Replace each word's word2vec vector in the dichotic listening pipeline with a random vector drawn from the same embedding distribution, keeping word boundaries and word order unchanged. If match-mismatch accuracy stays near 82%, the semantic advantage is an artifact of word-level alignment; if it drops toward the envelope's 77%, the EEG is genuinely tracking word meaning.
Extended reading notes
Core claim
The central discovery is that in a demanding two-talker (dichotic) listening situation, EEG signals carry a stronger signature of the attended story's text than of its acoustic envelope, while in ordinary natural listening the two are comparable. The paper establishes this through a unified match-mismatch model (STEM3) with separate acoustic, semantic, and EEG branches, and reports that the semantic branch reaches 82.00% match-mismatch accuracy versus 77.45% for the envelope branch; the multimodal branch reaches 83.76%. It further finds that accurate word boundaries matter for all modalities (random boundaries do not help), and that subjects attending to the right ear are decoded more accurately than those attending to the left ear, consistent with a right-ear advantage. The authors interpret these results as evidence that the brain prioritizes semantic content over low-level acoustics under competing-talker conditions.
Load-bearing premise
The argument assumes that word2vec embeddings of the attended story's transcript capture the semantic cues the EEG is actually encoding, and that the unattended story's embeddings serve as a genuinely mismatched semantic stimulus; if word2vec is not a faithful stand-in for the brain's semantics, the measured semantic advantage could be an artifact of word-level timing rather than meaning.
Editorial extensions
If this is right
- Under dichotic listening, the match-mismatch accuracy of the multimodal model (83.76%) can serve as a segment-level estimate of how well the attended stream can be identified from EEG alone, a relevant signal for neuro-steered hearing aids.
- Since random word boundaries do not reproduce the benefit of true boundaries, any practical decoder should use accurate word segmentation, not just fixed windows.
- In natural listening, acoustic and semantic cues are interchangeable for this task, so models that drop one modality may still match performance in easy conditions.
- The right-ear advantage appears not only in behavior but also in the EEG match-mismatch accuracy, which could make attention decoders more reliable when the attended ear is the right ear.
Reading between the lines
- Because word2vec embeddings encode word identity and co-occurrence statistics as well as meaning, the semantic advantage in dichotic listening may partly reflect lexical or positional alignment; a control with meaning-scrambled word vectors would separate true semantic encoding from word-level timing.
- The trial selection removed 35% of dichotic trials based on comprehension scores, so the reported accuracies may overstate performance for inattentive or distracted listeners; testing on the excluded trials would quantify that gap.
- The same three-branch architecture could be extended to other acoustic features (pitch, phonetic features) and to contextual embeddings (e.g., transformer-based), which would map a fuller hierarchy of what EEG tracks in competing speech.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a match-mismatch (MM) classification framework, STEM3, to compare how acoustic (speech envelope) and semantic (word2vec text embeddings) features of speech are encoded in EEG, under both natural and dichotic listening. The authors report that in natural listening (Dataset Small) acoustic and semantic cues yield similar MM accuracy (93.63% vs 93.24%), that in dichotic listening semantic cues outperform acoustic cues (82.00% vs 77.45%), that a multimodal combination reaches 83.76% on the dichotic task, that word boundary information significantly helps, and that a right-ear advantage is observed. The proposed model is compared with several prior MM models, and the authors claim large relative improvements.
Significance. If the findings are robust, the study offers a useful quantitative benchmark for comparing the neural encoding of acoustic and semantic speech features, and it extends the match-mismatch paradigm to dichotic listening, a relevant direction for auditory attention decoding. Strengths of the paper are its use of public datasets, a clearly specified model architecture, and several ablation studies (word boundaries, brain regions, loss functions). However, the central dichotic-listening claim (semantic over acoustic advantage) is threatened by a subject-identity confound, because attention side is fixed per subject and the MM task can be solved by recognizing the subject's constant attended story rather than by time-resolved stimulus-response matching. The behavioral trial-selection step and the baseline comparisons further weaken the evidence. These issues need to be addressed before the specific conclusions can be considered reliable.
major comments (4)
- [II-A, II-D, IV-C, Table I] In the dichotic dataset each participant is assigned to attend either the left or right ear for all 30 trials, so the attended story, the unattended story, and the speaker identities are constant within each subject. The MM task (Section II-D) defines the positive pair as the attended story segment and the negative pair as the unattended story segment. Because the training and test folds are split by trials rather than by subjects (Section II-H, 'Training and Evaluation Setup'), the model can exploit a subject-specific signature (e.g., overall EEG characteristics, or the speaker identity of the attended story) to classify correctly without performing time-resolved matching between the stimulus and the EEG. The reported semantic advantage (82.00% vs 77.45% in Table I) may therefore reflect that the two stories' word2vec text representations are more separable than the two male speakers' envelopes, rather than a stronger neural encoding of semantics. The paper does not provide any control experiment, such as leave-subject-out evaluation or a within-subject epoch-shuffle test, and the Limitations section (IV-F) does not list this confound. This is load-bearing for the abstract claim (c) that semantic cues are significantly better in dichotic listening, so it must be addressed.
- [II-H, III-B] The trial-selection procedure removes 35% of dichotic trials based on behavioral comprehension thresholds (attended score >60% and unattended score <40%). Because this selection is applied before any train/test split, the reported accuracies are for a subset of trials that are behaviorally 'clean', which can inflate MM accuracy and may also differentially affect the acoustic and semantic conditions. More importantly, the selection uses information from the test trials themselves (their post-hoc comprehension scores), which is not available in a real-world auditory attention decoding setting; the results therefore generalize only to trials where the listener demonstrably attended as instructed. The paper should report results on the full trial set, or at least justify why the selected subset is the appropriate evaluation population.
- [Table I, Section III-B] The claimed relative improvements over baselines are not consistent with the numbers in Table I. For example, for the Natural Dataset - Small, the acoustic STEM3 result of 93.63% versus the baseline [12] of 65.23% corresponds to a relative improvement of (93.63-65.23)/65.23 = 43.5%, not the reported 79%; for the Dichotic Dataset, (77.45-62.00)/62.00 = 24.9%, not 50.73%. The reader cannot reproduce the stated improvements from the reported numbers. Additionally, the Wang et al. [17] results on Natural Dataset - Small (36.33%) and Natural Dataset - Large (50.33%) are at or below chance (50%), which is suspicious and suggests that the comparison protocol may not be matched (e.g., different segment durations, trial selection, or model inputs). The paper should clarify the exact definition of 'relative improvement' and ensure that baselines are evaluated under the same train/test scheme and input features.
- [III-A, II-I, Figure 4] The choice of the similarity function and the hyperparameter λ (modality mixing weight) is made using the dichotic listening data (Figure 4), and the best configuration (Sim1 with λ=0.5) is then used for all subsequent experiments, including the final reported accuracies. If this selection is performed on the same data that is later used for evaluation, the reported results on the dichotic dataset are partly the result of tuning on the test set. The authors should state whether the λ and similarity-function selection was done on a separate validation set or within each cross-validation fold, and if it was not, they should acknowledge the potential optimistic bias.
minor comments (4)
- [II-H] The description of the trial-selection figure (Figure 3) says 'incorrectly responding to the unattended side', but the selection criterion in the text is a score of 40% or less rather than 'incorrect'. Please make the wording consistent.
- [II-G, Table II] The 'Without Word Boundary' ablation is described only as not using word boundaries, but it is unclear what temporal pooling scheme is used instead (e.g., fixed-length pooling, no pooling, or frame-level classification). Since word2vec features are inherently word-level, it is also unclear what 'without word boundary' means for the text branch. Please specify the exact architecture used in this ablation.
- [IV-B] The sentence 'In natural listening condition (large), we find the acoustic features to be significantly better than textual features' should be rephrased as 'we find that the acoustic features...' for clarity. Also, the discussion should note that the natural-listening conclusion of 'similar' levels is based only on Dataset Small, since the large dataset shows a significant acoustic advantage.
- [II-F] The word2vec embeddings are described as capturing semantics of words, but the paper does not discuss that these embeddings are pretrained on Google News and may not reflect the specific narrative context of the audiobook stories. This is a model choice that should be acknowledged when interpreting 'semantic cues'.
Circularity Check
No circularity in the STEM3 derivation; the dichotic attended-side confound is a validity risk, not a circular step.
full rationale
The paper's derivation chain is empirical and self-contained against external benchmarks. The MM task uses public speech-EEG datasets, external word2vec embeddings, and held-out trial folds; no parameter is fitted to a subset and then reported as a prediction, and no result is defined in terms of the conclusion it supports. The only self-citations (Soman et al. [5], [22]) are used as background or architectural motivation, and the word-boundary contribution is independently tested by ablations that compare no boundaries, random boundaries, and true word boundaries, so the self-citation is not load-bearing. The natural-dataset results are also not forced: on SparrKULee, acoustic features significantly outperform text features, which shows the framework can produce modality differences in either direction. The strongest concern in the dichotic design, that each subject's attended side is fixed across all trials (Section II-A) so a model could classify the constant attended-story label via subject-specific EEG signatures, is a validity/shortcut confound rather than a circularity: there is no equation in the paper where an output is identical to an input by construction, and no fitted value is renamed as a prediction. Therefore no circular step is present.
Assumptions & free parameters
free parameters (3)
- lambda (modality mixing weight) =
0.5 chosen as best; sampled from {0.0, 0.5, 1.0} during training
- Trial selection thresholds =
attended score > 60%, unattended score < 40%
- EEG downsampling to 64 Hz and word-pooling at 64/3 Hz =
64 Hz
assumptions (3)
- domain assumption word2vec embeddings of the transcribed attended story reliably represent the attended 'semantic' stream
- domain assumption The match-mismatch paradigm measures stimulus-response encoding rather than purely label-driven artifacts
- domain assumption The Broderick and SparrKULee data, including their word alignment and sentence segmentation, are accurate
invented entities (1)
-
STEM3 architecture
Cite this review
Pith. "Pith review of Uncovering the role of semantic and acoustic cues in normal and dichotic listening." pith.science (2026). https://pith.science/paper/BNOYF2PG
@misc{pith2026241111308,
author = {Pith},
title = {Pith review of: Uncovering the role of semantic and acoustic cues in normal and dichotic listening},
year = {2026},
howpublished = {\url{https://pith.science/paper/BNOYF2PG}},
note = {Machine review of arXiv:2411.11308}
}
read the original abstract
Speech comprehension is an involuntary task for the healthy human brain, yet the understanding of the mechanisms underlying this brain functionality remains obscure. In this paper, we aim to quantify the role of acoustic and semantic information streams in complex listening conditions. We propose a paradigm to understand the encoding of the speech cues in electroencephalogram (EEG) data, by designing a match-mismatch (MM) classification task. The MM task involves identifying whether the stimulus (speech) and response (EEG) correspond to each other. We build a multimodal deep-learning based sequence model STEM, which is input with acoustic stimulus (speech envelope), semantic stimulus (textual representations of speech), and the neural response (EEG data). We perform extensive experiments on two separate conditions, i) natural passive listening and, ii) a dichotic listening requiring auditory attention. Using the MM task as the analysis framework, we observe that - a) speech perception is fragmented based on word boundaries, b) acoustic and semantic cues offer similar levels of MM task performance in natural listening conditions, and c) semantic cues offer significantly improved MM classification over acoustic cues in dichotic listening task. The comparison of the STEM with previously proposed MM models shows significant performance improvements for the proposed approach. The analysis and understanding from this study allows the quantification of the roles played by acoustic and semantic cues in diverse listening tasks and in providing further evidences of right-ear advantage in dichotic listening.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[12]
An LSTM based architecture to relate speech stimulus to EEG,
M. J. Monesi, B. Accou, J. Montoya-Martinez, T. Francart, and H. Van Hamme, “An LSTM based architecture to relate speech stimulus to EEG,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 941–945. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, JUNE 2025 10
work page 2020
-
[17]
B. Wang, X. Xu, Z. Zhang, H. Zhu, Y . Yan, X. Wu, and J. Chen, “Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording,” in 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW). IEEE, 2024, pp. 111–112
work page 2024
-
[1]
Low-frequency cor- tical entrainment to speech reflects phoneme-level processing,
G. M. Di Liberto, J. A. O’sullivan, and E. C. Lalor, “Low-frequency cor- tical entrainment to speech reflects phoneme-level processing,” Current Biology, vol. 25, no. 19, pp. 2457–2465, 2015
work page 2015
-
[2]
Speech rhythms and their neural foundations,
D. Poeppel and M. F. Assaneo, “Speech rhythms and their neural foundations,” Nature Reviews Neuroscience , vol. 21, no. 6, pp. 322– 334, 2020
work page 2020
-
[3]
S. Sanei and J. A. Chambers, EEG Signal Processing . John Wiley & Sons, 2013
work page 2013
-
[4]
M. P. Broderick, A. J. Anderson, G. M. Di Liberto, M. J. Crosse, and E. C. Lalor, “Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narrative speech,” Current Biology, vol. 28, no. 5, pp. 803–809, 2018
work page 2018
-
[5]
An EEG study on the brain representations in language learning,
A. Soman, C. Madhavan, K. Sarkar, and S. Ganapathy, “An EEG study on the brain representations in language learning,” Biomedical Physics & Engineering Express , vol. 5, no. 2, p. 025041, 2019
work page 2019
-
[6]
Neural coding of continuous speech in auditory cortex during monaural and dichotic listening,
N. Ding and J. Z. Simon, “Neural coding of continuous speech in auditory cortex during monaural and dichotic listening,” Journal of Neurophysiology, vol. 107, no. 1, pp. 78–89, 2012
work page 2012
Show all 54 references
-
[7]
The multivariate temporal response function (mTRF) toolbox: a MATLAB toolbox for relating neural signals to continuous stimuli,
M. J. Crosse, G. M. Di Liberto, A. Bednar, and E. C. Lalor, “The multivariate temporal response function (mTRF) toolbox: a MATLAB toolbox for relating neural signals to continuous stimuli,” Frontiers in Human Neuroscience, vol. 10, p. 604, 2016
2016
-
[8]
A comparison of regularization methods in forward and backward models for auditory attention decoding,
D. D. Wong, S. A. Fuglsang, J. Hjortkjær, E. Ceolini, M. Slaney, and A. De Cheveigne, “A comparison of regularization methods in forward and backward models for auditory attention decoding,” Frontiers in Neuroscience, vol. 12, p. 531, 2018
2018
-
[9]
Speech intelligibility predicted from neural entrainment of the speech envelope,
J. Vanthornhout, L. Decruy, J. Wouters, J. Z. Simon, and T. Francart, “Speech intelligibility predicted from neural entrainment of the speech envelope,” Journal of the Association for Research in Otolaryngology , vol. 19, pp. 181–191, 2018
2018
-
[10]
Decoding the auditory brain with canonical component analysis,
A. de Cheveign ´e, D. D. Wong, G. M. Di Liberto, J. Hjortkjær, M. Slaney, and E. Lalor, “Decoding the auditory brain with canonical component analysis,” NeuroImage, vol. 172, pp. 206–216, 2018
2018
-
[11]
Is there chaos in the brain? i. concepts of nonlinear dynamics and methods of investigation,
P. Faure and H. Korn, “Is there chaos in the brain? i. concepts of nonlinear dynamics and methods of investigation,” Comptes Rendus de l’Acad´emie des Sciences-Series III-Sciences de la Vie , vol. 324, no. 9, pp. 773–793, 2001
2001
-
[13]
Deep correlation analysis for audio-EEG decoding,
J. R. Katthi and S. Ganapathy, “Deep correlation analysis for audio-EEG decoding,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 29, pp. 2742–2753, 2021
2021
-
[14]
Audi- tory stimulus-response modeling with a match-mismatch task,
A. De Cheveign ´e, M. Slaney, S. A. Fuglsang, and J. Hjortkjaer, “Audi- tory stimulus-response modeling with a match-mismatch task,” Journal of Neural Engineering , vol. 18, no. 4, p. 046040, 2021
2021
-
[15]
Modeling the relationship between acoustic stimulus and EEG with a dilated convolutional neural network,
B. Accou, M. J. Monesi, J. Montoya, T. Francart et al. , “Modeling the relationship between acoustic stimulus and EEG with a dilated convolutional neural network,” in2020 28th European Signal Processing Conference (EUSIPCO). IEEE, 2021, pp. 1175–1179
2021
-
[16]
Relating the fundamental frequency of speech with EEG using a dilated convo- lutional network,
C. Puffay, J. Van Canneyt, J. Vanthornhout, T. Francart et al., “Relating the fundamental frequency of speech with EEG using a dilated convo- lutional network,” arXiv preprint arXiv:2207.01963 , 2022
2022 arXiv
-
[18]
Decoding the attended speech stream with multi-channel EEG: implications for online, daily-life applications,
B. Mirkovic, S. Debener, M. Jaeger, and M. De V os, “Decoding the attended speech stream with multi-channel EEG: implications for online, daily-life applications,” Journal of Neural Engineering , vol. 12, no. 4, p. 046007, 2015
2015
-
[19]
Decoding envelope and frequency-following EEG responses to continuous speech using deep neural networks,
M. Thornton, D. P. Mandic, and T. Reichenbach, “Decoding envelope and frequency-following EEG responses to continuous speech using deep neural networks,” IEEE Open Journal of Signal Processing , 2024
2024
-
[20]
Detecting gamma-band responses to the speech envelope for the ICASSP 2024 auditory EEG decoding signal processing grand chal- lenge,
M. Thornton, J. Auernheimer, C. Jehn, D. Mandic, and T. Reichenbach, “Detecting gamma-band responses to the speech envelope for the ICASSP 2024 auditory EEG decoding signal processing grand chal- lenge,” in 2024 IEEE International Conference on Acoustics, Speech, and Signal Pr...
2024
-
[21]
Multi- head attention and GRU for improved match-mismatch classification of speech stimulus and EEG response,
M. Borsdorf, S. Pahuja, G. Ivucic, S. Cai, H. Li, and T. Schultz, “Multi- head attention and GRU for improved match-mismatch classification of speech stimulus and EEG response,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASS...
2023
-
[22]
Enhancing the EEG speech match mismatch tasks with word boundaries,
A. Soman, V . Sinha, and S. Ganapathy, “Enhancing the EEG speech match mismatch tasks with word boundaries,” Proceedings of Inter- speech 2023, pp. 526–530, 2023
2023
-
[23]
Semantic processing of unattended speech in dichotic listening,
J. Aydelott, Z. Jamaluddin, and S. Nixon Pearce, “Semantic processing of unattended speech in dichotic listening,”The Journal of the Acoustical Society of America , vol. 138, no. 2, pp. 964–975, 2015
2015
-
[24]
Top-down modulation of dichotic listening affects inter- hemispheric connectivity: an electroencephalography study,
O. Elyamany, J. Iffland, D. Lockhofen, S. Steinmann, G. Leicht, and C. Mulert, “Top-down modulation of dichotic listening affects inter- hemispheric connectivity: an electroencephalography study,” Frontiers in Neuroscience, vol. 18, p. 1424746, 2024
2024
-
[25]
Sex-related differences in selective auditory attention in dichotic listening with different levels of difficulty: fMRI data,
L. Mayorova and A. Kushnir, “Sex-related differences in selective auditory attention in dichotic listening with different levels of difficulty: fMRI data,” Neuroscience and Behavioral Physiology , vol. 54, no. 1, pp. 102–111, 2024
2024
-
[26]
The development of the right ear advantage in dichotic listening with focused attention,
G. Geffen, “The development of the right ear advantage in dichotic listening with focused attention,” Cortex, vol. 14, no. 2, pp. 169–177, 1978
1978
-
[27]
Neurophysiological evaluation of right-ear advantage during dichotic listening,
K. Tanaka, B. Ross, S. Kuriki, T. Harashima, C. Obuchi, and H. Okamoto, “Neurophysiological evaluation of right-ear advantage during dichotic listening,” Frontiers in psychology, vol. 12, p. 696263, 2021
2021
-
[28]
Human-robot interaction through real-time auditory and visual multiple-talker tracking,
H. G. Okuno, K. Nakadai, K. I. Hidai, H. Mizoguchi, and H. Ki- tano, “Human-robot interaction through real-time auditory and visual multiple-talker tracking,” in Proceedings 2001 IEEE/RSJ International Conference on Intelligent Robots and Systems. Expanding the Societal Role o...
2001
-
[29]
Selection of the closest sound source for robot auditory attention in multi-source scenarios,
Q. Nguyen and J. Choi, “Selection of the closest sound source for robot auditory attention in multi-source scenarios,” Journal of Intelligent & Robotic Systems, vol. 83, pp. 239–251, 2016
2016
-
[30]
Hearing device with brain-wave dependent audio process- ing
T. Lunner, “Hearing device with brain-wave dependent audio process- ing.” 2013
2013
-
[31]
Global hearing health care: new findings and perspectives,
B. S. Wilson, D. L. Tucci, M. H. Merson, and G. M. O’Donoghue, “Global hearing health care: new findings and perspectives,” The Lancet, vol. 390, no. 10111, pp. 2503–2515, 2017
2017
-
[32]
Cochlear implants: system design, integration, and evaluation,
F.-G. Zeng, S. Rebscher, W. Harrison, X. Sun, and H. Feng, “Cochlear implants: system design, integration, and evaluation,” IEEE Reviews in Biomedical Engineering, vol. 1, pp. 115–142, 2008
2008
-
[33]
The silent impact of hearing loss: Using longitudinal data to explore the effects on depression and social activity restriction among older people,
C. C. Andrade, C. R. Pereira, and P. A. Da Silva, “The silent impact of hearing loss: Using longitudinal data to explore the effects on depression and social activity restriction among older people,” Ageing & Society , vol. 38, no. 12, pp. 2468–2489, 2018
2018
-
[34]
EEG-based attention-driven speech enhancement for noisy speech mixtures using N-fold multi-channel Wiener filters,
N. Das, S. Van Eyndhoven, T. Francart, and A. Bertrand, “EEG-based attention-driven speech enhancement for noisy speech mixtures using N-fold multi-channel Wiener filters,” in 2017 25th European Signal Processing Conference (EUSIPCO). IEEE, 2017, pp. 1660–1664
2017
-
[35]
EEG-informed attended speaker extraction from recorded speech mixtures with ap- plication in neuro-steered hearing prostheses,
S. Van Eyndhoven, T. Francart, and A. Bertrand, “EEG-informed attended speaker extraction from recorded speech mixtures with ap- plication in neuro-steered hearing prostheses,” IEEE Transactions on Biomedical Engineering, vol. 64, no. 5, pp. 1045–1056, 2016
2016
-
[36]
Neuro-steered hearing devices: decoding auditory attention from the brain,
S. Geirnaert, S. Vandecappelle, E. Alickovic, A. de Cheveign ´e, E. Lalor, B. T. Meyer, S. Miran, T. Francart, and A. Bertrand, “Neuro-steered hearing devices: decoding auditory attention from the brain,” arXiv e- prints, pp. arXiv–2008, 2020
2008
-
[37]
EEG-based auditory attention detec- tion and its possible future applications for passive BCI,
J. Belo, M. Clerc, and D. Sch ¨on, “EEG-based auditory attention detec- tion and its possible future applications for passive BCI,” Frontiers in Computer Science, vol. 3, p. 661178, 2021
2021
-
[38]
Attention enhancement system using virtual reality and EEG biofeedback,
B. H. Cho, J.-M. Lee, J. Ku, D. P. Jang, J. Kim, I.-Y . Kim, J.-H. Lee, and S. I. Kim, “Attention enhancement system using virtual reality and EEG biofeedback,” in Proceedings IEEE Virtual Reality 2002 . IEEE, 2002, pp. 156–163
2002
-
[39]
SparrKULee: A speech-evoked auditory response reposi- tory of the KU Leuven, containing EEG of 85 participants,
B. Accou, L. Bollens, M. Gillis, W. Verheijen, H. Van Hamme, and T. Francart, “SparrKULee: A speech-evoked auditory response reposi- tory of the KU Leuven, containing EEG of 85 participants,” BioRxiv, pp. 2023–07, 2023
2023
-
[40]
Prosodylab-aligner: A tool for forced alignment of laboratory speech,
K. Gorman, J. Howell, and M. Wagner, “Prosodylab-aligner: A tool for forced alignment of laboratory speech,” Canadian Acoustics , vol. 39, no. 3, pp. 192–193, 2011
2011
-
[41]
EEGLAB: an open source toolbox for analysis of single-trial EEG dynamics including independent component analysis,
A. Delorme and S. Makeig, “EEGLAB: an open source toolbox for analysis of single-trial EEG dynamics including independent component analysis,” Journal of Neuroscience Methods , vol. 134, no. 1, pp. 9–21, 2004
2004
-
[42]
Extracting different levels of speech information from EEG using an LSTM-based model,
M. Jalilpour Monesi, B. Accou et al. , “Extracting different levels of speech information from EEG using an LSTM-based model,” Proceed- ings of Interspeech 2021 , pp. 526–530, 2021
2021
-
[43]
Speech recognition with primarily temporal cues,
R. V . Shannon, F.-G. Zeng, V . Kamath, J. Wygonski, and M. Ekelid, “Speech recognition with primarily temporal cues,” Science, vol. 270, no. 5234, pp. 303–304, 1995
1995
-
[44]
Chimaeric sounds reveal dichotomies in auditory perception,
Z. M. Smith, B. Delgutte, and A. J. Oxenham, “Chimaeric sounds reveal dichotomies in auditory perception,” Nature, vol. 416, no. 6876, pp. 87– 90, 2002
2002
-
[45]
Cortical measures of phoneme-level speech encoding correlate with the perceived clarity of natural speech,
G. M. Di Liberto, M. J. Crosse, and E. C. Lalor, “Cortical measures of phoneme-level speech encoding correlate with the perceived clarity of natural speech,” Eneuro, vol. 5, no. 2, 2018
2018
-
[46]
Neural speech tracking in the theta and in the delta frequency band differentially encode clarity and comprehension of speech in noise,
O. Etard and T. Reichenbach, “Neural speech tracking in the theta and in the delta frequency band differentially encode clarity and comprehension of speech in noise,” Journal of Neuroscience, vol. 39, no. 29, pp. 5750– 5759, 2019
2019
-
[47]
Efficient estimation of word representations in vector space,
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” arXiv preprint arXiv:1301.3781 , 2013
2013 arXiv
-
[48]
Distributed representations of words and phrases and their composi- tionality,
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their composi- tionality,” Advances in Neural Information Processing Systems , vol. 26, 2013
2013
-
[49]
Combining local and global features into a Siamese network for sentence similarity,
Y . Li, D. Zhou, and W. Zhao, “Combining local and global features into a Siamese network for sentence similarity,” IEEE Access , vol. 8, pp. 75 437–75 447, 2020
2020
-
[50]
Adam: A method for stochastic optimization,
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980 , 2014
2014 arXiv
-
[51]
Statistical comparisons of classifiers over multiple data sets,
J. Dem ˇsar, “Statistical comparisons of classifiers over multiple data sets,” The Journal of Machine learning research , vol. 7, pp. 1–30, 2006
2006
-
[52]
fMRI speech tracking in primary and non-primary auditory cortex while listening to noisy scenes,
L. Hausfeld, I. M. Hamers, and E. Formisano, “fMRI speech tracking in primary and non-primary auditory cortex while listening to noisy scenes,” Communications biology, vol. 7, no. 1, p. 1217, 2024
2024
-
[53]
Temporal dynamics of selective attention during dichotic listening,
B. Ross, S. A. Hillyard, and T. W. Picton, “Temporal dynamics of selective attention during dichotic listening,” Cerebral Cortex, vol. 20, no. 6, pp. 1360–1371, 2010
2010
-
[54]
Combined fmri region-and network- analysis reveal new insights of top-down modulation of bottom-up processes in auditory laterality,
K. Kazimierczak, A. R. Craven, L. Ersland, K. Specht, M. L. Dumitru, L. B. Sandøy, and K. Hugdahl, “Combined fmri region-and network- analysis reveal new insights of top-down modulation of bottom-up processes in auditory laterality,” Frontiers in Behavioral Neuroscience , vol....
2022
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.