Pith. sign in

REVIEW 4 major objections 3 minor 44 references

SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding

T0 review · 4 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Contrastive matching between SEEG brain signals and audio decodes Mandarin words above chance, and a single sensorimotor electrode matches the full array.

desk verdict Useful new Mandarin SEEG dataset and a sensible contrastive decoding demo, but the unreported acoustic contamination check leaves the neural claim unproven. read the letter →

arxiv 2505.19652 v1 pith:Z3IEN6NX submitted 2025-05-26 cs.HC cs.SDeess.AS

classification cs.HCcs.SDeess.AS
keywords stereo-electroencephalographyspeechdecodingBCIcontrastivelearningMandarinChinesesensorimotorcortexdetectionintracranialrecordingsword-level
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that stereo-electroencephalography (SEEG), the depth electrodes implanted for epilepsy monitoring, can decode spoken Mandarin words from brain activity alone. The authors collected a new dataset from eight epilepsy patients reading 48 Chinese words while SEEG and audio were recorded simultaneously, and they propose a contrastive matching method that aligns SEEG representations with audio representations. They report decoding accuracies above chance in both speech/non-speech detection and 48-way word decoding, and an electrode-by-electrode analysis in which a single sensorimotor cortex electrode performs roughly as well as the full electrode array. If these claims hold, they imply that Mandarin speech decoding is feasible with relatively sparse intracranial coverage, and that audio can serve as a natural training signal for brain decoders.

What carries the argument

The central object is SACM, a two-branch contrastive matching framework: a dilated residual convolutional network maps SEEG segments into embeddings, while a frozen pretrained self-supervised speech encoder maps audio segments into the same embedding space. Training uses a symmetric InfoNCE loss with a temperature parameter to pull matching SEEG-audio pairs together and push non-matching pairs apart, and decoding is nearest-neighbor retrieval by cosine similarity between the test SEEG embedding and candidate audio embeddings. Speech detection is handled separately by a compact convolutional classifier operating on high-gamma band envelopes. The contrastive objective is what lets audio act as a rich teaching signal for the sparse neural data.

What would settle it

Apply the acoustic contamination analysis to the raw SEEG and report the correlation between SEEG and microphone signal during speech and silence; if SEEG channels track the audio waveform, or if decoding remains above chance when audio is swapped or misaligned, the result would reflect leaked sound rather than neural speech activity.

Watch

Extended reading notes

Core claim

The central claim is that SEEG signals contain decodable information for word-level Mandarin speech, and that a contrastive learning framework called SACM can extract it by matching neural representations to audio representations. In the speech detection task, the full electrode array reached an average accuracy near 71 percent against a random baseline near 50 percent. In the 48-word decoding task, the full array reached top-5 accuracy near 16.8 percent against a chance level near 10.1 percent, and a single sensorimotor cortex electrode matched or exceeded the full array for several subjects. The paper further links decoding success to the discriminability of the patient's own speech features, suggesting that pronunciation intention modulates the neural signal.

Load-bearing premise

The central claim assumes the recorded brain signals are free of audio leakage from the microphone, since the paper says it checked for contamination but does not report what the check found.

Editorial extensions

If this is right

  • Mandarin Chinese word-level speech decoding from SEEG is feasible, extending speech neuroprosthesis results beyond non-tonal languages.
  • A single sensorimotor cortex electrode may be sufficient for basic speech detection and word decoding, which would make future implants less invasive and easier to place.
  • Synchronized audio can serve as a free supervisory signal for training neural decoders via contrastive matching, reducing reliance on hand-labeled trials.
  • Decoding performance varies strongly across individuals, with pronunciation clarity and electrode coverage as limiting factors, so patient-specific calibration will remain important.
  • The new dataset provides a testbed for mapping Mandarin-specific and language-general speech regions in the human brain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension beyond the paper: because electrode placement was driven by clinical needs rather than speech regions, the single-sensorimotor-electrode result may be strongest for patients whose monitoring happens to cover that region, and a prospective study with SMC-targeted implants would test it directly.
  • The same contrastive setup could be applied to imagined or attempted speech, where the patient does not vocalize, by using audio only as a pretraining target and retrieving from the shared audio-embedding space at inference; the paper does not attempt this.
  • The observed link between speech-feature discriminability and decoding accuracy suggests a correction step: normalizing or filtering audio features by articulation clarity could reveal how much of the neural signal is genuinely motor rather than acoustic.
  • A stronger validation would be to show decoding survives when the audio stream is withheld or adversarially distorted at test time, proving the model relies on neural rather than leaked acoustic information.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper introduces HUST-MIND, a new stereo-electroencephalography (SEEG) and synchronized audio dataset collected from eight drug-resistant epilepsy patients while they read 48 Mandarin Chinese monosyllabic words. It proposes SACM, a CLIP-inspired contrastive learning framework that maps SEEG segments and pre-trained HuBERT audio features into a shared space and retrieves the matching audio segment at test time. The authors report speech detection results using EEGNet and word/initial decoding results using SACM, claiming accuracies significantly above shuffled-label chance levels, and they further claim that a single sensorimotor cortex (SMC) electrode achieves performance comparable to the full electrode array. The manuscript includes ethical approval details, a code repository link, and a dataset available upon request.

Significance. If the central claims hold, this is a useful contribution to intracranial speech decoding: it provides a new tonal-language SEEG dataset, demonstrates a contrastive multimodal decoding approach, and offers evidence about the role of the sensorimotor cortex in Mandarin speech production. The paper's strengths include within-subject held-out evaluation, shuffled-label baselines, paired t-tests for the full-array versus random comparison, and public code. However, the load-bearing acoustic contamination check is asserted but never reported, the 'comparable' single-electrode claim is not supported statistically, and the statistical reporting is too thin to establish the significance claims as stated. The contribution is therefore plausible but not yet fully supported.

major comments (4)
  1. [§3.2.1, §3.2.2] The acoustic contamination check is asserted but its outcome is never reported: no per-subject or per-channel contamination metrics, no statement of whether any channels were flagged or removed, and no analysis of decoding results after excluding contaminated channels. Because all SEEG segments are time-locked to vocalized audio, an unreported acoustic leak into the recordings would allow the speech detection and SACM retrieval pipelines to succeed without any neural decoding. Please report the actual contamination assessment results and, if channels were removed, repeat Tables 3–5 with and without those channels.
  2. [§5.2, Tables 3–5] The claim that a single SMC electrode performs comparably to the full array is not supported by the reported statistics. The SMC row averages are computed over the six subjects with SMC electrodes, while the FULL row averages are computed over all eight subjects; for example, in Table 3 (CAR) SMC=70.82 from n=6 versus FULL=71.09 from n=8. The reported mean absolute differences therefore conflate subject subsets and ignore between-seed variability. Please compare SMC and FULL on the identical subject subset, with standard deviations or confidence intervals across the six random seeds, and provide an equivalence test (e.g., two one-sided tests) or at least paired differences with confidence intervals.
  3. [§5.3, Tables 4–5] The statistical evidence for 'significantly exceeding chance' is under-reported. The asterisks appear only on the average FULL row, with no p-values, no per-subject significance, no correction for multiple comparisons, and no effect sizes. Given n=8 and the visible between-subject variability (Table 4 CAR FULL ranges from 8.16% to 30.30%), the single aggregate paired t-test is not sufficient to support the abstract's unqualified claim. Report mean ± SD over seeds, per-subject or at least per-table p-values, and clarify whether the test uses eight subjects, six seeds, or both.
  4. [§5.3, Abstract and Conclusions] All decoding results are top-5 accuracies, but the abstract and conclusions state 'decoding accuracies significantly exceeding chance' without this qualification. Top-5 retrieval is a substantially weaker claim than exact word identification; please report top-1 accuracy as well, or explicitly qualify every statement about decoding accuracy with the top-5 metric.
minor comments (3)
  1. [Algorithm 1] The word 'back-propogation' is misspelled; it should be 'back-propagation'.
  2. [Figure 6] The y-axis label 'Mean Squa ed Amplitude' contains a typo; it should be 'Mean Squared Amplitude'.
  3. [Figure 7] The y-axis label 'SEEG P ower' contains a typo; it should be 'SEEG Power'.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: SACM is trained on held-out blocks against fixed external audio features, and the only unverified step (acoustic contamination check) is an empirical validity risk, not a circularity.

full rationale

The paper's central claims — speech detection and word/initial decoding — are evaluated on data partitions that are not used to fit the reported parameters. Speech detection uses 5-fold cross-validation on balanced speech/non-speech segments, with a shuffled-label random baseline. Speech decoding uses a session-wise 8:1:1 split (first eight blocks for training, remaining two for validation/test) and compares against a shuffled-feature random reference. The contrastive objective (Eq. 1) aligns SEEG representations to representations from a fixed, pre-trained HuBERT encoder; the audio encoder is not trained on the paper's data, and test retrieval (Eq. 4) selects among held-out audio segments. This is a standard supervised alignment setup, not a reduction of the prediction to a fitted parameter or to the definition of the input. The self-citations in the manuscript (Refs. 30 and 44) appear in related-work and future-work contexts and do not carry the derivation of any reported result; there is no imported 'uniqueness theorem' and no ansatz that is justified only by the authors' prior work. The one notable gap is the acoustic contamination check (Section 3.2.1 and 3.2.2): the paper states the check was performed using the MATLAB Contamination Analysis Package but reports no per-subject or per-channel outcome. If audio leaked into the SEEG channels, the decoding and single-electrode results would be artifactual. However, this is an unverified empirical assumption about data quality, not a circular argument: the paper's equations and evaluation protocol do not define the result in terms of the claim, and the missing check does not make the prediction equivalent to an input. The derivation chain is therefore self-contained with respect to circularity, and the score remains near the bottom of the 0-2 no-significant-circularity range.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The model introduces no new physical entities. The main assumptions are that the high-gamma SEEG signal is speech-related and not contaminated by audio, that HuBERT features are suitable labels for Mandarin words, and that trial alignment is accurate. Hyperparameters are set by hand without systematic tuning.

free parameters (3)
  • Temperature tau = 0.05
    Controls sharpness of the InfoNCE softmax; set by hand in Section 5.1, no sensitivity analysis.
  • SACM learning rate = 0.0003
    Set by hand in Section 5.1; affects training but not fitted to test.
  • EEGNet hyperparameters (F1, D, F2) = 16, 4, 64
    Standard settings from EEGNet paper; used as given, not tuned on this dataset.
assumptions (5)
  • domain assumption The 70-170 Hz high-gamma band contains the neural signal relevant to speech production.
    Invoked in Section 3.2.1 when band-pass filtering SEEG; not independently verified here.
  • domain assumption The acoustic contamination check removes all audio-related artifacts from the SEEG, so remaining signal is neural.
    Invoked in Section 3.2.1; outcome of the check is not reported.
  • domain assumption HuBERT speech features preserve word identity for Mandarin monosyllables with tones.
    Used in Section 4.2 as fixed audio encoder; HuBERT is pre-trained on speech, but its Mandarin tone handling is not explicitly validated.
  • standard math InfoNCE contrastive learning with cosine similarity yields a generalizable mapping from SEEG to audio.
    Core training objective in Equations 1-3; standard in deep learning.
  • domain assumption Each trial's audio segment is the true utterance of the displayed word and is time-aligned with the SEEG segment.
    Assumed in Section 3.2 and used for positive pair construction.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding." pith.science (2026). https://pith.science/paper/Z3IEN6NX

@misc{pith2026250519652,
  author       = {Pith},
  title        = {Pith review of: SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z3IEN6NX}},
  note         = {Machine review of arXiv:2505.19652}
}
read the original abstract

Speech disorders such as dysarthria and anarthria can severely impair the patient's ability to communicate verbally. Speech decoding brain-computer interfaces (BCIs) offer a potential alternative by directly translating speech intentions into spoken words, serving as speech neuroprostheses. This paper reports an experimental protocol for Mandarin Chinese speech decoding BCIs, along with the corresponding decoding algorithms. Stereo-electroencephalography (SEEG) and synchronized audio data were collected from eight drug-resistant epilepsy patients as they conducted a word-level reading task. The proposed SEEG and Audio Contrastive Matching (SACM), a contrastive learning-based framework, achieved decoding accuracies significantly exceeding chance levels in both speech detection and speech decoding tasks. Electrode-wise analysis revealed that a single sensorimotor cortex electrode achieved performance comparable to that of the full electrode array. These findings provide valuable insights for developing more accurate online speech decoding BCIs.

Figures

Figures reproduced from arXiv: 2505.19652 by the authors.

Figure 1
Figure 1. A speech decoding BCI. For patients with neurological conditions [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Electrode implantation locations for the eight subjects. Each red dot represents a recording contact, and lighter-colored dots indicate contacts located on [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Data collection setup, where SEEG and audio signals were recorded [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The SACM framework. During the training stage, SEEG and audio representations are extracted using separate neural networks. The features are [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Speech and non-speech segments in audio and SEEG data in the [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Mean squared amplitude of speech and non-speech audio segments. [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Average power of each electrode between 70-170 Hz for Subject 1. [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Confusion matrices of SEEG word decoding and speech feature clas [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 43 canonical work pages

  1. [1]

    A. B. Silva, K. T. Littlejohn, J. R. Liu, D. A. Moses, E. F. Chang, The speech neuroprosthesis, Nature Reviews Neuroscience 25 (7) (2024) 473– 492

  2. [2]

    Fedorenko, A

    E. Fedorenko, A. A. Ivanova, T. I. Regev, The language network as a natu- ral kind within the broader landscape of the human brain, Nature Reviews Neuroscience 25 (5) (2024) 289–312

  3. [3]

    Tomik, R

    B. Tomik, R. J. Guiloff, Dysarthria in amyotrophic lateral sclerosis: A review, Amyotrophic Lateral Sclerosis 11 (1-2) (2010) 4–15

  4. [4]

    Smith, M

    E. Smith, M. Delargy, Locked-in syndrome, British Medical Journal 330 (7488) (2005) 406–409. 9

  5. [5]

    S. H. Felgoise, V . Zaccheo, J. Duff, Z. Simmons, Verbal communica- tion impacts quality of life in patients with amyotrophic lateral sclerosis, Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration 17 (3-4) (2016) 179–183

  6. [6]

    J. R. Wolpaw, N. Birbaumer, D. J. McFarland, G. Pfurtscheller, T. M. Vaughan, Brain-computer interfaces for communication and control, Clinical Neurophysiology 113 (6) (2002) 767–791

  7. [7]

    E. C. Leuthardt, G. Schalk, J. R. Wolpaw, J. G. Ojemann, D. W. Moran, A brain-computer interface using electrocorticographic signals in humans, Journal of Neural Engineering 1 (2) (2004) 63–71

  8. [8]

    G. K. Anumanchipalli, J. Chartier, E. F. Chang, Speech synthesis from neural decoding of spoken sentences, Nature 568 (7753) (2019) 493–498

Show all 44 references
  1. [9]

    D. A. Moses, S. L. Metzger, J. R. Liu, G. K. Anumanchipalli, J. G. Makin, P. F. Sun, J. Chartier, M. E. Dougherty, P. M. Liu, G. M. Abrams, A. Tu- Chan, K. Ganguly, E. F. Chang, Neuroprosthesis for decoding speech in a paralyzed person with anarthria, New England Journal of Me...

  2. [10]

    Y . Liu, Z. Zhao, M. Xu, H. Yu, Y . Zhu, J. Zhang, L. Bu, X. Zhang, J. Lu, Y . Li, D. Ming, J. Wu, Decoding and synthesizing tonal language speech from brain activity, Science Advances 9 (23) (2023) eadh0478

  3. [11]

    S. L. Metzger, K. T. Littlejohn, A. B. Silva, D. A. Moses, M. P. Seaton, R. Wang, M. E. Dougherty, J. R. Liu, P. Wu, M. A. Berger, I. Zhuravleva, A. Tu-Chan, K. Ganguly, G. K. Anumanchipalli, E. F. Chang, A high- performance neuroprosthesis for speech decoding and avatar contr...

  4. [12]

    F. R. Willett, E. M. Kunz, C. Fan, D. T. Avansino, G. H. Wilson, E. Y . Choi, F. Kamdar, M. F. Glasser, L. R. Hochberg, S. Druckmann, K. V . Shenoy, J. M. Henderson, A high-performance speech neuroprosthesis, Nature 620 (7976) (2023) 1031–1036

  5. [13]

    N. S. Card, M. Wairagkar, C. Iacobacci, X. Hou, T. Singer-Clark, F. R. Willett, E. M. Kunz, C. Fan, M. V . Nia, D. R. Deo, A. Srinivasan, E. Y . Choi, M. F. Glasser, L. R. Hochberg, J. M. Henderson, K. Shahlaie, S. D. Stavisky, D. M. Brandman, An accurate and rapidly calibrati...

  6. [14]

    Zhang, Z

    D. Zhang, Z. Wang, Y . Qian, Z. Zhao, Y . Liu, X. Hao, W. Li, S. Lu, H. Zhu, L. Chen, K. Xu, Y . Li, J. Lu, A brain-to-text framework for de- coding natural tonal sentences, Cell Reports 43 (11) (2024) 114924

  7. [15]

    irritative

    J. Talairach, J. Bancaud, Lesion, “irritative” zone and epileptogenic focus, Confinia Neurologica 27 (1) (1966) 91–94

  8. [16]

    Ryvlin, J

    P. Ryvlin, J. H. Cross, S. Rheims, Epilepsy surgery in children and adults, The Lancet Neurology 13 (11) (2014) 1114–1126

  9. [17]

    Bourdillon, J

    P. Bourdillon, J. Isnard, H. Catenoix, A. Montavont, S. Rheims, P. Ryvlin, K. Ostrowsky-Coste, F. Mauguiere, M. Guénot, Stereo electroencephalography-guided radiofrequency thermocoagulation (SEEG-guided RF-TC) in drug-resistant focal epilepsy: Results from a 10-year experience...

  10. [18]

    Cossu, F

    M. Cossu, F. Cardinale, G. Casaceli, L. Castana, A. Consales, P. D’Orio, G. L. Russo, Stereo-EEG-guided radiofrequency thermocoagulations, Epilepsia 58 (S1) (2017) 66–72

  11. [19]

    Angrick, M

    M. Angrick, M. C. Ottenhoff, L. Diener, D. Ivucic, G. Ivucic, S. Goulis, J. Saal, A. J. Colon, L. Wagner, D. J. Krusienski, P. L. Kubben, T. Schultz, C. Herff, Real-time synthesis of imagined speech processes from min- imally invasive recordings of neural activity, Communicati...

  12. [20]

    P. Z. Soroush, M. Angrick, J. J. Shih, T. Schultz, D. J. Krusienski, Speech activity detection from stereotactic EEG, in: 2021 IEEE Int’l Conf. on Systems, Man, and Cybernetics, Melbourne, Australia, 2021, pp. 3402– 3407

  13. [21]

    Angrick, M

    M. Angrick, M. Ottenhoff, S. Goulis, A. J. Colon, L. Wagner, D. J. Krusienski, P. L. Kubben, T. Schultz, C. Herff, Speech synthesis from stereotactic EEG using an electrode shaft dependent multi-input convolu- tional neural network approach, in: Annual Int’l Conf. IEEE Enginee...

  14. [22]

    Verwoert, M

    M. Verwoert, M. C. Ottenhoff, S. Goulis, A. J. Colon, L. Wagner, S. Tou- sseyn, J. P. van Dijk, P. L. Kubben, C. Herff, Dataset of speech production in intracranial electroencephalography, Scientific Data 9 (1) (2022) 434

  15. [23]

    T. M. Thomas, A. Singh, L. Bullock, D. Liang, C. W. Morse, X. Sch- erschligt, J. P. Seymour, N. Tandon, Decoding articulatory and phonetic components of naturalistic continuous speech from the distributed lan- guage network, Journal of Neural Engineering 20 (4) (2023) 046030

  16. [24]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, I. Sutskever, Learning transferable visual models from natural language supervision, in: Proc. of the 38th Int’l Conf. on Machine Learning, Virtual, 2021, ...

  17. [25]

    N. E. Crone, L. Hao, J. J. Hart, D. Boatman, R. P. Lesser, R. Irizarry, B. Gordon, Electrocorticographic gamma activity during word production in spoken and sign language, Neurology 57 (11) (2001) 2045–2053

  18. [26]

    K. E. Bouchard, N. Mesgarani, K. Johnson, E. F. Chang, Functional or- ganization of human sensorimotor cortex for speech articulation, Nature 495 (7441) (2013) 327–332

  19. [27]

    Chartier, G

    J. Chartier, G. K. Anumanchipalli, K. Johnson, E. F. Chang, Encoding of articulatory kinematic trajectories in human speech sensorimotor cortex, Neuron 98 (5) (2018) 1042–1054

  20. [28]

    F. R. Willett, D. T. Avansino, L. R. Hochberg, J. M. Henderson, K. V . Shenoy, High-performance brain-to-text communication via handwriting, Nature 593 (7858) (2021) 249–254

  21. [29]

    Lee, D.-H

    K.-W. Lee, D.-H. Lee, S.-J. Kim, S.-W. Lee, Decoding neural correla- tion of language-specific imagined speech using EEG signals, in: Annual Int’l Conf. IEEE Engineering in Medicine and Biology Society, Glasgow, United Kingdom, 2022, pp. 1977–1980

  22. [30]

    S. Li, H. Wang, X. Chen, D. Wu, Multimodal brain-computer interfaces: AI-powered decoding methodologies, arXiv preprint arXiv:2502.02830 (2025)

  23. [31]

    Endong, R

    X. Endong, R. Gaoqi, X. Xiaoyue, Z. Jiaojiao, The construction of the BCC corpus in the age of big data, Corpus Linguistics 3 (1) (2016) 93– 109

  24. [32]

    Roussel, G

    P. Roussel, G. L. Godais, F. Bocquelet, M. Palma, H. Jiang, S. Zhang, A.-L. Giraud, P. Mégevand, K. Miller, J. Gehrig, C. Kell, P. Kahane, S. Chabardés, B. Yvert, Observation and assessment of acoustic contami- nation of electrophysiological brain signals during speech product...

  25. [33]

    G. Li, S. Jiang, S. E. Paraskevopoulou, M. Wang, Y . Xu, Z. Wu, L. Chen, D. Zhang, G. Schalk, Optimal referencing for stereo- electroencephalographic (SEEG) recordings, NeuroImage 183 (2018) 327–335

  26. [34]

    Giannakopoulos, pyAudioAnalysis: An open-source python library for audio signal analysis, PloS one 10 (12) (2015) e0144610

    T. Giannakopoulos, pyAudioAnalysis: An open-source python library for audio signal analysis, PloS one 10 (12) (2015) e0144610

  27. [35]

    V . J. Lawhern, A. J. Solon, N. R. Waytowich, S. M. Gordon, C. P. Hung, B. J. Lance, EEGNet: a compact convolutional neural network for EEG- based brain-computer interfaces, Journal of Neural Engineering 15 (5) (2018) 056013

  28. [36]

    Défossez, C

    A. Défossez, C. Caucheteux, J. Rapin, O. Kabeli, J.-R. King, Decoding speech perception from non-invasive brain recordings, Nature Machine Intelligence 5 (10) (2023) 1097–1107

  29. [37]

    Benchetrit, H

    Y . Benchetrit, H. J. Banville, J.-R. King, Brain decoding: toward real-time reconstruction of visual perception, in: The 12th Int’l Conf. on Learning Representations, Vienna, Austria, 2024

  30. [38]

    Q. Zhou, C. Du, S. Wang, H. He, CLIP-MUSED: CLIP-guided multi- subject visual neural information semantic decoding, in: The 12th Int’l Conf. on Learning Representations, Vienna, Austria, 2024

  31. [39]

    Y . Song, B. Liu, X. Li, N. Shi, Y . Wang, X. Gao, Decoding natural images from EEG for object recognition, in: The 12th Int’l Conf. on Learning Representations, Vienna, Austria, 2024

  32. [40]

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, A. Mohamed, HuBERT: self-supervised speech representation learning by masked prediction of hidden units, IEEE/ACM Trans. on Audio, Speech, and Language Processing 29 (2021) 3451–3460

  33. [41]

    F. Wang, H. Liu, Understanding the behaviour of contrastive loss, in: Proc. IEEE/CVF Conf. on Computer Vision and Pattern Recognition, Nashville, TN, 2021, pp. 2495–2504

  34. [42]

    A. G. Huth, W. A. D. Heer, T. L. Griffiths, F. E. Theunissen, J. L. Gallant, Natural speech reveals the semantic maps that tile human cerebral cortex, Nature 532 (7600) (2016) 453–458

  35. [43]

    X. Chen, R. Wang, A. Khalilian-Gourtani, L. Yu, P. Dugan, D. Friedman, W. Doyle, O. Devinsky, Y . Wang, A. Flinker, A neural speech decoding framework leveraging deep learning and speech synthesis, Nature Ma- chine Intelligence 6 (4) (2024) 467–480

  36. [44]

    Z. Jia, H. Wang, Y . Shen, F. Hu, J. An, K. Shu, D. Wu, Magne- toencephalography (MEG) based non-invasive Chinese speech decoding, IEEE Trans. on Cognitive and Developmental Systems, submitted (2025). 10

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.