Pith. sign in

REVIEW 2 cited by

Decoding individual words from non-invasive brain recordings across 723 participants

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.17829 v1 pith:7UZH3ZKV submitted 2024-12-11 eess.SP cs.LG

classification eess.SPcs.LG
keywords participantswordsdecodingnon-invasiveacrossbraindatadecode
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning has recently enabled the decoding of language from the neural activity of a few participants with electrodes implanted inside their brain. However, reliably decoding words from non-invasive recordings remains an open challenge. To tackle this issue, we introduce a novel deep learning pipeline to decode individual words from non-invasive electro- (EEG) and magneto-encephalography (MEG) signals. We train and evaluate our approach on an unprecedentedly large number of participants (723) exposed to five million words either written or spoken in English, French or Dutch. Our model outperforms existing methods consistently across participants, devices, languages, and tasks, and can decode words absent from the training set. Our analyses highlight the importance of the recording device and experimental protocol: MEG and reading are easier to decode than EEG and listening, respectively, and it is preferable to collect a large amount of data per participant than to repeat stimuli across a large number of participants. Furthermore, decoding performance consistently increases with the amount of (i) data used for training and (ii) data used for averaging during testing. Finally, single-word predictions show that our model effectively relies on word semantics but also captures syntactic and surface properties such as part-of-speech, word length and even individual letters, especially in the reading condition. Overall, our findings delineate the path and remaining challenges towards building non-invasive brain decoders for natural language.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset

    cs.LG 2025-06 conditional novelty 6.0 of 10

    The 2025 PNPL competition presents over 50 hours of within-subject MEG data, defines speech detection and phoneme classification benchmarks with F1-macro scoring, and reports reference baselines of 68.04% and 60.39%.

  2. Dynadiff: Single-stage Decoding of Images from Continuously Evolving fMRI

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single-stage, LoRA-finetuned diffusion model decodes seen images directly from continuous BOLD fMRI time series and beats previous pipelines on semantic metrics.

Pith tools