Pith. sign in

REVIEW 2 cited by

Toward a realistic model of speech processing in the brain with self-supervised learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.01685 v2 pith:PNUIYFVZ submitted 2022-06-03 q-bio.NC cs.AIcs.CL

classification q-bio.NCcs.AIcs.CL
keywords brainspeechprocessingself-supervisedalgorithmsfunctionalaccountacquisition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Several deep neural networks have recently been shown to generate activations similar to those of the brain in response to the same input. These algorithms, however, remain largely implausible: they require (1) extraordinarily large amounts of data, (2) unobtainable supervised labels, (3) textual rather than raw sensory input, and / or (4) implausibly large memory (e.g. thousands of contextual words). These elements highlight the need to identify algorithms that, under these limitations, would suffice to account for both behavioral and brain responses. Focusing on the issue of speech processing, we here hypothesize that self-supervised algorithms trained on the raw waveform constitute a promising candidate. Specifically, we compare a recent self-supervised architecture, Wav2Vec 2.0, to the brain activity of 412 English, French, and Mandarin individuals recorded with functional Magnetic Resonance Imaging (fMRI), while they listened to ~1h of audio books. Our results are four-fold. First, we show that this algorithm learns brain-like representations with as little as 600 hours of unlabelled speech -- a quantity comparable to what infants can be exposed to during language acquisition. Second, its functional hierarchy aligns with the cortical hierarchy of speech processing. Third, different training regimes reveal a functional specialization akin to the cortex: Wav2Vec 2.0 learns sound-generic, speech-specific and language-specific representations similar to those of the prefrontal and temporal cortices. Fourth, we confirm the similarity of this specialization with the behavior of 386 additional participants. These elements, resulting from the largest neuroimaging benchmark to date, show how self-supervised learning can account for a rich organization of speech processing in the brain, and thus delineate a path to identify the laws of language acquisition which shape the human brain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disentangling the Factors of Convergence between Brains and Computer Vision Models

    cs.AI 2025-08 unverdicted novelty 7.0 of 10

    By systematically varying model size, training amount, and image type in DINOv3 vision transformers, this paper shows that brain similarity increases with scale and human-centric data and emerges in a characteristic t...

  2. From Thought to Action: How a Hierarchy of Neural Dynamics Supports Language Production

    q-bio.NC 2025-02 conditional novelty 6.0 of 10

    During typing, the brain sequentially represents sentence context, then words, syllables, and letters, and these representations overlap in time and are carried by neural codes that change faster for lower-level features.

Pith tools