Pith. sign in

REVIEW 1 cited by

Self-supervised models of audio effectively explain human cortical responses to speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14252 v1 pith:VFQX3AAN submitted 2022-05-27 cs.CL

classification cs.CL
keywords modelsself-supervisedauditoryhumanacousticlayersprocessingspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either hand-constructed acoustic filters or representations from supervised audio neural networks. In this work, we capitalize on the progress of self-supervised speech representation learning (SSL) to create new state-of-the-art models of the human auditory system. Compared against acoustic baselines, phonemic features, and supervised models, representations from the middle layers of self-supervised models (APC, wav2vec, wav2vec 2.0, and HuBERT) consistently yield the best prediction performance for fMRI recordings within the auditory cortex (AC). Brain areas involved in low-level auditory processing exhibit a preference for earlier SSL model layers, whereas higher-level semantic areas prefer later layers. We show that these trends are due to the models' ability to encode information at multiple linguistic levels (acoustic, phonetic, and lexical) along their representation depth. Overall, these results show that self-supervised models effectively capture the hierarchy of information relevant to different stages of speech processing in human cortex.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting Artificial Neural Network Representations to Learn Recognition Model for Music Identification from Brain Recordings

    q-bio.NC 2024-12 conditional novelty 5.0 of 10

    Training EEG encoders with an auxiliary InfoNCE loss that predicts a co-trained music encoder's representation improves 10-song EEG identification accuracy on the NMED-T dataset.

Pith tools