Pith. sign in

REVIEW 2 cited by

Scaling laws for language encoding models in fMRI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11863 v4 pith:UQSBLJK5 submitted 2023-05-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelslanguagebrainencodingscalingfmriperformancesize
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Representations from transformer-based unidirectional language models are known to be effective at predicting brain responses to natural language. However, most studies comparing language models to brains have used GPT-2 or similarly sized language models. Here we tested whether larger open-source models such as those from the OPT and LLaMA families are better at predicting brain responses recorded using fMRI. Mirroring scaling results from other contexts, we found that brain prediction performance scales logarithmically with model size from 125M to 30B parameter models, with ~15% increased encoding performance as measured by correlation with a held-out test set across 3 subjects. Similar logarithmic behavior was observed when scaling the size of the fMRI training set. We also characterized scaling for acoustic encoding models that use HuBERT, WavLM, and Whisper, and we found comparable improvements with model size. A noise ceiling analysis of these large, high-performance encoding models showed that performance is nearing the theoretical maximum for brain areas such as the precuneus and higher auditory cortex. These results suggest that increasing scale in both models and data will yield incredibly effective models of language processing in the brain, enabling better scientific understanding as well as applications such as decoding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Neural Foundation Models for Vision: Aligning EEG, MEG, and fMRI Representations for Decoding, Encoding, and Modality Conversion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A contrastive model aligns EEG, MEG, and fMRI activity to CLIP image embeddings, enabling image retrieval from brain signals, neural retrieval from images, and cross-modal neural retrieval.

  2. Bridging Auditory Perception and Language Comprehension through MEG-Driven Encoding Models

    q-bio.NC 2024-12 conditional novelty 4.0 of 10

    Text embeddings from CLIP and GPT-2 predict MEG responses to spoken stories better than audio features, with text effects strongest over frontal sensors and audio effects over lateral temporal sensors.

Pith tools