Pith. sign in

REVIEW 1 cited by

Matching domain experts by training from scratch on domain knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09395 v2 pith:6K3RCTXP submitted 2024-05-15 q-bio.NC cs.AIcs.CL

classification q-bio.NCcs.AIcs.CL
keywords neurosciencetrainedllmsperformancesmallliteraturemodelsresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, large language models (LLMs) have outperformed human experts in predicting the results of neuroscience experiments (Luo et al., 2024). What is the basis for this performance? One possibility is that statistical patterns in that specific scientific literature, as opposed to emergent reasoning abilities arising from broader training, underlie LLMs' performance. To evaluate this possibility, we trained (next word prediction) a relatively small 124M-parameter GPT-2 model on 1.3 billion tokens of domain-specific knowledge. Despite being orders of magnitude smaller than larger LLMs trained on trillions of tokens, small models achieved expert-level performance in predicting neuroscience results. Small models trained on the neuroscience literature succeeded when they were trained from scratch using a tokenizer specifically trained on neuroscience text or when the neuroscience literature was used to finetune a pretrained GPT-2. Our results indicate that expert-level performance may be attained by even small LLMs through domain-specific, auto-regressive training approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Human-Like Processing: Large Language Models Perform Equivalently on Forward and Backward Scientific Text

    cs.CL 2024-11 conditional novelty 5.0 of 10

    GPT-2 models trained on character-reversed neuroscience text perform as well on a neuroscience abstract-selection benchmark as models trained on normal text, despite higher perplexity.

Pith tools