Pith. sign in

REVIEW 3 cited by

Instruction-tuning Aligns LLMs to the Human Brain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.00575 v2 pith:4AYETPF6 submitted 2023-12-01 cs.CL

classification cs.CL
keywords alignmentbrainhumaninstruction-tuningllmsknowledgelanguageworld
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction-tuning is a widely adopted finetuning method that enables large language models (LLMs) to generate output that more closely resembles human responses. However, no studies have shown that instruction-tuning actually teaches LLMs to process language in a similar manner as humans. We investigate the effect of instruction-tuning on aligning LLM and human language processing mechanisms in two ways: (1) brain alignment, the similarity of LLM internal representations to neural activity in the human language system, and (2) behavioral alignment, the similarity of LLM and human behavior on a reading task. We assess 25 vanilla and instruction-tuned LLMs on three datasets involving humans reading naturalistic stories and sentences, and find that instruction-tuning generally enhances brain alignment (~6%), but has no similar effect on behavioral alignment. To identify factors underlying this improvement in brain alignment, we compute correlations between brain alignment and various LLM properties, such as model size, problem-solving, and world knowledge understanding. Notably, we find a strong positive correlation between brain alignment and model size (r = 0.95), as well as performance on tasks requiring world knowledge (r = 0.81). Our results demonstrate that instruction-tuning LLMs improves both world knowledge representations and brain alignment, suggesting that the mechanisms that encode world knowledge in LLMs also improve representational alignment to the human brain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5 citations worldwide. Full citation record

  1. Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Targeted perturbation of reward-anticipatory units in VLMs induces anhedonia-like effort avoidance and clinical-scale score drops without impairing baseline task competence.

  2. Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation

    cs.AI 2025-06 reject novelty 6.0 of 10

    LLM-based long-horizon event simulation, used as a reward signal, is claimed to improve safety alignment and indirect-harm detection, but evaluation confounds simulation with the capability of the external projector model.

  3. Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

    q-bio.NC 2025-05 conditional novelty 6.0 of 10

    Instruction-tuned multimodal LLMs predict fMRI responses to natural images better than vision-only models and on par with CLIP, though most explained variance is shared across instructions.

Pith tools