Pith. sign in

Paper Citation Record · LEDGER

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

As of 8 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.20029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20029 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:53.714175Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6008b149-f63a-466c-a1c2-7c0720922511 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.463444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.463444Z digest=sha256:524fec13846699dd2f256891d3f21eef1b374d8574fd405a8a4d7ff6715fb5bd

Observation 5a3b562f-ae32-4eff-851d-50c54f913806 · outbound

This paper cites InstructBLIP (Dai et al.,.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) InstructBLIP (Dai et al.,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.237159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.323012Z digest=sha256:9aa09a450a9a3246d53888e2c001a07f8775658d527a19a3ee959b9381770987

Observation 4272d690-7d15-4aae-97b4-ef46650da518 · outbound

This paper cites Yes,” “No,.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Yes,” “No,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:53.950680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.714175Z digest=sha256:233659d807accb8fdfd19397c5f8748b3b6354b499cdc49dd3c406647e6cd0db

Observation 1a9ad026-2a66-4868-8c9b-e735d6bf5ea1 · outbound

This paper cites The Llama 3 Herd of Models.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.923241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.923241Z digest=sha256:a5e871e852f9c2a1f23d92767427bcd9131e261876bc94c46200e0f9f1e78797

Observation bd294af3-9fcd-4021-b3eb-bc9e08df4189 · outbound

This paper cites Microsoft coco: Common objects in context.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Microsoft coco: Common objects in context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.296677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.296677Z digest=sha256:b04d418892d3e1f35fb9ec8c14c8f5696e8e70b7e234b5ded7aa190137efe8dd

Observation ef5420d5-bf60-4a57-aec7-a18bbfa206fc · outbound

This paper cites an unresolved cited work.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:56.492727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:52.486214Z digest=sha256:9c03f36ef3edf395b121a5cd734f5a1ff56263e6ab6155ae01165065809ae45f

Observation d1bd5683-8ad1-4bfa-a72d-2b274f026b7b · outbound

This paper cites Speech language models lack important brain-relevant semantics.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Speech language models lack important brain-relevant semantics

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:56.210466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:52.556342Z digest=sha256:dc6665a16453b59c579cb5f5d381f701a7e36fd2c9672fdb0fcfde638271c2ba

Observation 32fc21fb-8b78-4d9e-9d17-1bafb46267b0 · outbound

This paper cites Tuning in to neural encoding: Linking human brain and artificial supervised representations of language.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Tuning in to neural encoding: Linking human brain and artificial supervised representations of language

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.703135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.703135Z digest=sha256:07257fd4b225fcacf4f77f4448df448c45ee0ab1406a4c3887a95bd5ee404c7b

Observation 30d43f35-4313-4945-b3d5-643980756172 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) LLaMA: Open and Efficient Foundation Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.788312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.788312Z digest=sha256:f6db2f8f2ced910f258bfa35573b20af879b0ad5134b6aacccb1357a445bc86c

Observation 40ee46a3-fcc9-4991-9ee2-c181b98ac617 · outbound

This paper cites Natural language supervision with a large and diverse dataset builds better models of human high-level visual cor- tex.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Natural language supervision with a large and diverse dataset builds better models of human high-level visual cor- tex

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.788664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:52.944682Z digest=sha256:4ee834af14f069d67dea06bd884d7f355d65d38e114fab4624bb2275fea73d40

Observation 0154a279-ab1c-4c37-a938-4adaf1defe66 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Transformers: State-of-the-art natural language processing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.603814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.035379Z digest=sha256:b276c6ba01d93ec011717e448a2a94ac5965cdec5e794994959979630687b707

Observation e5e9885c-13ff-4e8d-b582-05c08cc75bab · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:53.137091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:53.137091Z digest=sha256:24399ccc792d60034be67d5e1ddc161b91ed429a64b26eb4c92e73560ba353ef

Observation b27bc93b-42a4-40a4-a7d5-a0b7defa4b0a · outbound

This paper cites mid temporal lobe bodies.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) mid temporal lobe bodies

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.391941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.218111Z digest=sha256:0cc17410634fa412e1e9df6a353d108a2b3c9b60ae59eb0f094c279daa361f8c

Observation 01491c0b-48a0-4842-ade4-c7e0d4003240 · outbound

This paper cites an unresolved cited work.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:55.116310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.385595Z digest=sha256:d3abfeb3da576857aa457470d663a665d762336557a6258ea9fce1fa477453bb

Observation 2d486ccb-729d-4194-bb47-984abd3823a9 · outbound

This paper cites an unresolved cited work.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:54.987578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.471257Z digest=sha256:11e5761155ca6fe3fd739b4e1a321bfc63c7ffc02d356b11ae82964aef381b5b

Observation fcfa8868-7021-4d47-8262-27de04c40461 · outbound

This paper cites The color bar highlights color codes for each instruction.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) The color bar highlights color codes for each instruction

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:54.702164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.511816Z digest=sha256:dac7c765b060bb4b8339cb143088ee5a1abd26b734c81bf8f1026f8a38304e80

Observation a993ba63-c9ec-4e8e-8890-80fabe7ef333 · outbound

This paper cites an unresolved cited work.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Unresolved cited work

Reference 25

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:04:54.163990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.630159Z digest=sha256:f119cc97aed2c7e7dfcb0bb14efc73c84d9d3fb47ca5bce3b3a3e2d7657d72c4

Observation 4a698fe1-db21-4942-8ed8-3644532d760d · outbound

This paper cites The color bar highlights color codes for each layer.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) The color bar highlights color codes for each layer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:54.428394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:53.570042Z digest=sha256:e76230d9544190901667452528250a00badbe63b7826294bf96492b9afb72fa7

Observation 49f1b797-11f5-486f-bd89-5ac289dd2bd0 · outbound

This paper cites Shared computational principles for language processing in humans and deep language models.Nature Neuroscience, 25(3):369–380,.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Shared computational principles for language processing in humans and deep language models.Nature Neuroscience, 25(3):369–380,

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:56.870712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:52.053045Z digest=sha256:a745f1ba85986e3d3a47a7980c4b5b2f56d790e1f103a2d95a4290195b518376

Observation 077b16f1-8bfc-4cd6-850a-d2ce1188c6a1 · outbound

This paper cites Mistral 7B.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Mistral 7B

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.192451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.192451Z digest=sha256:e827fff628c9324abb44500d8ca76e277bb24b2416c3cc3c24bb8f41706c8004

Observation ded41309-6046-4826-9a9a-4b67125ee344 · outbound

This paper cites Visual representations in the human brain are aligned with large language models.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Visual representations in the human brain are aligned with large language models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.812008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.812008Z digest=sha256:548c5a626c19aa7e96ab5057de54a9fe03f20a942f38eb997a99617ee65fa7e9

Observation 4031bc71-7040-47ef-b0ef-5a4a88e636d4 · outbound

This paper cites What can 1.8 billion regressions tell us about the pressures shaping high-level visual representation in brains and machines? bioRxiv, pp.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) What can 1.8 billion regressions tell us about the pressures shaping high-level visual representation in brains and machines? bioRxiv, pp

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.690010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.690010Z digest=sha256:00c48ecd4dad1f1c2a55b446696ce7b1196ec34e8837fe7547c981d73a3ded55

Observation b3c81966-98d0-43c4-ab4e-1a65c5eec772 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Learning Transferable Visual Models From Natural Language Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.633416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.633416Z digest=sha256:caad282977228c300c6e7f0b42ff9b1da4370c6068ea677002a1810de8f49a16

Observation 2da6bc0d-5468-4a17-8ea0-d5438a48a744 · outbound

This paper cites Neural taskonomy: Inferring the similarity of task- derived representations from brain activity.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Neural taskonomy: Inferring the similarity of task- derived representations from brain activity

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.980113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:52.844290Z digest=sha256:dc90ee4ee1b3b8f7b5dd3ec9019a39aa335cd10003e90ac4380e54b1e2591bdb

Observation 2b172c6a-0b48-41b3-9ff3-20874950eab3 · outbound

This paper cites Instruction-tuning Aligns LLMs to the Human Brain.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Instruction-tuning Aligns LLMs to the Human Brain

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.556068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.556068Z digest=sha256:1f7417f35a9cb1c0213c3405014ea5e0e1e50658583cd6565829c1e96d7a4aaa

Observation 0960fd90-1235-4a58-b508-589fceb87291 · outbound

This paper cites The brain tells a story: Unveiling distinct representations of semantic content in speech, objects, and stories in the human brain with large language models.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) The brain tells a story: Unveiling distinct representations of semantic content in speech, objects, and stories in the human brain with large language models

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:56.721573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:04:52.401068Z digest=sha256:21ab44e90bc9dcd211a9ad6ab99a8bbaf0ee98927d482b22f77251e0e2a0d5cc

Pith citing papers

No inbound Pith citation observations are available.