Pith. sign in

Paper Citation Record · LEDGER

The in-context inductive biases of vision-language models differ across modalities

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2502.01530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01530 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:04:35.449560Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:38:00.558288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 632f3338-6350-4f67-8c87-82e17ddce842 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

The in-context inductive biases of vision-language models differ across modalities Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.327832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.327832Z digest=sha256:3b1353cdb5b23a8e97af7572283120e536f3de3242101f115e58b147876974a9

Observation 29d0f044-b06e-44dd-8eef-d7a3706fb66f · outbound

This paper cites Colour-name versus shape-name learning in young children.

The in-context inductive biases of vision-language models differ across modalities Colour-name versus shape-name learning in young children

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.759882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.332776Z digest=sha256:14d4a482c902e12c5359941767e7fc7061fa805073915b20404750526ba5a1c0

Observation 4ea83377-4818-4216-a14f-f0b6d31930d4 · outbound

This paper cites Language Models are Few-Shot Learners.

The in-context inductive biases of vision-language models differ across modalities Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.336573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.336573Z digest=sha256:a59ed5d6e2a671e212d2499b06066e2189a7c728d27099d8e79092f18ceacdbb

Observation 1b099067-f2a2-4a7d-be50-53140eb69d77 · outbound

This paper cites Transformers generalize differently from information stored in context vs in weights.

The in-context inductive biases of vision-language models differ across modalities Transformers generalize differently from information stored in context vs in weights

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.745176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.340817Z digest=sha256:08634a1e47c574a99899593a1409351cf118ffae9ac06c2554548df4e7d73af7

Observation 4e96c286-cfdd-4e2f-8f61-85c6f9002d54 · outbound

This paper cites The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements.

The in-context inductive biases of vision-language models differ across modalities The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.732719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.349461Z digest=sha256:21c0d595f821d548d50c2be71bb514913c83260c85768ef5865e388d17c5bfb1

Observation d5eadd47-6556-4f83-a50a-1ec1702b9579 · outbound

This paper cites Dreamsim: Learning new dimensions of human visual similarity using synthetic data.

The in-context inductive biases of vision-language models differ across modalities Dreamsim: Learning new dimensions of human visual similarity using synthetic data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.719158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.355141Z digest=sha256:6d40c524ce18c0c248773e91ba36900cc4c03ae332075d4b54fe511db535fcf3

Observation aa95eb3f-1628-4291-87da-c49cba0aa8d3 · outbound

This paper cites Ordering adjectives in referential communication.

The in-context inductive biases of vision-language models differ across modalities Ordering adjectives in referential communication

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.706032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.359616Z digest=sha256:7167b17a3351a0a20f234bd89d5002758ba71397001ddb5909ea532648941792

Observation 89178649-e182-40a8-8d17-59889033142a · outbound

This paper cites Can We Talk Models Into Seeing the World Differently?.

The in-context inductive biases of vision-language models differ across modalities Can We Talk Models Into Seeing the World Differently?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.363001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.363001Z digest=sha256:71c27fab02e072ab337b22f003fc455757f516b78a56fd57d787be0d9aa30076

Observation 919f6bc9-a5dd-4998-ab2a-aa2ffe5cfe16 · outbound

This paper cites Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness.

The in-context inductive biases of vision-language models differ across modalities Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.366822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.366822Z digest=sha256:81cf395040bbf5a9d0db630b0c8611d6c63af8f654b95dc6c8c45095a4ed67f2

Observation 9b6bdda7-6d07-4b76-a214-ac6f42ebfa02 · outbound

This paper cites Shortcut learning in deep neural networks.

The in-context inductive biases of vision-language models differ across modalities Shortcut learning in deep neural networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.370467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.370467Z digest=sha256:5b7ac184b34a632dcb0551388d071dc7fe296b6a883a3212b25c3612fdc6e168

Observation ed4d4125-e2b0-4a67-80f9-2b11c493c7a9 · outbound

This paper cites Partial success in closing the gap between human and machine vision.

The in-context inductive biases of vision-language models differ across modalities Partial success in closing the gap between human and machine vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.374251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.374251Z digest=sha256:a9d46252c407db1dd2cb39b7d227be2ce11483d72e1dcb3c7cb767e34372e81a

Observation 5c12d530-980f-4d81-bdc0-1643d2ff597f · outbound

This paper cites Logic and conversation.

The in-context inductive biases of vision-language models differ across modalities Logic and conversation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.674507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.377893Z digest=sha256:3fca80f0057a775f5e90cead0fe37dbf5547b986040788ef61fe5282b7a5bbd3

Observation 3a864aec-8413-4d71-951d-e12c1d91c422 · outbound

This paper cites Revealing the multidimensional mental representations of natural objects underlying human similarity judgements.

The in-context inductive biases of vision-language models differ across modalities Revealing the multidimensional mental representations of natural objects underlying human similarity judgements

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.663386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.383206Z digest=sha256:e0d7998d10cd9facd9a8247ae7878741818ce8db8a0247906066bc2850a33ef0

Observation be61e304-e86a-44f1-b1c8-7701f53ed20f · outbound

This paper cites What shapes feature representations? exploring datasets, architectures, and training.

The in-context inductive biases of vision-language models differ across modalities What shapes feature representations? exploring datasets, architectures, and training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.650489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.389182Z digest=sha256:8460987f9ca48d0e584e6d26ea7b3c62acc98147161d82b07455708748b21543

Observation a1404e10-6e47-4563-9899-1ca548c6b98d · outbound

This paper cites The broader spectrum of in-context learning.

The in-context inductive biases of vision-language models differ across modalities The broader spectrum of in-context learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.394368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.394368Z digest=sha256:a894f2d492c53ccc4f40e06fd848896d8c96e7002f2ea8abb538d73d5a3a9fff

Observation 815f8d0e-199a-4082-a2fc-e0c1c85c004b · outbound

This paper cites The importance of shape in early lexical learning.

The in-context inductive biases of vision-language models differ across modalities The importance of shape in early lexical learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.638485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.398462Z digest=sha256:b4e57d5b921de9d2897dcb37909a6bd8d8b5b01a2d3b6b314a8f7f3f566cc095

Observation 949b7e2e-0fcb-4415-bd04-dd239d8e855f · outbound

This paper cites Aligning Machine and Human Visual Representations across Abstraction Levels.

The in-context inductive biases of vision-language models differ across modalities Aligning Machine and Human Visual Representations across Abstraction Levels

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.404895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.404895Z digest=sha256:cc7f4e873fc305fcf78ac3878b75d9243dd509523dc2b80e12b00c75b27ec8d3

Observation 56766abc-6cfe-4bbc-9c7e-adae61da79d4 · outbound

This paper cites Adversarial training for free! Advances in neural information processing systems, 32, 2019.

The in-context inductive biases of vision-language models differ across modalities Adversarial training for free! Advances in neural information processing systems, 32, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.412082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.412082Z digest=sha256:10d34380ac9f3552fcb6d5dec3e13c0b2f42615f113be20a1d10b85693d20993

Observation 02f32883-6932-4d63-916f-9d85c6e4a1b1 · outbound

This paper cites Intriguing properties of neural networks.

The in-context inductive biases of vision-language models differ across modalities Intriguing properties of neural networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.416921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.416921Z digest=sha256:7fb2f13c1521dd60c63ae318465d1a25d388131c8296ddc2e04659c9fadb3b88

Observation 3dd74afd-215a-4f9c-aea2-66a248d522d4 · outbound

This paper cites What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models.

The in-context inductive biases of vision-language models differ across modalities What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.616811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.421381Z digest=sha256:8dcdb0a30dca817ca31ffa956fd12818142906042fade6709e11d9349cd0bf52

Observation 23296d28-0426-4f7e-b581-d67b34d7db99 · outbound

This paper cites Larger language models do in-context learning differently.

The in-context inductive biases of vision-language models differ across modalities Larger language models do in-context learning differently

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.425399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.425399Z digest=sha256:975ffd500df30966a1ac6e86033b0c2cc652f75ce601658021e7830fcb1602b0

Observation 2c06b120-441d-44ef-b34e-7e0eff0858d3 · outbound

This paper cites Word learning as bayesian inference.

The in-context inductive biases of vision-language models differ across modalities Word learning as bayesian inference

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.604365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.429873Z digest=sha256:db851cdf3996d2bbd31823d6d340162bd2c6d61b890e55af036c53c55c5ba584

Observation 0fa84074-350b-4d98-b226-d267066c5e05 · outbound

This paper cites What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023.

The in-context inductive biases of vision-language models differ across modalities What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.433581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.433581Z digest=sha256:0a333dd7057e59ae4d0a0296ddfc35740e122114e576879d5095f0138b7c94d8

Observation b3d06aa3-6000-40c4-8135-710786633bab · outbound

This paper cites write newline.

The in-context inductive biases of vision-language models differ across modalities write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.437148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.437148Z digest=sha256:82fa70fbc56083268c1f1ede7c9ba05c93dddca1913a5650d72a2d486e983827

Observation 8ac483a6-3e46-4cfc-b126-80dbfd4a62d6 · outbound

This paper cites @esa (Ref.

The in-context inductive biases of vision-language models differ across modalities @esa (Ref

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.441637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.441637Z digest=sha256:902c9eed406ee59806fd8fd3b68d3d7271a66d72b29d5b6694c510ce72126e73

Observation 8b0d6d83-99cb-4a88-a8e6-dbfcbc83db96 · outbound

This paper cites an unresolved cited work.

The in-context inductive biases of vision-language models differ across modalities Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.445663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.445663Z digest=sha256:5e48f65e887fb68e8969fb6ca980704b4c83ec3838a882c849e387f088772c7d

Observation 9337145e-ba16-43af-83e0-bb3fd8fcbe85 · outbound

This paper cites an unresolved cited work.

The in-context inductive biases of vision-language models differ across modalities Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.449560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.449560Z digest=sha256:80e2b388feef214c2c6b0db4a4246337191c175ff31a9a4d317a99850ce3c1f0

Pith citing papers

Observation d06f6240-a08d-4c04-9a8b-60082a742f1f · inbound

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs cites this paper.

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs The in-context inductive biases of vision-language models differ across modalities

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-01T17:38:00.558288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:38:00.558288Z digest=sha256:c72b9b4068a4928c224ed12d40351a75b98d9862aef60ab3ab5a46072e70e3d3