Pith. sign in

Paper Citation Record · LEDGER

The in-context inductive biases of vision-language models differ across modalities

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2502.01530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01530 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:04:35.449560Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:38:00.558288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 632f3338-6350-4f67-8c87-82e17ddce842 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

The in-context inductive biases of vision-language models differ across modalities Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.327832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.327832Z digest=sha256:877106dfa8df6390f408631e7823cf0f31c8e229053e4ce363460957e3a882ea

Observation 29d0f044-b06e-44dd-8eef-d7a3706fb66f · outbound

This paper cites Colour-name versus shape-name learning in young children.

The in-context inductive biases of vision-language models differ across modalities Colour-name versus shape-name learning in young children

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.759882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.332776Z digest=sha256:df52971320d79d9aab8ace6c717abfc7c700ad20741db406d253bc7dfc62bed4

Observation 4ea83377-4818-4216-a14f-f0b6d31930d4 · outbound

This paper cites Language Models are Few-Shot Learners.

The in-context inductive biases of vision-language models differ across modalities Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.336573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.336573Z digest=sha256:38e2543464f4f445270c2e43600dc945a8e2e58beab2fd44a2a709a700bf274b

Observation 1b099067-f2a2-4a7d-be50-53140eb69d77 · outbound

This paper cites Transformers generalize differently from information stored in context vs in weights.

The in-context inductive biases of vision-language models differ across modalities Transformers generalize differently from information stored in context vs in weights

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.745176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.340817Z digest=sha256:cca55f9644badc138b69f87d2e4786b3755d52792ee780bd7aa9ef6cea285c2d

Observation 4e96c286-cfdd-4e2f-8f61-85c6f9002d54 · outbound

This paper cites The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements.

The in-context inductive biases of vision-language models differ across modalities The different representational frameworks underpinning abstract and concrete knowledge: Evidence from odd-one-out judgements

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.732719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.349461Z digest=sha256:d566efb9e4d614758dbe63cd2f354a8c2adbdf289624c9f779bca91ae64983de

Observation d5eadd47-6556-4f83-a50a-1ec1702b9579 · outbound

This paper cites Dreamsim: Learning new dimensions of human visual similarity using synthetic data.

The in-context inductive biases of vision-language models differ across modalities Dreamsim: Learning new dimensions of human visual similarity using synthetic data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.719158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.355141Z digest=sha256:94e4992045095e7ac11d1d2fbecded38c4bcc508f51ab6194eb2391000124ec4

Observation aa95eb3f-1628-4291-87da-c49cba0aa8d3 · outbound

This paper cites Ordering adjectives in referential communication.

The in-context inductive biases of vision-language models differ across modalities Ordering adjectives in referential communication

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.706032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.359616Z digest=sha256:84840dbc7b166e6c8d5cb8b0cb74c899f63d08fdd9998741831eb9fac652151c

Observation 89178649-e182-40a8-8d17-59889033142a · outbound

This paper cites Can We Talk Models Into Seeing the World Differently?.

The in-context inductive biases of vision-language models differ across modalities Can We Talk Models Into Seeing the World Differently?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.363001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.363001Z digest=sha256:fbdf6c21bf3b3ab00bb3603e2e32d05c653bafa3989f63d8b350efc2f882c783

Observation 919f6bc9-a5dd-4998-ab2a-aa2ffe5cfe16 · outbound

This paper cites Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness.

The in-context inductive biases of vision-language models differ across modalities Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.366822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.366822Z digest=sha256:09bd99064413b4257b8898deb72acf40e3eae86d5781b3ce6bfe3b64ec21811b

Observation 9b6bdda7-6d07-4b76-a214-ac6f42ebfa02 · outbound

This paper cites Shortcut learning in deep neural networks.

The in-context inductive biases of vision-language models differ across modalities Shortcut learning in deep neural networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.370467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.370467Z digest=sha256:3038e7ce2b4379f3976ce69805c61b68861de16c3cf6a38a83aa1e3bcf8a487c

Observation ed4d4125-e2b0-4a67-80f9-2b11c493c7a9 · outbound

This paper cites Partial success in closing the gap between human and machine vision.

The in-context inductive biases of vision-language models differ across modalities Partial success in closing the gap between human and machine vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.374251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.374251Z digest=sha256:130b8afb048061cacc14804d39e1b3499373b001f77ca93c46a06e7d4f2c4084

Observation 5c12d530-980f-4d81-bdc0-1643d2ff597f · outbound

This paper cites Logic and conversation.

The in-context inductive biases of vision-language models differ across modalities Logic and conversation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.674507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.377893Z digest=sha256:532a3f2d3d89e8116f405b82f22d772525aea5d2e0b0562bd40dd30ed7d2f120

Observation 3a864aec-8413-4d71-951d-e12c1d91c422 · outbound

This paper cites Revealing the multidimensional mental representations of natural objects underlying human similarity judgements.

The in-context inductive biases of vision-language models differ across modalities Revealing the multidimensional mental representations of natural objects underlying human similarity judgements

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.663386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.383206Z digest=sha256:acdf8a408d87629dc08a8df285dfc6680a7419ca405cf6ff55cfa1f3bc1c0cd6

Observation be61e304-e86a-44f1-b1c8-7701f53ed20f · outbound

This paper cites What shapes feature representations? exploring datasets, architectures, and training.

The in-context inductive biases of vision-language models differ across modalities What shapes feature representations? exploring datasets, architectures, and training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.650489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.389182Z digest=sha256:a5a64cd277feccdc64f449dbac97567c8f7dbcfcaca1272a7cffd2ec5bd75d76

Observation a1404e10-6e47-4563-9899-1ca548c6b98d · outbound

This paper cites The broader spectrum of in-context learning.

The in-context inductive biases of vision-language models differ across modalities The broader spectrum of in-context learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.394368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.394368Z digest=sha256:a9b383450878847a588c814f2b478c978c5f36987042b8f91978ead2fba33bbc

Observation 815f8d0e-199a-4082-a2fc-e0c1c85c004b · outbound

This paper cites The importance of shape in early lexical learning.

The in-context inductive biases of vision-language models differ across modalities The importance of shape in early lexical learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.638485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.398462Z digest=sha256:d45b3c96726515809c96c2129a77829bdfd1d0d5626b1b8641c75e572c4b20f2

Observation 949b7e2e-0fcb-4415-bd04-dd239d8e855f · outbound

This paper cites Aligning Machine and Human Visual Representations across Abstraction Levels.

The in-context inductive biases of vision-language models differ across modalities Aligning Machine and Human Visual Representations across Abstraction Levels

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.404895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.404895Z digest=sha256:d9962e3b21d45b4c06a4b803e630b564fe5660c7287d5d5c6399b17fca566350

Observation 56766abc-6cfe-4bbc-9c7e-adae61da79d4 · outbound

This paper cites Adversarial training for free! Advances in neural information processing systems, 32, 2019.

The in-context inductive biases of vision-language models differ across modalities Adversarial training for free! Advances in neural information processing systems, 32, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.412082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.412082Z digest=sha256:fee6377f4ac12bd163acd35da00fdbef37986a9011d33d728d4c14e2aeec66ab

Observation 02f32883-6932-4d63-916f-9d85c6e4a1b1 · outbound

This paper cites Intriguing properties of neural networks.

The in-context inductive biases of vision-language models differ across modalities Intriguing properties of neural networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.416921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.416921Z digest=sha256:25a7e422b733b5c157c88a61c697ad4afc7ff9f6c3f43a8476a2fb873e8703cc

Observation 3dd74afd-215a-4f9c-aea2-66a248d522d4 · outbound

This paper cites What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models.

The in-context inductive biases of vision-language models differ across modalities What does kiki look like? cross-modal associations between speech sounds and visual shapes in vision-and-language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.616811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.421381Z digest=sha256:44613ff767eeedf96dda854631731f6a52aed842f89c37b68ca0845c5dfde454

Observation 23296d28-0426-4f7e-b581-d67b34d7db99 · outbound

This paper cites Larger language models do in-context learning differently.

The in-context inductive biases of vision-language models differ across modalities Larger language models do in-context learning differently

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.425399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.425399Z digest=sha256:4f30188d7c0664c4f5bff633eac6a50f900984696e86d8f85e767794192753c7

Observation 2c06b120-441d-44ef-b34e-7e0eff0858d3 · outbound

This paper cites Word learning as bayesian inference.

The in-context inductive biases of vision-language models differ across modalities Word learning as bayesian inference

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T15:04:35.604365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T15:04:35.429873Z digest=sha256:86fbb5e3c2196fbc3b45284a2ab2adcf31138f10e25964613c5846f711e77b07

Observation 0fa84074-350b-4d98-b226-d267066c5e05 · outbound

This paper cites What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023.

The in-context inductive biases of vision-language models differ across modalities What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36: 0 17773--17794, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.433581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.433581Z digest=sha256:b6a80dfdb10223d9a8a0a14f88876cf8a66f5d1f28c440bca8a277c2b8f36a24

Observation b3d06aa3-6000-40c4-8135-710786633bab · outbound

This paper cites write newline.

The in-context inductive biases of vision-language models differ across modalities write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.437148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.437148Z digest=sha256:4d3814a57cc6cb0e3461e8327993aa83d681b982e84b1826c88a3ccc46b7840e

Observation 8ac483a6-3e46-4cfc-b126-80dbfd4a62d6 · outbound

This paper cites @esa (Ref.

The in-context inductive biases of vision-language models differ across modalities @esa (Ref

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.441637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.441637Z digest=sha256:c075ca9b43935709497cbee8e82131afc104ebd37272d02c45ff6384c04dfb88

Observation 8b0d6d83-99cb-4a88-a8e6-dbfcbc83db96 · outbound

This paper cites an unresolved cited work.

The in-context inductive biases of vision-language models differ across modalities Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.445663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.445663Z digest=sha256:f3c5a35e855cf24b314fbf1b077aa8401d5ba7453513f7d3e110a1dc63d2d6c6

Observation 9337145e-ba16-43af-83e0-bb3fd8fcbe85 · outbound

This paper cites an unresolved cited work.

The in-context inductive biases of vision-language models differ across modalities Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:35.449560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:04:35.449560Z digest=sha256:e27fa1f8af08de6bff1f18bbb0ff1644e0bb3d4d13a318049c6e48d294a4b4d1

Pith citing papers

Observation d06f6240-a08d-4c04-9a8b-60082a742f1f · inbound

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs cites this paper.

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs The in-context inductive biases of vision-language models differ across modalities

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-01T17:38:00.558288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:38:00.558288Z digest=sha256:af21a51f6ae4eb610c9190862312a9d92b3244e62bfcf23da425e634f477bca6