Pith. sign in

Paper Citation Record · LEDGER

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

As of 14 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2607.26326.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26326 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:13:15.367545Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved70
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ccf8c7c-3489-4d68-adc6-90d2b643b223 · outbound

This paper cites an unresolved cited work.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:09.842361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:09.842361Z digest=sha256:e219f1a4046e5185e3bc6b462f97c9b2200a3e4e177515f2e26f75f99738c4d2

Observation b7b75f1f-a283-4aa7-821a-33b7136ef599 · outbound

This paper cites Gemma 4 Technical Report.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Gemma 4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:09.923293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:09.923293Z digest=sha256:3eeeae5e5e91144fb93ba9d0ea916047aa32f8b4720da4ce7370e31a69639161

Observation b487c4ec-0c62-4f9b-aa35-1d2d52ed0182 · outbound

This paper cites 2025 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2025 , url=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.012366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.012366Z digest=sha256:493fb0efdf06d37663821ff1316658d5fa6de17e5c024c8a48dfa344cada54a4

Observation 910369bf-2559-42a9-aa8a-bb44823d356f · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning , pages =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Proceedings of the 38th International Conference on Machine Learning , pages =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.095484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.095484Z digest=sha256:6e36d14d17e3335a35f08f6eff9f0d555a64f94f0169827e6309e0bcdc91f40e

Observation b423aee0-981d-42a7-9cd5-a6d51cd9b0c2 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.155985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.155985Z digest=sha256:f8b77ff2f5611118ee68d7de414f757de388fbfab0e49e347bedc3fa5fed265e

Observation d3628eac-303f-4e75-895a-f180bf1761a1 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.225148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.225148Z digest=sha256:d91d8ad5020db4c9e72ef17caa73618c35206204c29b79a7d139b418d9af1077

Observation 84005604-d1cd-4dd6-89a8-894510332df7 · outbound

This paper cites (No Title) , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models (No Title) , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.298000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.298000Z digest=sha256:0bc5857e85d4334fc9ce63c1e09547814c1e8f2f21cfcc094a5211ad7636c933

Observation 95976a2e-11ce-4173-b4d4-91158602ff6e · outbound

This paper cites 2022 , editor =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2022 , editor =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.384660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.384660Z digest=sha256:74a2d1d1b8356797e06f8590e8037a127be13ea8bf4347630aa3645faa033d12

Observation b6fff76e-b15b-4274-ba0d-cc8185c4b826 · outbound

This paper cites Flamingo:.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Flamingo:

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.469896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.469896Z digest=sha256:19aed609f0e0b653fc5d6f553a1e54dfe2baa10dc2dd370d32027c512e4d53c8

Observation aa11c4b2-4e56-41a9-99f5-aff3889b1504 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Advances in Neural Information Processing Systems , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.533427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.533427Z digest=sha256:bc37c697eb00414d9626b5b7bba57a79303acf40b094c9366480d686dbabf169

Observation f9db17bb-3a6a-4819-bc23-035718b2b3f5 · outbound

This paper cites Advances in Neural Information Processing Systems , editor =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Advances in Neural Information Processing Systems , editor =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.598017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.598017Z digest=sha256:42752f8631a56147c30077874eb680680cacaf4049576c131571934c2067b8b7

Observation 18ecdcb4-f389-4f5e-b456-267b36449c61 · outbound

This paper cites 2023 , editor =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2023 , editor =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.664105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.664105Z digest=sha256:31013a0213bbc3983fa12d6ad7a72eae13d9aa31b24143a070732fe184d03d2e

Observation 1b6d5171-d043-4a90-88c9-d8a9080012fb · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.735781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.735781Z digest=sha256:c19e0ff7ecfede5cec3b9542dcb7e81edfd5fd0d19f91c1981b5b31f5afecb99

Observation 62a46523-e4b3-470e-a7f6-09f85696c35b · outbound

This paper cites 2024 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2024 , url=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.801958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.801958Z digest=sha256:0e9ea64338848f2ced1935f6e869f4e552ff1c8369c50e1a89f06b17d4a80d01

Observation 4f59c5cc-450a-4115-a22c-c6f7ed76bde2 · outbound

This paper cites Eyes Wide Shut?.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Eyes Wide Shut?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.882439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.882439Z digest=sha256:9bec09f33d5d1ab3c7e129b5d97d91e5620f1105437b5e8a4cbc2ede123c7e20

Observation b1403138-f2d6-45c5-89e3-cd1fc1194b53 · outbound

This paper cites 2025 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2025 , url=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:10.969920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:10.969920Z digest=sha256:01037a73b299d5c116260c78399d32612240766385f1a8dbdce2c72b8015a81e

Observation 4b6a5762-daaa-424c-9f5e-81436e38fd5c · outbound

This paper cites 2026 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2026 , url=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.059086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.059086Z digest=sha256:4264a62ecd579bb59333b2a70b3a49862032bd12ed55244e071995c51f7852d1

Observation 2eb86cb6-c5a3-450c-beb2-30fc93f09761 · outbound

This paper cites arXiv preprint arXiv:2603.03276 , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models arXiv preprint arXiv:2603.03276 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.139457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.139457Z digest=sha256:7b339550f99791c6505e16cef65f9035d4556540913a5df21970141fed92fdc8

Observation 4f136fb6-08b8-419a-9a78-10f237ebd3e4 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models The Fourteenth International Conference on Learning Representations , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.221493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.221493Z digest=sha256:ae501c88ddcde5d28e2ccccf97a46b5e5d9e3e4117db29ad01e678c6327064ea

Observation 664085ba-13fd-4f02-aad8-862a92176332 · outbound

This paper cites arXiv preprint arXiv:2512.02014 , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models arXiv preprint arXiv:2512.02014 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.280707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.280707Z digest=sha256:75e7fdaaa140ac978e7701c57a76de0e953089a9ba4f45090384a9da0e63feaf

Observation 92631c34-1e16-47e3-ab05-989c76851296 · outbound

This paper cites 2024 , doi =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2024 , doi =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.338143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.338143Z digest=sha256:80c5cc2032c819529d09bc63ff676b66086566e8fa14e0ac37d79d36bc076419

Observation 9211bc7d-11af-42f5-b615-e1e5926d1f7a · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Diffusion Transformers with Representation Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.391077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.391077Z digest=sha256:d2f804bafc3c2345a112a1b0857a89ce8663176a870002120d20442bd9ab273f

Observation 1fe5fe08-ffc9-41ca-abba-0cdd746e2954 · outbound

This paper cites Hidden in plain sight:.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Hidden in plain sight:

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.469442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.469442Z digest=sha256:60911dba036e7ba404e0546cce601487467218e010b58638d3f2c7e6775e6176

Observation d47ce39c-c931-44a8-9850-0de634ea6718 · outbound

This paper cites 2026 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2026 , url=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.549910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.549910Z digest=sha256:fe81e83f3003baf7b730f4421ad2193868827f13ca6a041af2e3d98db8bc3dba

Observation 414fb9d9-d763-4a43-8672-7c7b9db70454 · outbound

This paper cites Context-faithful Prompting for Large Language Models.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Context-faithful Prompting for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.625364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.625364Z digest=sha256:a8f23c1fa4cd87e7f088d2548eed2976d9df83ae4c29cf969ec236a364df2e33

Observation 9fa0abda-4a42-4b5f-8094-728d8024a6ec · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models The Fourteenth International Conference on Learning Representations , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.703423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.703423Z digest=sha256:bd1c4fd82d8a4cdc892b0f4621775b4462fdd484a338d37fd9a6b143a09a504a

Observation 832e7f14-080d-4a76-9621-eeb22e79db27 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models The Eleventh International Conference on Learning Representations , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.782392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.782392Z digest=sha256:b3d260792ceb8a483db3cafdb2ed86f7a56191a7bae0d7d0486746dc0721e605

Observation 0a85dc38-d652-4898-8c37-eba9bf85723d · outbound

This paper cites Copyright Violations and Large Language Models.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Copyright Violations and Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.861495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.861495Z digest=sha256:e25caf3d7d8d535a9d47c0913377a49bafbb01073556a184f21ad97444bce873

Observation 1d35855a-6928-4ec3-b79d-c6913fa663bf · outbound

This paper cites Findings of the association for computational linguistics: ACL 2023 , pages=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Findings of the association for computational linguistics: ACL 2023 , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.923173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.923173Z digest=sha256:27d744a5bf9d2b4122a3510a701fb51b8ac418b2b2b46423304cfce01f0810a4

Observation e2512a93-2ef9-4149-83d9-2e80c8419fe1 · outbound

This paper cites Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:11.994621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:11.994621Z digest=sha256:012df6da7d50a99a6312d4c8f7650218aa9cdc99d7e7cd35264cb557330d66d4

Observation 0c537931-47fe-4536-8821-6ffa8a5824ea · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models The Thirteenth International Conference on Learning Representations , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.092133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.092133Z digest=sha256:fd49ecde07aeaf4838f125d7a2d34cfb739b62869ffaf4f0cad439295c740dd4

Observation d0cdb115-e6cb-49cd-a1f7-a5adf2845606 · outbound

This paper cites Glass and Pengcheng He , booktitle=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Glass and Pengcheng He , booktitle=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.149111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.149111Z digest=sha256:6d53f78bdbad003396bdc88f3beab2aac7b5bf7a84652f1887327cb908b185a6

Observation e5297999-aa78-45fe-b4fb-d58a124fb1a6 · outbound

This paper cites Aho and Jeffrey D.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Aho and Jeffrey D

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.208109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.208109Z digest=sha256:3c4fa56df83fda83e1f74844110a8e5fd9300016691327af257d3e5650c5accd

Observation 47f809a0-8fb6-4174-8c68-a2a8de79d583 · outbound

This paper cites an unresolved cited work.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.287628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.287628Z digest=sha256:18eab66ea4254aea499cb164c70062ee25b3bb699c783440929d22a954cbf18e

Observation ec2c1c73-a7d6-4886-b6fb-425e99672c94 · outbound

This paper cites Chandra and Dexter C.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Chandra and Dexter C

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.353819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.353819Z digest=sha256:986a207c9f67fde46f098a0a2122f7db3ce5c772ee9f4a48187cae1344b75d43

Observation fb841ff3-0226-49ce-b123-0fd26c8bdfc7 · outbound

This paper cites Scalable training of.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Scalable training of

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.438844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.438844Z digest=sha256:7f5529af3381e0b5668be9564031a1fce4fa8876cd50360038cb77ed7ea69bae

Observation 67bff66b-4601-4b71-af52-162f30fd843d · outbound

This paper cites an unresolved cited work.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.545066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.545066Z digest=sha256:bc14741947c8aee23cbd624caf62f82ef80c63530708aa383a5cbfca63dc84d3

Observation 74c925e2-4f10-4bb6-8e8c-57dfa12a591b · outbound

This paper cites Tetreault , title =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Tetreault , title =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.615882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.615882Z digest=sha256:fef45c01bc53319c7e5e25c805120dec2cd19e77a7e8700bbf2c1b3d143f1287

Observation 9c988ef6-6f60-4b6b-a8c3-5597a231edcf · outbound

This paper cites A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.686570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.686570Z digest=sha256:a6c819f710f7d5e686ce254fa5a3a6988422047fa953fa15bf417c8c9f4b1bf3

Observation c9e73f2c-1660-462d-a86f-0088bc9446c6 · outbound

This paper cites 2024 , doi=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2024 , doi=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.758309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.758309Z digest=sha256:27a1c2e6217342cf44d6d29fade34e765fc953fe9e353094f51ced018946eb3b

Observation 9175f5fc-4d48-4e19-831e-acac83db6132 · outbound

This paper cites 2024 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2024 , url=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.834873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.834873Z digest=sha256:a35107d697ce26e5142c0461634d765d069dce1b795388569e20c88ddeec6abc

Observation 25c9e551-5903-4b19-8a14-0eb5d6779a4f · outbound

This paper cites Transfer between Modalities with.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Transfer between Modalities with

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.882556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.882556Z digest=sha256:be49acdbc27a54aec1b5f2ee61c5e0a94ad55ffd87c88d9795d84c4fae298f43

Observation cca88b54-17f8-4a69-82f2-cbf1173cc5d7 · outbound

This paper cites 2025 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2025 , url=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:12.939453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:12.939453Z digest=sha256:3e44b75b2e1a92c128137401572ba6e6dfc03a8463065d2617d67bc10afa7f15

Observation b2255361-d862-4368-acd4-e031019e462a · outbound

This paper cites 2026 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2026 , url=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.001910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.001910Z digest=sha256:27f84b208b7b1afcaa80e78044b88d1cfbaf757ed9809e7b42d4516cfae047f4

Observation 5e8e5c8b-c4e8-452c-a8f8-0bbdab1d3d81 · outbound

This paper cites 2025 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2025 , url=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.054833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.054833Z digest=sha256:666e890b6358a1345343f8fc83acd08fcc293d7bc903a858699cd17a7e2d5412

Observation e7faeabd-24cf-4d46-819a-c49160be9bc0 · outbound

This paper cites Auto-encoding variational.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Auto-encoding variational

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.115747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.115747Z digest=sha256:665fc0d8965e0191bb9eca3e6c05b35cc6e0db3f37d9e975a6e3069e3f4790a7

Observation f45236fe-4298-4598-be09-663110e02368 · outbound

This paper cites International Conference on Learning Representations , volume=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models International Conference on Learning Representations , volume=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.182783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.182783Z digest=sha256:7ab1396c20c0c16f29b6b5d5bfa85ff34d1bbcb695109049b132511f8dc055fe

Observation 4e5cbc8a-5af0-4400-9a32-3005e251cbad · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models The Fourteenth International Conference on Learning Representations , year=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.244115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.244115Z digest=sha256:d68efa8ffeb0d73c6a9aae17735fbc8da8dd010911338b66997f0a4ad2e7c908

Observation 0cd07c96-6e73-49c5-b34e-10d3633de3e4 · outbound

This paper cites , author=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models , author=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.310093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.310093Z digest=sha256:333a971d841dd4ba58f0439be7dfed020c74d8983639e9ce6bbcd1d8e832dc52

Observation b6083585-9c5e-4f09-add9-dcc09c9bbf75 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models The Fourteenth International Conference on Learning Representations , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.361805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.361805Z digest=sha256:a04d0060d8e4a1dd0115b982d6cc274c68910e43755d2c4e504cbcc3ae5cac3e

Observation 5d861501-2f06-4f0b-9eb4-55de9c4c8232 · outbound

This paper cites VL ind-Bench: Measuring Language Priors in Large Vision-Language Models.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models VL ind-Bench: Measuring Language Priors in Large Vision-Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.423914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.423914Z digest=sha256:06a6c03e9e3bc49f7f4bc07c0c09a43de5a04cf731556f14d16562d6f3570c76

Observation 08519e15-7219-4f0c-b7bb-d4ff1ab24777 · outbound

This paper cites 2025 , url=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2025 , url=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.505397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.505397Z digest=sha256:cd6c5068b4a7bfbd27c34f085073150a0f2aed95d08b760213f2284a5354b253

Observation f0a677dd-9dc0-4dc5-88fd-fa8d6487ee01 · outbound

This paper cites and Bar, Amir and Singh, Ritambhara and Eickhoff, Carsten.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models and Bar, Amir and Singh, Ritambhara and Eickhoff, Carsten

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.539568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.539568Z digest=sha256:8413d4132200a8944749da92e5d531eb18bb00cd4ba9d06b1fa77f931382c81c

Observation 97dcdc19-2745-4bc6-bbf9-805f71b60bad · outbound

This paper cites ROME : Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models ROME : Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense

Reference 54

Resolution
verified exact
doi, observed 2026-08-01T00:16:14.874294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-01T00:13:13.556259Z digest=sha256:4f090fd9480093c9d4ab14594d3176c46c40a9ea3fc4195493d735a8f99d51d2

Observation f1a95946-80d5-47b9-9ff4-eb9f75b297cf · outbound

This paper cites Characterizing Mechanisms for Factual Recall in Language Models.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Characterizing Mechanisms for Factual Recall in Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.657915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.657915Z digest=sha256:fa1f6dbf3065610c86d9ee7b55beeebd5322bcc10e717f0869c9038e47824ef5

Observation 2f99b3a7-a02b-4770-abb6-8f94928dc089 · outbound

This paper cites Activation Scaling for Steering and Interpreting Language Models.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Activation Scaling for Steering and Interpreting Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.758516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.758516Z digest=sha256:6c9c5f8de6cd1df4e79418ac0479807a2c8da6cfba7bd3e46df5721868991d27

Observation 333d0f76-a160-49a6-bc16-8a518c9199ac · outbound

This paper cites A Glitch in the Matrix ? Locating and Detecting Language Model Grounding with Fakepedia.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models A Glitch in the Matrix ? Locating and Detecting Language Model Grounding with Fakepedia

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.842678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.842678Z digest=sha256:0688afc13745d31c46f831af5fdae5527c80ad32060a112c2ee247bf8001a8ae

Observation 29fe374e-bfa1-49a1-99f2-37a5a6ac3f6d · outbound

This paper cites International Conference on Machine Learning (.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models International Conference on Machine Learning (

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.890107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.890107Z digest=sha256:07c0910b412c1e038ef0585c831b4e4f19d7f1fb55e5d45989d77e89f5ceeb0a

Observation 43c93d88-791a-4587-a1c2-5c90c6d9f638 · outbound

This paper cites 2009 , pages=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models 2009 , pages=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:13.928509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:13.928509Z digest=sha256:0df4a03baff90613625c634ba79212bcfb72ce5dafb05633c54ebb17a1d15a3f

Observation 44f8bd7d-e5fc-4e0d-b98a-f913eb2635a7 · outbound

This paper cites Interpretability in the Wild:.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Interpretability in the Wild:

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.009568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.009568Z digest=sha256:9867821eea3c3a32196282a06bc70a48945ed1f4dfbac02e3d29cb2f357a471b

Observation 612c4b22-e8ad-4b63-92d6-d4c4279b5772 · outbound

This paper cites Probing Visual Language Priors in.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Probing Visual Language Priors in

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.088574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.088574Z digest=sha256:64364023c3b929f358e900a0558d55d897c499e05fae91dff48f5f0878c29fe7

Observation 81a85eb0-27f6-42f1-861c-71d6473581c8 · outbound

This paper cites Steering Language Models With Activation Engineering.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Steering Language Models With Activation Engineering

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.164799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.164799Z digest=sha256:78679778f09a2d4f39e38025a141c52c0b3238bbfe68ee39eaa7f55ecb699f32

Observation 00ffaf42-7708-4387-8627-5c562e590520 · outbound

This paper cites and Wang, Zifan and Mallen, Alex and Basart, Steven and Koyejo, Sanmi and Song, Dawn and Fredrikson, Matt and Kolter, J.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models and Wang, Zifan and Mallen, Alex and Basart, Steven and Koyejo, Sanmi and Song, Dawn and Fredrikson, Matt and Kolter, J

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.273315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.273315Z digest=sha256:d3287078349ee7e026e656125ff1444bc6bf8a51ecb7ff7c825999a9e190e127

Observation a7388b18-5804-4b6d-b312-dc28ed9eba99 · outbound

This paper cites Proceedings of the Third Conference on Causal Learning and Reasoning , pages =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Proceedings of the Third Conference on Causal Learning and Reasoning , pages =

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.331436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.331436Z digest=sha256:ed4a618e39187d099c988c3471e7016c744d3d81d823cdd7390580c9f7d80920

Observation 1e338a5d-47e9-40d9-a15e-2e0344dd05cb · outbound

This paper cites Entity-Based Knowledge Conflicts in Question Answering.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Entity-Based Knowledge Conflicts in Question Answering

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.425106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.425106Z digest=sha256:bd878b955ded753560e99efad8583a53d3f5ed7fdc4a31dcaf745d4dadb8d010

Observation 6142a99f-3150-4e3e-b9b2-5480d5d10b3f · outbound

This paper cites Workshop on Responsibly Building the Next Generation of Multimodal Foundational Models , year=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Workshop on Responsibly Building the Next Generation of Multimodal Foundational Models , year=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.533536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.533536Z digest=sha256:45177039276e54ab845c2888b112814dcf827dd9b5a080e094b699ac0d2c9663

Observation 706ee78c-78ca-4427-80b4-f1dd60c88a62 · outbound

This paper cites Understanding Retrieval Robustness for Retrieval-augmented Image Captioning.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Understanding Retrieval Robustness for Retrieval-augmented Image Captioning

Reference 67

Resolution
verified exact
doi, observed 2026-08-01T00:16:14.719348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-01T00:13:14.663572Z digest=sha256:a6c693361b4d3f9dd5804b67b13f7ad029db024b1627b20dab341949feb1a867

Observation 86c0618a-5e43-42dc-8352-4c7314ec181a · outbound

This paper cites Position:.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Position:

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.769727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.769727Z digest=sha256:77d37f02a70b1c91966092826f529c56e5274eee4f7428c30b0c028c977b7a00

Observation 36f6e457-a25f-4f20-9a94-2d9240f1f508 · outbound

This paper cites Do Vision and Language Models Share Concepts?.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Do Vision and Language Models Share Concepts?

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:14.910017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:14.910017Z digest=sha256:74ba10319ae12b79be77452099eafdb6b5a7fd1df8bede096ad73800e73507b3

Observation 9f0fbfb4-ce1a-44ad-bd3c-5c35e2dde9a7 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , month =.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , month =

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:15.056648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:15.056648Z digest=sha256:b2c921f7ac6f9070b7a43e03075a38a5221748fc0fb54c811b14959ddb461320

Observation 3d0e2560-ab0a-459d-916d-e9f684a2c139 · outbound

This paper cites F oodie QA : A Multimodal Dataset for Fine-Grained Understanding of C hinese Food Culture.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models F oodie QA : A Multimodal Dataset for Fine-Grained Understanding of C hinese Food Culture

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:15.179630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:15.179630Z digest=sha256:7f38f7706c42ec72c414ba41dd4cff2c2a5ad03d0a89aec53ac59def738fc18a

Observation a08553a4-1e69-4ce1-8779-0663cbeb8c51 · outbound

This paper cites What if Othello -Playing Language Models Could See?.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models What if Othello -Playing Language Models Could See?

Reference 72

Resolution
verified exact
doi, observed 2026-08-01T00:16:14.550006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-01T00:13:15.289611Z digest=sha256:5bf323598c01a9dd00ff2e434009014b3fb4bfe2357ed6f3cef7c0f75ed6243b

Observation 27fe734f-35dc-4279-8cfc-8d09ce1fd985 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:15.367545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:13:15.367545Z digest=sha256:f8e31dcf4cc1afd68de2e76c54c1d5e5347a2d0328d526b8a64c8b0cda5906de

Pith citing papers

No inbound Pith citation observations are available.