Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models Can't See the Obvious

As of 18 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2507.04741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04741 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:44:42.681512Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:59:19.379119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:06:03.287917Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved22
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd12751b-47ca-4096-a191-89bda0a7648c · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Vision-Language Models Can't See the Obvious Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:37.709893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:37.709893Z digest=sha256:84eeb27ccf6a902ced74ae385eb628be42436fdcf68c2d71111ad5930f32542f

Observation 14030453-a82d-4e78-8056-cc33a3e8e7f4 · outbound

This paper cites GPT-4 Technical Report.

Vision-Language Models Can't See the Obvious GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:37.742263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:37.742263Z digest=sha256:cb8a0f1a8715bfbb00c617379c828e2a9dc5f00817fc115b4ec8dbfd92d18b1c

Observation c84e1d8b-fd01-4970-8d70-d88fc1c91ee9 · outbound

This paper cites Nocaps: Novel object captioning at scale.

Vision-Language Models Can't See the Obvious Nocaps: Novel object captioning at scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:48.926655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:37.818524Z digest=sha256:0f9ddff6cf75a860e983d8d6af06e6cc338ae3e76a68955e9a96e585589ae2ff

Observation d9902289-7f7d-4263-81bb-a492e0057036 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Vision-Language Models Can't See the Obvious Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:48.787498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:37.873272Z digest=sha256:6e96582ab73d61df6b7588de8bc24c67b4ac04a2de026ad9ff3f88c7c49fc8d3

Observation 9245c91e-3f91-4a07-a082-51bfa2d0ad38 · outbound

This paper cites Claude 3.5 sonnet.

Vision-Language Models Can't See the Obvious Claude 3.5 sonnet

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:48.603654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:37.942975Z digest=sha256:6c7bee187d976e55cde9b8995eb43f4242021872d8c8ac60f2cfd765b845526b

Observation 01da0eb5-e139-462a-8a73-b0c20fedacac · outbound

This paper cites Turning visual search time on its head.

Vision-Language Models Can't See the Obvious Turning visual search time on its head

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:48.460746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:37.996493Z digest=sha256:5d00035789ee2d21da5edde6034e550264c67cb60b3e8e060cdd68cc99146135

Observation 6e58f10e-1d99-417a-a725-3919fb13421b · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Vision-Language Models Can't See the Obvious PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:38.051046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:38.051046Z digest=sha256:3803e47aa64b67651156f5bb2b910c000a2766767753454fa834801d1ebbf1b9

Observation 85ec1332-0188-49f3-8417-6e2c4ea76d3b · outbound

This paper cites Scene text visual question answering.

Vision-Language Models Can't See the Obvious Scene text visual question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:48.288258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:38.114146Z digest=sha256:a6eb396e35ae8c6393c9f85d73d773749d8e39f41a6505f104417a3b6540d003

Observation acaf3167-99b5-4a3c-9459-617788618ca5 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Vision-Language Models Can't See the Obvious $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:38.169220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:38.169220Z digest=sha256:1fbf5761ba69cd51aefbbce9c7ffbcf5977988fca7b46e7b3a6769ec3e7128ec

Observation 7f2421e0-e15e-4496-ba4c-52e29bad11e6 · outbound

This paper cites Omni3d: A large benchmark and model for 3d object detection in the wild.

Vision-Language Models Can't See the Obvious Omni3d: A large benchmark and model for 3d object detection in the wild

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:48.143683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:38.230708Z digest=sha256:d8cffcde0d98b244c3f82d9c190455e5ba3fb1787024a55bf6902533d782a691

Observation a3e07d3b-54ea-481e-993b-17d919c953f4 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Vision-Language Models Can't See the Obvious Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:38.292602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:38.292602Z digest=sha256:bced271ca8dd7ef68887914a53c865e980bee4e3e458d5f30ab9691c753360bd

Observation 66303033-7847-46cf-9af2-7a3889c0c8ff · outbound

This paper cites Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks.

Vision-Language Models Can't See the Obvious Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:47.983123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:38.347843Z digest=sha256:565b87a29bc2494eaab2b58b790802fb645d7aa87580aa6fc49173c5b07f5db0

Observation b53fa82b-0c77-4ecf-9072-06482e2ac240 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Vision-Language Models Can't See the Obvious NVLM: Open Frontier-Class Multimodal LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:38.410248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:38.410248Z digest=sha256:8988d18a14648aa8a8b18fb2142642f835d90f41829aec3d1cbc5b6256933619

Observation 7ab4640c-af0a-4d83-bb6a-defc78e4bcb7 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Vision-Language Models Can't See the Obvious Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:38.476458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:38.476458Z digest=sha256:2afe0afe2feec3912f42637d9b8585c14ed7f6850181df6b91e6bce1bd1c4443

Observation f8edecab-ee5a-4e73-af10-f8528b260f00 · outbound

This paper cites The Llama 3 Herd of Models.

Vision-Language Models Can't See the Obvious The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:38.533337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:38.533337Z digest=sha256:7e4e3e4fb76c3307b51555612c47c330672ddc9f963c2f73635819c28e23a13d

Observation 93ac14f5-a627-458c-9e9d-2e6e900a1aa0 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Vision-Language Models Can't See the Obvious MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:38.555170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:38.555170Z digest=sha256:1f51b9b66a2179f30be17b631157cc188f5346a6e76ef607b8fedb346583ff92

Observation a852a5ac-b1c6-4c94-abb1-3a3a1798a530 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Vision-Language Models Can't See the Obvious Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:47.824009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:38.683920Z digest=sha256:f86ff1a94b8b71c60ec973379d1f129b651addb22fe6af59a14d39f7c6b3eda4

Observation 69e60933-1d29-4596-96be-06efc1af00a7 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Vision-Language Models Can't See the Obvious Vizwiz grand challenge: Answering visual questions from blind people

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:47.666770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:38.820165Z digest=sha256:41b890faff7bf10a3396520393b218aaebc1b8573dcf71cbd034abdd29c55ad4

Observation 4fd64c2a-609c-4055-9494-2344bed1e8c9 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Vision-Language Models Can't See the Obvious Cogagent: A visual language model for gui agents

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:47.543544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:38.933059Z digest=sha256:7466cebd6473ec3cfe779a6735caf92a6fbbc9e876f3fb147a8a998ef294d4ca

Observation b79a28f9-eb19-423d-8203-96c7161629a1 · outbound

This paper cites Minicpm: Un- veiling the potential of small language models with scalable training strategies.

Vision-Language Models Can't See the Obvious Minicpm: Un- veiling the potential of small language models with scalable training strategies

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:44:39.020480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:39.020480Z digest=sha256:0030f29f7aca8bd1fdb5411379a29df7a77f5425ba3603ac5aa975ef12612193

Observation 4185bdb5-6312-4327-902f-77d227b6d58a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and com- positional question answering.

Vision-Language Models Can't See the Obvious Gqa: A new dataset for real-world visual reasoning and com- positional question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:47.320871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:39.083159Z digest=sha256:c9d7e61f940e393c76b938fde3e34b3c1d441e631bf963af569aeb9125d2421b

Observation eaf7e95c-5853-4744-88c7-3a63c6dd5ddd · outbound

This paper cites A diagram is worth a dozen images.

Vision-Language Models Can't See the Obvious A diagram is worth a dozen images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:47.138713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:39.183198Z digest=sha256:cfb8e2ea51b134328490feff1552b35daa80b3097fedc72a491810bc5193969d

Observation 70d38a84-c250-49c0-9dcc-e97c8efaaca8 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Vision-Language Models Can't See the Obvious OpenVLA: An Open-Source Vision-Language-Action Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:39.353436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:39.353436Z digest=sha256:68e926a94529274c66daf526d31a23237bebfc738c90307d4bfb3004cf56a540

Observation 7baea62f-772e-4f11-8f89-19d417b2316b · outbound

This paper cites Do Saliency Models Detect Odd-One-Out Targets? New Datasets and Evaluations.

Vision-Language Models Can't See the Obvious Do Saliency Models Detect Odd-One-Out Targets? New Datasets and Evaluations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:39.472454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:39.472454Z digest=sha256:3a58928564550f325c06e45ab881406b1c28b23e98432de37aa5fe5ec4985a41

Observation 280158b9-84d2-4448-93ab-a08974e26520 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

Vision-Language Models Can't See the Obvious Building and better understanding vision-language models: insights and future directions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:39.601722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:39.601722Z digest=sha256:90ac6e907a1d5bec5f7eeb42560e3fdc24a4d904e9c1a45a6a3a6310a2059eae

Observation fc82e7c1-710b-488b-8e50-98a377738b7c · outbound

This paper cites What matters when building vision-language models?.

Vision-Language Models Can't See the Obvious What matters when building vision-language models?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:39.806355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:39.806355Z digest=sha256:1b3120ab6582ff7f864da34651f573689470320adf4e06a9197e1abcc0864d15

Observation 2c3af3bd-b253-46e6-92b9-a6814b72ecb6 · outbound

This paper cites Seed- bench: Benchmarking multimodal large language models.

Vision-Language Models Can't See the Obvious Seed- bench: Benchmarking multimodal large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:46.988909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:39.925626Z digest=sha256:f685912044cbb5a52a556a1f684b4bc390a83d6a4d4703266081587d97db7086

Observation d36e798a-9e8f-4642-b26c-6f4d23a73377 · outbound

This paper cites Evaluating ob- ject hallucination in large vision-language models.

Vision-Language Models Can't See the Obvious Evaluating ob- ject hallucination in large vision-language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:46.839109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:40.048744Z digest=sha256:5b8a801e162d257a39674104ccd4b290bf6a3489cd0ac34714bf457f8744a22a

Observation b11e4fac-9660-4513-b0f3-dbb799e01229 · outbound

This paper cites Vila: On pre- training for visual language models.

Vision-Language Models Can't See the Obvious Vila: On pre- training for visual language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:46.627895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:40.212638Z digest=sha256:cb13b6c01508b4b6070aa4c0c06a462cdf11ecc01308a4044b13791662324b3e

Observation 0172889d-6d31-4f20-81f1-ba300481fc30 · outbound

This paper cites Microsoft coco: Com- mon objects in context.

Vision-Language Models Can't See the Obvious Microsoft coco: Com- mon objects in context

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:46.456530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:40.319698Z digest=sha256:37945a56624962f0dd3d6956e7d7717bb5a4253a0926ea20d9426eaea5f720ec

Observation cf61eb70-d750-4e9b-88d3-ec67164caa64 · outbound

This paper cites Visual instruction tuning.

Vision-Language Models Can't See the Obvious Visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:46.228014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:40.458836Z digest=sha256:c215ae156ed23c0818c546db84a2facdb647966db70e73b6e7f5e212c2c7b7b4

Observation 94263374-54bb-438d-8038-f7c4efd109f1 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Vision-Language Models Can't See the Obvious MMBench: Is Your Multi-modal Model an All-around Player?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:40.616829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:40.616829Z digest=sha256:9dc62b2d772195c28ca64b7b97b0ed96a92a4d75af937734d9505ec509ca1c0f

Observation 99d0bad0-eda6-4edf-9cfc-0bbff6d80ae8 · outbound

This paper cites Learn to explain: Multi- modal reasoning via thought chains for science ques- tion answering.

Vision-Language Models Can't See the Obvious Learn to explain: Multi- modal reasoning via thought chains for science ques- tion answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:46.062014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:40.734841Z digest=sha256:5325b3906a1a4e930fa301545330380761b6276e94015b773c9346ebc880be49

Observation 8a45a6f8-99c7-4389-8bfd-b67c878aea0e · outbound

This paper cites Mathvista: Evaluating math reasoning in visual contexts with gpt- 4v, bard, and other large multimodal models.

Vision-Language Models Can't See the Obvious Mathvista: Evaluating math reasoning in visual contexts with gpt- 4v, bard, and other large multimodal models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:45.903231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:40.859398Z digest=sha256:650be737874873e1376f2371526fdf514e1a179abc4f04f285dd55eed0239834

Observation f1d3dbf4-af42-4153-aa0f-aadd4f58b836 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Vision-Language Models Can't See the Obvious ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:41.014044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:41.014044Z digest=sha256:7f7bdecd535945aec5373ac6097d768eab55c4489356b134f2d47f57c5cdd5d9

Observation c7d4289e-651b-4e57-8e0c-6df8e3cd5099 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Vision-Language Models Can't See the Obvious Docvqa: A dataset for vqa on document images

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:45.723442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.081835Z digest=sha256:94adb873283dea74aa327978e0619eb5935bf10825ddb2edb2a37bca84f24f74

Observation d07c4494-5902-40ae-9b15-7bc080e31455 · outbound

This paper cites Info- graphicvqa.

Vision-Language Models Can't See the Obvious Info- graphicvqa

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:45.559020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.166811Z digest=sha256:4083f81de80d9b591af172ad219b45f9ec16eb606473e79935c4351fbb7bb1de

Observation 763fe03c-eabb-4a20-b96d-b791228021e5 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vi- sion with open, customizable modelsy.

Vision-Language Models Can't See the Obvious Llama 3.2: Revolutionizing edge ai and vi- sion with open, customizable modelsy

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:45.442472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.267499Z digest=sha256:52f4bcd4679ec07d0b165ad10e845c04371c455396948789e372edd64cee032a

Observation cdf1c0a9-bec6-4f2f-94d9-7a191168e8a4 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Vision-Language Models Can't See the Obvious Ocr-vqa: Visual question answering by reading text in images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:45.283847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.334651Z digest=sha256:7432ba2c508c452444eaa187d1159a42b673506a5d716fdc089afc9082626c18

Observation 77168ac6-96d4-4a6f-b57e-209e8e3f7c67 · outbound

This paper cites Mind children: The future of robot and human intelligence.

Vision-Language Models Can't See the Obvious Mind children: The future of robot and human intelligence

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:45.147605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.401301Z digest=sha256:7152f7ebffebbff41caf75778adaff06c418988257b912591f708280c911bfda

Observation f9af5b3d-b0b3-4a02-ab56-90d1d4907437 · outbound

This paper cites ScreenAgent: A Vision Language Model-driven Computer Control Agent.

Vision-Language Models Can't See the Obvious ScreenAgent: A Vision Language Model-driven Computer Control Agent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:41.444033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:41.444033Z digest=sha256:99456e58065a1580b2e028c14b963f59322914ab6f77855937b414d2b5f18ef1

Observation cb40e48f-4ad3-404b-a693-273af561f915 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Can't See the Obvious Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:44:44.991913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.489524Z digest=sha256:4ce1844c7c843b06dd1afb096e54d5c8984c1906190aca4ea3beeb74eb0a2f6b

Observation 30a73b3a-b0c3-4f3f-be7b-5d710f83ddaa · outbound

This paper cites Learning transferable visual models from natural language supervision.

Vision-Language Models Can't See the Obvious Learning transferable visual models from natural language supervision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:44.845666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.539206Z digest=sha256:fe77b43c32d7223ae9da91e893727dafedf7747ebc7915da14f29abe6f6baa50

Observation 53770063-7796-46d4-a93e-6356347c5a43 · outbound

This paper cites A- okvqa: A benchmark for visual question answering using world knowledge.

Vision-Language Models Can't See the Obvious A- okvqa: A benchmark for visual question answering using world knowledge

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:44.699076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.596641Z digest=sha256:752de6eb15f231f2de81a9011d534693032dfe10b9accdd8757e64babfa054f4

Observation 58ef7527-0c2e-4c45-bc2b-9d75b2604e1e · outbound

This paper cites Towards vqa models that can read.

Vision-Language Models Can't See the Obvious Towards vqa models that can read

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:44.559632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.649314Z digest=sha256:d1ea654c5e473bf03fe642ab781248e0d94be09cd4fe80638f329a1b8d79ebc7

Observation b7ceb6ce-2ee2-47d9-b46d-23e11e69a6c0 · outbound

This paper cites Internvl2: Better than the best — ex- panding performance boundaries of open-source mul- timodal models with the progressive scaling strat- egy.

Vision-Language Models Can't See the Obvious Internvl2: Better than the best — ex- panding performance boundaries of open-source mul- timodal models with the progressive scaling strat- egy

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:44.437364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.701633Z digest=sha256:9a78b867f85f536fc92e97b7f150ede81d8eda60d8f8ddedd9943552b2dc7100

Observation 7b86b4c5-7926-4b88-a168-55593c6c5e42 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Vision-Language Models Can't See the Obvious Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:41.746239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:41.746239Z digest=sha256:22dcb81456b217be390859085e60c0ea680163d558e11dc06ed02a3722240271

Observation b467029b-f322-49e9-b0ea-64172cfb66a6 · outbound

This paper cites Eyes wide shut? ex- ploring the visual shortcomings of multimodal llms.

Vision-Language Models Can't See the Obvious Eyes wide shut? ex- ploring the visual shortcomings of multimodal llms

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:44.320333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.796290Z digest=sha256:65e91e28c2c14bf74804b72a5bd3b360aa08335b0dc0af0bd15e11adc8748af1

Observation 21317282-71bd-4a2f-9b38-e93fbf8e20c9 · outbound

This paper cites A feature- integration theory of attention.

Vision-Language Models Can't See the Obvious A feature- integration theory of attention

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:44.202666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.870184Z digest=sha256:6040d29f379d5c46c619f2f2ef05d6d3c708ee49683d0257afd0e0320f568b09

Observation 8b0563c5-8821-43c4-85a9-cadf84cee1ad · outbound

This paper cites Measuring mul- timodal mathematical reasoning with math-vision dataset, 2024.

Vision-Language Models Can't See the Obvious Measuring mul- timodal mathematical reasoning with math-vision dataset, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:44.074777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:41.933031Z digest=sha256:6c7444b4b0a771bcd10ad0bfbdf6455247c3a94d016d427f8fd35b2413173f6f

Observation 84e9e85b-4037-4e21-a62f-8039d6d8de67 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Vision-Language Models Can't See the Obvious Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:41.992244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:41.992244Z digest=sha256:155b7afbc2aaf4008dddc8e46da56c88265323e7a62aa80ea90eac28e5e452b5

Observation d9ab4423-8c2b-48e9-81dd-9e660a6bc669 · outbound

This paper cites Five factors that guide attention in visual search.

Vision-Language Models Can't See the Obvious Five factors that guide attention in visual search

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:43.967292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:42.046127Z digest=sha256:9cf7cc1b8e7cc6a3184a4b9c87890224533d6cd394215984920d709dcfa8e4ba

Observation 8e1fbd81-329d-4957-8a5a-bcce8c3636e8 · outbound

This paper cites Grok-1.5v.

Vision-Language Models Can't See the Obvious Grok-1.5v

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:43.805297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:42.099279Z digest=sha256:4048587439999fe8d42428c0c5a5a07a2a8961d168f6e7131c695f2f3aefeb39

Observation 55856718-6c7f-4364-b869-accb4f33b636 · outbound

This paper cites Qwen2 Technical Report.

Vision-Language Models Can't See the Obvious Qwen2 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:42.166888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:42.166888Z digest=sha256:cca855068cc547dd33ccb1ffe46ba6e1d45d666a4c0155d2593decafdc22cbbe

Observation 6e3da194-c577-4740-b461-eb25861eeb62 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Vision-Language Models Can't See the Obvious MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:42.219769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:42.219769Z digest=sha256:35944061feeb71352827916463022776f3e5d4e7f186e5429cb07ed9336e9bc5

Observation da020f1b-8afb-4cac-b3a2-f168d68970d8 · outbound

This paper cites Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Vision-Language Models Can't See the Obvious Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:43.649705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:42.300156Z digest=sha256:044a34cc26def3868d664483563aecc2519084baed908b600e2ce55e33409c28

Observation a431cb4e-c191-4aa1-8326-56c07894d270 · outbound

This paper cites Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Vision-Language Models Can't See the Obvious Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:43.503591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:42.377530Z digest=sha256:9f409c7a0ff54d8a043abafe3f384b562d82a874d7b641a9c531f4fa1f77c4ab

Observation 4fbb2be1-551e-4ca5-a347-103e3972404a · outbound

This paper cites Sigmoid loss for language im- age pre-training.

Vision-Language Models Can't See the Obvious Sigmoid loss for language im- age pre-training

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:43.392391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:42.440463Z digest=sha256:741226f3e70b79b7169818a625135aeb9026bc1e404f227fff2381f2fa009499

Observation ef59860e-698c-4fba-934e-a74aba7d0f62 · outbound

This paper cites Se- mantic understanding of scenes through the ade20k dataset.

Vision-Language Models Can't See the Obvious Se- mantic understanding of scenes through the ade20k dataset

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:44:43.236260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:42.518402Z digest=sha256:78a0729cf6cb9b43c4a1d8b407719291acbae265540d67d6351564679bab23fa

Observation 680e15e7-5b94-4da6-b8d0-d018b9073a8a · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Vision-Language Models Can't See the Obvious MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:42.599032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:42.599032Z digest=sha256:b0d13b211f0c166091918dc4dac0ea4d421a0781c4bb6a2c71d5df98c7a80478

Observation 95e07439-2c88-47c1-a2bd-5328a736e6a6 · outbound

This paper cites The next most common range is > 25 distractors, with a similar count to the lowest range, reflecting the dataset’s coverage of highly complex scenarios.

Vision-Language Models Can't See the Obvious The next most common range is > 25 distractors, with a similar count to the lowest range, reflecting the dataset’s coverage of highly complex scenarios

Reference 600

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T19:44:43.079147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T19:44:42.681512Z digest=sha256:21a87b09b1d63637f6b354148a339b4b3bb5eb9b48dfe82d9f201d869fd28f84

Pith citing papers

Observation f6045e2f-5afe-426b-82fb-a8aead0e969b · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Vision-Language Models Can't See the Obvious

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:47.936124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:8da977323dc91762f4b4ef40a4e3811e66b373f7e9719187b2fd9d62ab166323

Observation c44bf1c6-25b3-4166-9b95-036afa91f29a · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs Vision-Language Models Can't See the Obvious

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.291857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:a7bbe8e05c5ce8f7e5e1c808cf1cff32fc30078adc60a6937853d2969d959fcb