Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:44:42.681512Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2507.04741.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:44:42.681512Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T19:59:19.379119Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T10:06:03.287917Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd12751b-47ca-4096-a191-89bda0a7648c · outbound
Vision-Language Models Can't See the Obvious Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14030453-a82d-4e78-8056-cc33a3e8e7f4 · outbound
Vision-Language Models Can't See the Obvious GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c84e1d8b-fd01-4970-8d70-d88fc1c91ee9 · outbound
Vision-Language Models Can't See the Obvious Nocaps: Novel object captioning at scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9902289-7f7d-4263-81bb-a492e0057036 · outbound
Vision-Language Models Can't See the Obvious Flamingo: a visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9245c91e-3f91-4a07-a082-51bfa2d0ad38 · outbound
Vision-Language Models Can't See the Obvious Claude 3.5 sonnet
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 01da0eb5-e139-462a-8a73-b0c20fedacac · outbound
Vision-Language Models Can't See the Obvious Turning visual search time on its head
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6e58f10e-1d99-417a-a725-3919fb13421b · outbound
Vision-Language Models Can't See the Obvious PaliGemma: A versatile 3B VLM for transfer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ec1332-0188-49f3-8417-6e2c4ea76d3b · outbound
Vision-Language Models Can't See the Obvious Scene text visual question answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation acaf3167-99b5-4a3c-9459-617788618ca5 · outbound
Vision-Language Models Can't See the Obvious $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2421e0-e15e-4496-ba4c-52e29bad11e6 · outbound
Vision-Language Models Can't See the Obvious Omni3d: A large benchmark and model for 3d object detection in the wild
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a3e07d3b-54ea-481e-993b-17d919c953f4 · outbound
Vision-Language Models Can't See the Obvious Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66303033-7847-46cf-9af2-7a3889c0c8ff · outbound
Vision-Language Models Can't See the Obvious Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b53fa82b-0c77-4ecf-9072-06482e2ac240 · outbound
Vision-Language Models Can't See the Obvious NVLM: Open Frontier-Class Multimodal LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab4640c-af0a-4d83-bb6a-defc78e4bcb7 · outbound
Vision-Language Models Can't See the Obvious Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8edecab-ee5a-4e73-af10-f8528b260f00 · outbound
Vision-Language Models Can't See the Obvious The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93ac14f5-a627-458c-9e9d-2e6e900a1aa0 · outbound
Vision-Language Models Can't See the Obvious MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a852a5ac-b1c6-4c94-abb1-3a3a1798a530 · outbound
Vision-Language Models Can't See the Obvious Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 69e60933-1d29-4596-96be-06efc1af00a7 · outbound
Vision-Language Models Can't See the Obvious Vizwiz grand challenge: Answering visual questions from blind people
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4fd64c2a-609c-4055-9494-2344bed1e8c9 · outbound
Vision-Language Models Can't See the Obvious Cogagent: A visual language model for gui agents
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b79a28f9-eb19-423d-8203-96c7161629a1 · outbound
Vision-Language Models Can't See the Obvious Minicpm: Un- veiling the potential of small language models with scalable training strategies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4185bdb5-6312-4327-902f-77d227b6d58a · outbound
Vision-Language Models Can't See the Obvious Gqa: A new dataset for real-world visual reasoning and com- positional question answering
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eaf7e95c-5853-4744-88c7-3a63c6dd5ddd · outbound
Vision-Language Models Can't See the Obvious A diagram is worth a dozen images
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 70d38a84-c250-49c0-9dcc-e97c8efaaca8 · outbound
Vision-Language Models Can't See the Obvious OpenVLA: An Open-Source Vision-Language-Action Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7baea62f-772e-4f11-8f89-19d417b2316b · outbound
Vision-Language Models Can't See the Obvious Do Saliency Models Detect Odd-One-Out Targets? New Datasets and Evaluations
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 280158b9-84d2-4448-93ab-a08974e26520 · outbound
Vision-Language Models Can't See the Obvious Building and better understanding vision-language models: insights and future directions
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc82e7c1-710b-488b-8e50-98a377738b7c · outbound
Vision-Language Models Can't See the Obvious What matters when building vision-language models?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3af3bd-b253-46e6-92b9-a6814b72ecb6 · outbound
Vision-Language Models Can't See the Obvious Seed- bench: Benchmarking multimodal large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d36e798a-9e8f-4642-b26c-6f4d23a73377 · outbound
Vision-Language Models Can't See the Obvious Evaluating ob- ject hallucination in large vision-language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b11e4fac-9660-4513-b0f3-dbb799e01229 · outbound
Vision-Language Models Can't See the Obvious Vila: On pre- training for visual language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0172889d-6d31-4f20-81f1-ba300481fc30 · outbound
Vision-Language Models Can't See the Obvious Microsoft coco: Com- mon objects in context
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cf61eb70-d750-4e9b-88d3-ec67164caa64 · outbound
Vision-Language Models Can't See the Obvious Visual instruction tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 94263374-54bb-438d-8038-f7c4efd109f1 · outbound
Vision-Language Models Can't See the Obvious MMBench: Is Your Multi-modal Model an All-around Player?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d0bad0-eda6-4edf-9cfc-0bbff6d80ae8 · outbound
Vision-Language Models Can't See the Obvious Learn to explain: Multi- modal reasoning via thought chains for science ques- tion answering
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8a45a6f8-99c7-4389-8bfd-b67c878aea0e · outbound
Vision-Language Models Can't See the Obvious Mathvista: Evaluating math reasoning in visual contexts with gpt- 4v, bard, and other large multimodal models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f1d3dbf4-af42-4153-aa0f-aadd4f58b836 · outbound
Vision-Language Models Can't See the Obvious ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d4289e-651b-4e57-8e0c-6df8e3cd5099 · outbound
Vision-Language Models Can't See the Obvious Docvqa: A dataset for vqa on document images
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d07c4494-5902-40ae-9b15-7bc080e31455 · outbound
Vision-Language Models Can't See the Obvious Info- graphicvqa
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 763fe03c-eabb-4a20-b96d-b791228021e5 · outbound
Vision-Language Models Can't See the Obvious Llama 3.2: Revolutionizing edge ai and vi- sion with open, customizable modelsy
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cdf1c0a9-bec6-4f2f-94d9-7a191168e8a4 · outbound
Vision-Language Models Can't See the Obvious Ocr-vqa: Visual question answering by reading text in images
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 77168ac6-96d4-4a6f-b57e-209e8e3f7c67 · outbound
Vision-Language Models Can't See the Obvious Mind children: The future of robot and human intelligence
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f9af5b3d-b0b3-4a02-ab56-90d1d4907437 · outbound
Vision-Language Models Can't See the Obvious ScreenAgent: A Vision Language Model-driven Computer Control Agent
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb40e48f-4ad3-404b-a693-273af561f915 · outbound
Vision-Language Models Can't See the Obvious Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 30a73b3a-b0c3-4f3f-be7b-5d710f83ddaa · outbound
Vision-Language Models Can't See the Obvious Learning transferable visual models from natural language supervision
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 53770063-7796-46d4-a93e-6356347c5a43 · outbound
Vision-Language Models Can't See the Obvious A- okvqa: A benchmark for visual question answering using world knowledge
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 58ef7527-0c2e-4c45-bc2b-9d75b2604e1e · outbound
Vision-Language Models Can't See the Obvious Towards vqa models that can read
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b7ceb6ce-2ee2-47d9-b46d-23e11e69a6c0 · outbound
Vision-Language Models Can't See the Obvious Internvl2: Better than the best — ex- panding performance boundaries of open-source mul- timodal models with the progressive scaling strat- egy
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7b86b4c5-7926-4b88-a168-55593c6c5e42 · outbound
Vision-Language Models Can't See the Obvious Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b467029b-f322-49e9-b0ea-64172cfb66a6 · outbound
Vision-Language Models Can't See the Obvious Eyes wide shut? ex- ploring the visual shortcomings of multimodal llms
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 21317282-71bd-4a2f-9b38-e93fbf8e20c9 · outbound
Vision-Language Models Can't See the Obvious A feature- integration theory of attention
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8b0563c5-8821-43c4-85a9-cadf84cee1ad · outbound
Vision-Language Models Can't See the Obvious Measuring mul- timodal mathematical reasoning with math-vision dataset, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 84e9e85b-4037-4e21-a62f-8039d6d8de67 · outbound
Vision-Language Models Can't See the Obvious Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ab4423-8c2b-48e9-81dd-9e660a6bc669 · outbound
Vision-Language Models Can't See the Obvious Five factors that guide attention in visual search
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e1fbd81-329d-4957-8a5a-bcce8c3636e8 · outbound
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55856718-6c7f-4364-b869-accb4f33b636 · outbound
Vision-Language Models Can't See the Obvious Qwen2 Technical Report
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3da194-c577-4740-b461-eb25861eeb62 · outbound
Vision-Language Models Can't See the Obvious MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da020f1b-8afb-4cac-b3a2-f168d68970d8 · outbound
Vision-Language Models Can't See the Obvious Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a431cb4e-c191-4aa1-8326-56c07894d270 · outbound
Vision-Language Models Can't See the Obvious Mmmu: A mas- sive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4fbb2be1-551e-4ca5-a347-103e3972404a · outbound
Vision-Language Models Can't See the Obvious Sigmoid loss for language im- age pre-training
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ef59860e-698c-4fba-934e-a74aba7d0f62 · outbound
Vision-Language Models Can't See the Obvious Se- mantic understanding of scenes through the ade20k dataset
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 680e15e7-5b94-4da6-b8d0-d018b9073a8a · outbound
Vision-Language Models Can't See the Obvious MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e07439-2c88-47c1-a2bd-5328a736e6a6 · outbound
Vision-Language Models Can't See the Obvious The next most common range is > 25 distractors, with a similar count to the lowest range, reflecting the dataset’s coverage of highly complex scenarios
Reference 600
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f6045e2f-5afe-426b-82fb-a8aead0e969b · inbound
Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Vision-Language Models Can't See the Obvious
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c44bf1c6-25b3-4166-9b95-036afa91f29a · inbound
Counting to Four is still a Chore for VLMs Vision-Language Models Can't See the Obvious
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.