Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T05:55:09.188048Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 2 inbound Pith citation observations for arXiv:2507.01955.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T05:55:09.188048Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T23:12:32.101737Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-28T23:12:46.506585Z
78 of 78 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation af1fc42c-2f6e-4ea4-93bc-781bf5de67c6 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Slic superpixels compared to state-of-the-art superpixel methods
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c552708-69fc-4795-a042-9e9a2c03f8c4 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d644d0d-8561-4e14-a800-6aca7549da9c · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks The llama 3 herd of models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc4905d7-0b8d-4292-9cb9-a399ab984a99 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76698047-1279-4ede-9909-1ff84911fde7 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Flamingo: a visual language model for few-shot learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 796457a1-fa09-4710-a41e-c413d6899df4 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Introducing claude 3.5 sonnet
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4f32294-4695-4628-aaf1-deb3244d8dec · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 849e7cf6-7fe5-4925-a38a-ed56f0d71f87 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Qwen Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd9afcaa-8d04-42c8-8195-84a29332c955 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks PaliGemma: A versatile 3B VLM for transfer
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2857d08f-9d20-46ea-ba78-b043d324ede8 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Omni3d: A large benchmark and model for 3d object detection in the wild
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 925f2404-1be5-44df-a0cb-75b8b1d9e772 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks End-to- end object detection with transformers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8337363-346b-47ab-b363-adc6f363a953 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Evaluating Large Language Models Trained on Code
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eb24732f-0bf0-418c-b2d8-fe4667ea7b09 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks An Empirical Study of GPT-4o Image Generation Capabilities
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ee451af-a98f-4a60-be4a-ed43905084ed · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Self-icl: Zero-shot in-context learning with self- generated demonstrations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5a53400-4bc9-47b3-a2b8-01c4f7ac540a · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Reproducible scal- ing laws for contrastive language-image learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c39fabb-892a-4e64-8eb7-441385c8b786 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d4ad231-3973-4dce-ad72-6f8e6e600c80 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks RobustBench: a standardized adversarial robustness benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95dcfcdb-207a-4700-a145-00ca4333cc3c · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 730fb3d8-ed13-4f0e-8647-67f0203e295b · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c0671321-6117-43a8-ab20-b43f891fbbbb · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Find your inspiration
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84230edf-4209-4308-9141-436bcd19977b · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c79d67e7-4ba0-4d8c-8528-915765975396 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Explore vision capabilities with the gemini api
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ade1cdb5-0570-4e03-a98d-7f347cf903ee · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Gemini 2.0 flash
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6bbb666f-1d9f-4f5d-b1eb-55311d29b474 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df572c0c-a2e5-4fd4-86d2-271579942874 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Measuring Massive Multitask Language Understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3236a1d4-eee7-4b55-8888-743ea23c65d6 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks The many faces of robust- ness: A critical analysis of out-of-distribution generalization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ead8605-e0da-4ea8-a98c-dd2ddaf30320 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 02e8f7a4-7cd6-4b78-8421-92cceb6dec0c · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Llama multiple images
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8948bcd2-cb98-4b5b-a743-8de1248f04a9 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0050aef9-85f8-41d0-b6fb-84e79ea0f2d6 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Oneformer: One transformer to rule universal image segmentation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 973035ca-77dd-43cd-8391-1e740949f05d · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Many-Shot In-Context Learning in Multimodal Foundation Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8a5010f-9722-4133-bb60-b2db95d3b2ec · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chen, and Andrew Y
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4108af43-0196-432a-8b3b-576dadfce6b7 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks 3d common corruptions and data augmentation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3e3db48-7e1c-46c8-b4c3-c1a86f15a2ed · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks 3d com- mon corruptions for object recognition
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 036c8900-e6b2-4266-9847-e8efc8e45b8b · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Decomposed Prompting: A Modular Approach for Solving Complex Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c62bba2-d900-44ba-aeee-2bb0cc87b4bb · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Segment Anything
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 87106f7e-ac80-4372-9227-e5d629d9cbbf · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d8bfc312-e3fa-462e-8b82-cc85b5c18280 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Evaluating Object Hallucination in Large Vision-Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31b85b5e-3e4a-4ca4-a6e6-8d011e5bd977 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2633707-72e5-45a9-a7cd-99ca8e7ceba8 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Llava-next: Improved reason- ing, ocr, and world knowledge
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b60ef790-8b20-4ad7-96ea-a956880e5b0b · outbound
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f21c826-b6a9-46b9-adb4-1b22d7d3983c · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks 4M: Massively multimodal masked modeling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation caa1614e-ebf5-4ff0-9589-84e51b30f0ec · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Introducing openai o1
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8d810ea7-01d5-48ed-bba2-64f380570a8e · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Hello gpt-4o
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bde9b641-06a3-4a78-9d7e-5f65a9716d32 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Introducing 4o image generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7521322-5fc0-41de-8eb0-ad9449d1db3e · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Introducing openai o3 and o4-mini
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8bd60bd1-6dd6-41cf-8065-10bf0bc2c12c · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Vision language models are blind: Failing to translate detailed visual features into words
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f276678e-f04e-4304-96e6-e1db2b28a8c7 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8be94be3-92eb-45b7-be9d-0d662b598277 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bfa972c0-e570-4f0d-86ef-a64c9340edf0 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3baaf630-7cc1-4ff2-8d78-9f99823937a8 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f59fb832-03c7-465f-9852-983109d65a38 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Bernstein, Alexander C
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2a0a9a7d-cfb2-4517-9bf2-0bb3380aec2b · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Super- pixels: An evaluation of the state-of-the-art
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b46d22b-83c9-4b4a-a96e-162b889799e0 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4b59d9c-d342-4c57-a817-5f8602939e42 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Gemini: A Family of Highly Capable Multimodal Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1f2549cc-fb7c-44bc-95a2-27ecc561455e · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ffedd751-f1f1-40a0-92bd-419d0354ae2e · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cfcc742-0105-4fbb-bced-2d8361878310 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks The internet’s source for visuals
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 056c4c45-ed18-4d98-a27d-5da1f431d869 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Learning robust global representations by penalizing local predictive power
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f5817ac-39fd-43d8-8f80-93b363ef8914 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 213f451e-94d2-463f-99cf-2529fedae323 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ac40667-6aaa-441a-815a-faca4caf20a6 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chain-of- thought prompting elicits reasoning in large language models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae4d69db-d967-4b4b-95a1-1a9f2762b48b · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 69876831-09f2-44a3-b370-55ab9781abd9 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks V*: Guided visual search as a core mechanism in multimodal llms
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5e57172-4c45-4398-9a9c-3339072f57b7 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a23c936f-d569-4cab-8394-1ceed439aaf2 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 47e5d366-d3d8-48bd-94a6-4a69f495f551 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks The dawn of lmms: Preliminary explorations with gpt-4v(ision)
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0db5fbe9-c484-4b21-b977-62af7bc7cfb9 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Tree of thoughts: Deliberate problem solving with large language models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 376d96c5-4802-4c08-8b5a-3b4ffcd3b0af · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks A Survey on Multimodal Large Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af8d3448-6315-4891-abf2-c2aab21c68cf · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 822a7497-64f7-4414-9d5b-6cf82546187f · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks MM-LLMs: Recent Advances in MultiModal Large Language Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08aebd87-c1cb-4bf5-b47d-a4a3ffec27dc · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Scene parsing through ADE20K dataset
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 12eea387-8707-4645-84e3-3078f699f366 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96d4ef64-4828-49f6-9958-e2f55a64254f · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Detrs with collab- orative hybrid assignments training
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a25b1049-d68a-4101-a9fa-985b3c5b141b · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Learning ordinal relationships for mid-level vi- sion
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0615429-3242-496d-87d8-516a40145772 · outbound
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4da8a1cd-71eb-4304-a4b9-4ed8cf720607 · outbound
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks - Produce a raw image (same dimensions as input) whose pixel colors encode the normals as above
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3310b469-7a3f-417f-a91d-d20ba3c26157 · outbound
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f10f806c-317f-4939-8b37-230e2c825177 · inbound
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15d0897b-ccfe-4f79-a439-0d17abd1a35f · inbound
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.