Pith. sign in

Paper Citation Record · LEDGER

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

As of 6 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 2 inbound Pith citation observations for arXiv:2507.01955.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01955 v3

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T05:55:09.188048Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:12:32.101737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-28T23:12:46.506585Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact32
  • verified fuzzy43
  • unresolved0
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af1fc42c-2f6e-4ea4-93bc-781bf5de67c6 · outbound

This paper cites Slic superpixels compared to state-of-the-art superpixel methods.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Slic superpixels compared to state-of-the-art superpixel methods

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.633297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:0a685688f54f913a818fd40a8d45aec74dc56b6d8d10b034de7e56262919dde5

Observation 4c552708-69fc-4795-a042-9e9a2c03f8c4 · outbound

This paper cites GPT-4 Technical Report.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.152595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:f91aa36f97f33a6ffdc852aaf4cdd1b2276191f716419f7809c8b6b0ff2abc6c

Observation 9d644d0d-8561-4e14-a800-6aca7549da9c · outbound

This paper cites The llama 3 herd of models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks The llama 3 herd of models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.636410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:554a3b72447cf625f696c7ea50f97f4a03160ea660d23654134c24be303faa39

Observation cc4905d7-0b8d-4292-9cb9-a399ab984a99 · outbound

This paper cites UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.159043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:61d9482565743ccb94c5a3004a9ad6f3a6d55a6415a6883754857b2c65c06716

Observation 76698047-1279-4ede-9909-1ff84911fde7 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.630048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:b15f84cc6ea32a989bcc072cd90b38b7b441ea5fd29afd9ac04d64e4563aeca5

Observation 796457a1-fa09-4710-a41e-c413d6899df4 · outbound

This paper cites Introducing claude 3.5 sonnet.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Introducing claude 3.5 sonnet

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.627137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:ded059606f8dda4c832b28dfcc75ebc43fa2ad1802a00c2e5a4f3e630d24e6ce

Observation f4f32294-4695-4628-aaf1-deb3244d8dec · outbound

This paper cites 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.142513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:987cc8ea2311f36207c8e51c2ab040ce799a168df2a97a8ef2f0e8b06cd549b5

Observation 849e7cf6-7fe5-4925-a38a-ed56f0d71f87 · outbound

This paper cites Qwen Technical Report.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Qwen Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.137262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:8e74afaa10dc7aaa6f6ac2a28a7e9ce13e91e98fa65885fa8906551b6badb28e

Observation dd9afcaa-8d04-42c8-8195-84a29332c955 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks PaliGemma: A versatile 3B VLM for transfer

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.147550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:ca3b971c460182f50a8a42487a3417250a4b1cf536e6e994cd6d98d7cda81d23

Observation 2857d08f-9d20-46ea-ba78-b043d324ede8 · outbound

This paper cites Omni3d: A large benchmark and model for 3d object detection in the wild.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Omni3d: A large benchmark and model for 3d object detection in the wild

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.734208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:94e850470a1c915344f1cffa1f3a7ca55f599e5aea3c95773029d7026f53a2d7

Observation 925f2404-1be5-44df-a0cb-75b8b1d9e772 · outbound

This paper cites End-to- end object detection with transformers.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks End-to- end object detection with transformers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.737167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:3b1614e74975b7df67656cc89123d65ac9fc9d88bded503a94903d321ce943f8

Observation d8337363-346b-47ab-b363-adc6f363a953 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Evaluating Large Language Models Trained on Code

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.070966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:df973be3700de51d984afe1618ee80995853a1c2b45278313658fd304154359b

Observation eb24732f-0bf0-418c-b2d8-fe4667ea7b09 · outbound

This paper cites An Empirical Study of GPT-4o Image Generation Capabilities.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks An Empirical Study of GPT-4o Image Generation Capabilities

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.065260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:024350f7ff88f22257a2a19ddcb4bde191647de3a0e17f9578f47fa6718b7b1f

Observation 2ee451af-a98f-4a60-be4a-ed43905084ed · outbound

This paper cites Self-icl: Zero-shot in-context learning with self- generated demonstrations.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Self-icl: Zero-shot in-context learning with self- generated demonstrations

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.711639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:95ddf1fe24e226174a7e7b4d96579748c5e0b2a35b4903010cec1df2a60aa160

Observation a5a53400-4bc9-47b3-a2b8-01c4f7ac540a · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Reproducible scal- ing laws for contrastive language-image learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.728328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:7aa5fd9f3f46c99cb83211e53ae41588aeed04f2c0f92b88d90bddb1a7a64dc6

Observation 5c39fabb-892a-4e64-8eb7-441385c8b786 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.076312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:a44637557c980bc5ffb0f72173f26c32ee8c14db95404325c9f452b95135e47d

Observation 0d4ad231-3973-4dce-ad72-6f8e6e600c80 · outbound

This paper cites RobustBench: a standardized adversarial robustness benchmark.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks RobustBench: a standardized adversarial robustness benchmark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.095871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:9f2227bdcc3c876c0b7d1765c4db244e628c54c38703e74227ae2b37741ba54c

Observation 95dcfcdb-207a-4700-a145-00ca4333cc3c · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.705908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:d048568e159778707ea0d510e8897f2a76d95bb7eda71e14390247749ba16698

Observation 730fb3d8-ed13-4f0e-8647-67f0203e295b · outbound

This paper cites Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.699880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:6436a55c2b562aeda4896d49145a258af63e41fda6298e0c97cc1aadd7b3ec60

Observation c0671321-6117-43a8-ab20-b43f891fbbbb · outbound

This paper cites Find your inspiration.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Find your inspiration

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.724740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:e401013799be75d4a48c35e886ed39f64f3079b30e2f638f40c82ac505fcfed1

Observation 84230edf-4209-4308-9141-436bcd19977b · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.011869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:8a66ff98b623ba787f79aa8552017a7cb57fafa3680dc652dfcbece14961146e

Observation c79d67e7-4ba0-4d8c-8528-915765975396 · outbound

This paper cites Explore vision capabilities with the gemini api.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Explore vision capabilities with the gemini api

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.731642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:eedf6dc377692da9e7c0170afda6849c5dc4289d3fa1ecf24f7a69c156d13745

Observation ade1cdb5-0570-4e03-a98d-7f347cf903ee · outbound

This paper cites Gemini 2.0 flash.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Gemini 2.0 flash

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.740029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:b3b88d9c5f444dc09a04b8b01a9179bc76977eecc262e35f79724865fb13ebd2

Observation 6bbb666f-1d9f-4f5d-b1eb-55311d29b474 · outbound

This paper cites Benchmarking Neural Network Robustness to Common Corruptions and Perturbations.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Benchmarking Neural Network Robustness to Common Corruptions and Perturbations

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.007151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:dd53bc82872888e2f0578259b0a714aaa1e8256813e655b098acd650253554a5

Observation df572c0c-a2e5-4fd4-86d2-271579942874 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Measuring Massive Multitask Language Understanding

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.111288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:e86df467b15bf274158786d9f3bd086a365b15ca5b6a4de82e43d36c1dadbd46

Observation 3236a1d4-eee7-4b55-8888-743ea23c65d6 · outbound

This paper cites The many faces of robust- ness: A critical analysis of out-of-distribution generalization.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks The many faces of robust- ness: A critical analysis of out-of-distribution generalization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.742942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:e8a543b6920c168d4b510a2cf36051da8fe8474dbf1a079c61aeb23a74f93d68

Observation 7ead8605-e0da-4ea8-a98c-dd2ddaf30320 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.017947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:f526921d27f26d022b458438b43cfe78ccd15825f44550dd314711b0c6021d1d

Observation 02e8f7a4-7cd6-4b78-8421-92cceb6dec0c · outbound

This paper cites Llama multiple images.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Llama multiple images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.745902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:06cb48a095292c130c823d41a73903744160002f9e85695398202dd0f2785a7d

Observation 8948bcd2-cb98-4b5b-a743-8de1248f04a9 · outbound

This paper cites Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:07.979885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:7ff3b2a9d44c9b6243d681e77a3c32e4a8c4f1ac69d27465de491cf771d4175e

Observation 0050aef9-85f8-41d0-b6fb-84e79ea0f2d6 · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Oneformer: One transformer to rule universal image segmentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.758433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:e78fb58aefd60f0a11edb341e05b9dea9716be250f7837623728c6126172dfff

Observation 973035ca-77dd-43cd-8391-1e740949f05d · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:57:08.083275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:ce4c603032a380182a770b9bea607e7c2bdf3b915bb3550689fbb4abfd4043c3

Observation b8a5010f-9722-4133-bb60-b2db95d3b2ec · outbound

This paper cites Chen, and Andrew Y.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chen, and Andrew Y

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.702966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:4e26d23e4941c5674989aa3ada8b17b42e76150909c6da3d6027d6564ef50406

Observation 4108af43-0196-432a-8b3b-576dadfce6b7 · outbound

This paper cites 3d common corruptions and data augmentation.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks 3d common corruptions and data augmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.764254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:02e8a1efd4d7abe0a201b1d530b1a23c7d3c17c942945c5c85ae906d9477236a

Observation d3e3db48-7e1c-46c8-b4c3-c1a86f15a2ed · outbound

This paper cites 3d com- mon corruptions for object recognition.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks 3d com- mon corruptions for object recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.748514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:699c78c79898f8901d1c371de32f00ee61a197738befa6224b070b3afc22ea58

Observation 036c8900-e6b2-4266-9847-e8efc8e45b8b · outbound

This paper cites Decomposed Prompting: A Modular Approach for Solving Complex Tasks.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Decomposed Prompting: A Modular Approach for Solving Complex Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:14:05.332461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:a4790cf6010b0d10c471ab0b57975d04ddd4265ee5c692d81fb562a1c1e2327e

Observation 2c62bba2-d900-44ba-aeee-2bb0cc87b4bb · outbound

This paper cites Segment Anything.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Segment Anything

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.132363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:ae241f6e4fe644e1dd9366e40866776d6d4a3df6f1692d21e2c66d4fd51b8844

Observation 87106f7e-ac80-4372-9227-e5d629d9cbbf · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:07.985092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:62b8b5c1b95af7ab31f6ffbe036515d53c93a3b875e915689f2fbc5fd7f188dc

Observation d8bfc312-e3fa-462e-8b82-cc85b5c18280 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Evaluating Object Hallucination in Large Vision-Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.046330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:cc606933c6224ee4c997bd3becbe4d8a76da642a80d21e21dcbabe852c18ea29

Observation 31b85b5e-3e4a-4ca4-a6e6-8d011e5bd977 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.714845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:b3679737ca889779a831a14a8e3d290fe2942d040830e32536b81eb8cb3d5b4e

Observation b2633707-72e5-45a9-a7cd-99ca8e7ceba8 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Llava-next: Improved reason- ing, ocr, and world knowledge

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.691806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:610049ca17199d7d4ca2e78c24430154b95677fce38241602b7c2a010a9d67db

Observation b60ef790-8b20-4ad7-96ea-a956880e5b0b · outbound

This paper cites Llama 90b.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Llama 90b

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.694330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:309f521d69daefcb4cd79c81c39facb7a463a3c85cd23a2ac5cc40f17296e6c2

Observation 5f21c826-b6a9-46b9-adb4-1b22d7d3983c · outbound

This paper cites 4M: Massively multimodal masked modeling.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks 4M: Massively multimodal masked modeling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.720188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:abb3a5470696b90e208ae44e409489e6ec520af2d9227c17cc53518154997f0c

Observation caa1614e-ebf5-4ff0-9589-84e51b30f0ec · outbound

This paper cites Introducing openai o1.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Introducing openai o1

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.689259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:e014602d1192741c8d118713eaab2a72619cf91f7c5c16d659343410f90a6247

Observation 8d810ea7-01d5-48ed-bba2-64f380570a8e · outbound

This paper cites Hello gpt-4o.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Hello gpt-4o

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.755078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:41a28fd808587ada40f3a0b41537a0eca9d9d282e5759915cb489d4e38664192

Observation bde9b641-06a3-4a78-9d7e-5f65a9716d32 · outbound

This paper cites Introducing 4o image generation.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Introducing 4o image generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.685707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:1980049a432f57a369835848204534a54fe6fb746cc5627297694ea1859e8be5

Observation e7521322-5fc0-41de-8eb0-ad9449d1db3e · outbound

This paper cites Introducing openai o3 and o4-mini.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Introducing openai o3 and o4-mini

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.682787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:3f0161864d825448b5faeb8ea3bf16c93c1917cad18f8c38a2820c736d8c3ad8

Observation 8bd60bd1-6dd6-41cf-8065-10bf0bc2c12c · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Vision language models are blind: Failing to translate detailed visual features into words

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:07.991005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:d3665fd5eafa26fea9d4a3735fb2149929731d4b9c2ce302d1286304245dafa1

Observation f276678e-f04e-4304-96e6-e1db2b28a8c7 · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.709224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:9f7a758a1a8033cd54e82296b5521feacf05569339a099336f98d4409e99da0b

Observation 8be94be3-92eb-45b7-be9d-0d662b598277 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.034625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:806a8cddce10799eba8f049b7586c83ab28e280407457454b6c1e670cb8a6b80

Observation bfa972c0-e570-4f0d-86ef-a64c9340edf0 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.028629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:6db2d8891ee81ac9c7331933a0dcde54f2f60bc3b6a8fa026904a319efb9f6f8

Observation 3baaf630-7cc1-4ff2-8d78-9f99823937a8 · outbound

This paper cites Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.721474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:49bd402dee815538b6d07a0e92f25853d47dfd0bbee88d6be66f569812f800a0

Observation f59fb832-03c7-465f-9852-983109d65a38 · outbound

This paper cites Bernstein, Alexander C.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Bernstein, Alexander C

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.717995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:ec6af9a27ca66833934b708441f929d803d606678c1501282935a7a640b5a9ea

Observation 2a0a9a7d-cfb2-4517-9bf2-0bb3380aec2b · outbound

This paper cites Super- pixels: An evaluation of the state-of-the-art.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Super- pixels: An evaluation of the state-of-the-art

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.761451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:27631abb7cbd202122781921fbcfc25d76314efb4c289d9a93f6aa95456bb814

Observation 2b46d22b-83c9-4b4a-a96e-162b889799e0 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.040162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:ccbf07fbb6a555549c7d36692da7b52573912f4daf45af535fbb7e2d00ae5660

Observation a4b59d9c-d342-4c57-a817-5f8602939e42 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Gemini: A Family of Highly Capable Multimodal Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.059358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:8b1c92fa9a39f0f1d02f8a2f86204cb8dc9c523f5c9e1e8c164e3e44975ca509

Observation 1f2549cc-fb7c-44bc-95a2-27ecc561455e · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:07.996972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:52ef14087781630a4807d234e5146b01955375ea806b72b7c62dae5c12de1084

Observation ffedd751-f1f1-40a0-92bd-419d0354ae2e · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.023724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:2a3d30237d2ba0d587acc0b83924d575cd79b829bf97554a6aa1fd0ba6ca4dce

Observation 5cfcc742-0105-4fbb-bced-2d8361878310 · outbound

This paper cites The internet’s source for visuals.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks The internet’s source for visuals

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.677897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:f0c463b4b156c4480ffc709cd62f3fc6239409c76e26ed310b4121c42b7bd0fc

Observation 056c4c45-ed18-4d98-a27d-5da1f431d869 · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Learning robust global representations by penalizing local predictive power

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.751780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:bfa8819023524ec5a489a2cea4fe750e9ef42243863e2af3e71bd4b6b4ddf0d0

Observation 5f5817ac-39fd-43d8-8f80-93b363ef8914 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.002146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:e64a2b6091ab76fb061e01b072a9f77f5616399f553832017caafc15968dd5d6

Observation 213f451e-94d2-463f-99cf-2529fedae323 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.053030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:fc409369d00e66eb2e20e177a62d9e324d41499e5272939397fed8bcbf88b3ba

Observation 2ac40667-6aaa-441a-815a-faca4caf20a6 · outbound

This paper cites Chain-of- thought prompting elicits reasoning in large language models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Chain-of- thought prompting elicits reasoning in large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.675172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:613ad00e0d17acbed9599aa76207ed14f48756717ae45418455961a7bb65c669

Observation ae4d69db-d967-4b4b-95a1-1a9f2762b48b · outbound

This paper cites Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.686589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:e3e87846aa07a4b9f779b581a16dd9799c3017a49ba645e949b81666a3e38640

Observation 69876831-09f2-44a3-b370-55ab9781abd9 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks V*: Guided visual search as a core mechanism in multimodal llms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.672495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:3d8b8fd24a70303e6280dd19639b62eb0f937c5ff4674bb03a6592d2d38d59c3

Observation b5e57172-4c45-4398-9a9c-3339072f57b7 · outbound

This paper cites DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.122652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:622995651ce04002273a64e1f4177d72f2ea10e209f61789e7685a72cd1722e6

Observation a23c936f-d569-4cab-8394-1ceed439aaf2 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.127733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:94e4740728578cca9c05eddb9e502b087c327dfeeeae98761aecee1124d891ee

Observation 47e5d366-d3d8-48bd-94a6-4a69f495f551 · outbound

This paper cites The dawn of lmms: Preliminary explorations with gpt-4v(ision).

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks The dawn of lmms: Preliminary explorations with gpt-4v(ision)

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.669221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:cc79331759d7237fb4ca818768460297a7a3f02192d60712f9771742ceb2e8ad

Observation 0db5fbe9-c484-4b21-b977-62af7bc7cfb9 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Tree of thoughts: Deliberate problem solving with large language models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.680646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:d6a2a7b4ed584e0fa947991d952b2eaa213559c76a4a78c4b9313f3c02e56452

Observation 376d96c5-4802-4c08-8b5a-3b4ffcd3b0af · outbound

This paper cites A Survey on Multimodal Large Language Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks A Survey on Multimodal Large Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.106326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:6ab23450ff77ba8c1d2278af43848d12ffd13669e7272675998cc2511957231d

Observation af8d3448-6315-4891-abf2-c2aab21c68cf · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.665794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:18d7956c04f5aae7eb709710bbda0d33c7af593cb6af03138f6f7daae4358604

Observation 822a7497-64f7-4414-9d5b-6cf82546187f · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.117043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:ff4bd75801ba7684c2d8658d344efc0e13904b3d6667bf0da876df9d26d93024

Observation 08aebd87-c1cb-4bf5-b47d-a4a3ffec27dc · outbound

This paper cites Scene parsing through ADE20K dataset.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Scene parsing through ADE20K dataset

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.708902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:332ee22cebdfd9a9b82ae57937e3160b98ab8f9cf67c3f012bc42064f1acf590

Observation 12eea387-8707-4645-84e3-3078f699f366 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.101076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:23ce8e05fc077562a482eb5b0ad783a05662579db2ded527a1f32155d5f8181c

Observation 96d4ef64-4828-49f6-9958-e2f55a64254f · outbound

This paper cites Detrs with collab- orative hybrid assignments training.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Detrs with collab- orative hybrid assignments training

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.674916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:81e101d3370440bb78ceca08d3288f4ae40772e53a615ca4131556708d3b7b95

Observation a25b1049-d68a-4101-a9fa-985b3c5b141b · outbound

This paper cites Learning ordinal relationships for mid-level vi- sion.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Learning ordinal relationships for mid-level vi- sion

Reference 76

Resolution
malformed identifier
raw_fallback, observed 2026-05-19T06:13:00.683693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:1c6ef09c73cf526726354d7decc38cce7e03fa775c3f04235f973867017a3cc9

Observation b0615429-3242-496d-87d8-516a40145772 · outbound

This paper cites objectness.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks objectness

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.653908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:72968cca95b03879fc6c4ae940cda7ab7b1aedb06a958906050c5ad15ce8a38e

Observation 4da8a1cd-71eb-4304-a4b9-4ed8cf720607 · outbound

This paper cites - Produce a raw image (same dimensions as input) whose pixel colors encode the normals as above.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks - Produce a raw image (same dimensions as input) whose pixel colors encode the normals as above

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T06:13:00.647136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:91d8e947d6a4278112dcb1547a3cdcfc9fa256d3b9f36a8ce2a2ef73ccbbd18c

Observation 3310b469-7a3f-417f-a91d-d20ba3c26157 · outbound

This paper cites Listing 3.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Listing 3

Reference 79

Resolution
malformed identifier
raw_fallback, observed 2026-05-19T06:13:00.694038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:2e0d8d1a87c10017b31c9a3d21fdef5ba8d63329564e30304a4cfcc5b33c903d

Pith citing papers

Observation f10f806c-317f-4939-8b37-230e2c825177 · inbound

Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning cites this paper.

Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-09T22:39:15.088967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T22:24:44.884495Z digest=sha256:01deb9d4ebe74f5803536a9c92ccf9e1e530337553a05a6a9194c8264710985e

Observation 15d0897b-ccfe-4f79-a439-0d17abd1a35f · inbound

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education cites this paper.

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:12:46.508029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T23:12:32.101737Z digest=sha256:ff917a95a5f87939d5734ba809a53f5e18f4a76c904240cbe402d62158841283