Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:33.769358Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 2 inbound Pith citation observations for arXiv:2505.20753.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:33.769358Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T04:17:40.198357Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:55.004072Z
100 of 103 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f0c5ce36-7756-4d7f-99c4-076090780625 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Tallyqa: Answering complex counting questions
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 378366f4-92e8-4dfe-8de8-9cc157e94372 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Macmillan, 2005
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf90947-2c0e-48fd-b9a6-444821b05537 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ab5a2e-056d-4f9d-ab3c-b2824f777881 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Graph of thoughts: Solving elaborate problems with large language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7367a42-a8d2-4aba-b715-d790901e698b · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vizwiz: nearly real-time answers to visual questions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d52929-5863-4763-8f58-0d862d6a24ce · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Due: End-to-end document understanding benchmark
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a14884bd-d4e2-49a5-babe-6df5e63e82e7 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f9a441d-3597-404b-865a-d2b17460588e · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bda90484-709f-472d-b9fa-185e123158c1 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65924100-3232-4961-af83-da8e2ddea302 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db8759c-387c-4251-be4e-56fc6039b0cb · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gonzalez, Ion Stoica, and Eric P
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52b4758f-07a6-4826-a16b-2257cfbfeaff · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a573766-a33e-4c3b-a218-901eee60fb5d · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417f4e4c-8ac4-44a4-ab4c-d8069a1cf6ed · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 955f3862-cd6d-425b-9487-cacf88b4967d · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4562eb-e1eb-4298-ab9d-ce26fb8d5219 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Flacuna: Unleashing the problem solving power of vicuna using flan fine-tuning, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580d697d-f7a0-417f-9820-53c92980fb3e · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d33b65c-aaa8-43f7-ac12-8c5fefa27395 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a0534e-ba85-4e6b-94bb-03596443725c · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual programming: Compositional visual reasoning without training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64bfc413-9da2-4864-9bc9-332813dd0d7f · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models GREC: Generalized Referring Expression Comprehension
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608935c6-b7d3-4b0e-9d73-7fb1f6999ddf · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models spaCy: Industrial-strength Natural Language Processing in Python
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adcf352b-ff59-4a77-a590-39a1900d31a6 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual program distillation: Distilling tools and programmatic reasoning into vision-language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca6894f-45a7-4e25-b8b5-4e6c8779bbe1 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Hudson and Christopher D
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40686fc-98d3-4eb9-b2aa-989e442e7940 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vcoder: Versatile vision encoders for multimodal large lan- guage models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a3bc90a-502a-4109-a130-f4f4218e71e7 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc28b69b-2bdd-4f01-92cb-1e75ae9a1463 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Dvqa: Understanding data visualizations via question answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8d1255-7b28-40ec-9aeb-c7560ea16bc5 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Mdetr-modulated detection for end-to-end multi-modal understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c46aa50-7e9e-497d-8cdd-58f9336c2bdf · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models A diagram is worth a dozen images
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88d2eda-89bf-441f-a912-3f178702cb78 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ocr-free document understanding transformer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde02abc-151c-4d07-8be9-3d5787e222d9 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8733d8a4-d5f7-4e6d-8edb-c83824827eb1 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Shamma, Michael S
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d264e598-dba1-4a0c-8182-39169e1c6db4 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435ca4ef-4f62-4aeb-99fb-5769e0e1ef27 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaa4e6ba-fc9f-48f5-9b0b-facb283dbcf0 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a55e395e-5397-4627-8430-f5d8ea9f42d7 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Evaluating Object Hallucination in Large Vision-Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f1502bd-9178-44d1-928a-fb4d474499b0 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8473446e-f3dc-4a6b-9da8-71901626e012 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Monkey: Image resolution and text label are important things for large multi-modal models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c11fef41-281f-476d-8165-e17a14aa3f49 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Openorca: An open dataset of gpt augmented flan reasoning traces
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2e8eb45-df10-44a8-8d88-9f589bc5a246 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Microsoft coco: Common objects in context
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32adae86-aba9-41c4-98a9-81111160a3a0 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82437ce4-a71e-489f-8774-159e3913f8a1 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9833d0-f52b-4065-8c01-1211b45fa3a0 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Improved baselines with visual instruction tuning, 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63cea63f-afb5-4efd-bb9f-2b4d68737187 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Llava-next: Improved reasoning, ocr, and world knowledge, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b47bb85-f4b1-4233-ab38-4ddb71c2176b · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05240b21-1ac2-4421-a709-8474e902b52d · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea90225d-36a0-4628-b4d1-b53b74070a7d · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MMBench: Is Your Multi-modal Model an All-around Player?
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2adaeff-cf3f-4379-9212-7a1a5c377f61 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c56508f-e269-44b5-96d9-07a2bec943ec · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Decoupled Weight Decay Regularization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5289b8-a2f0-4180-8ac6-af8bf1b6a8b7 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a68b5598-0cb6-4120-bbc7-4155901aeac1 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93c279a6-7aaf-40bb-99c5-b7769f486fd0 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models WizardCoder: Empowering Code Large Language Models with Evol-Instruct
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb7fc0a8-3ee4-45b7-ba58-37f468155d54 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f18f105c-723e-4f69-bc5d-835bcd223f7d · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99fdb9fa-25a6-4dd7-960c-7f37b6558ffc · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models DocVQA: A Dataset for VQA on Document Images
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7fb0e1e-04da-47cc-bf80-2bb65dc0b676 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Infographicvqa
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a71263-a21a-4380-abef-d0154ff3ec05 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Schema theory revisited.Review of educational research, 75(4):531–566, 2005
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8731295-39b9-4461-a89c-a346ff2d16ed · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ocr-vqa: Visual question answering by reading text in images
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50c017f3-90c0-4df5-a43d-a94f52ab6a77 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Compositional chain of thought prompting for large multimodal models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4c16501-60db-4847-a6c4-4a5b7eb6b375 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Modeling context between objects for referring expression understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04158600-31a1-4870-a91c-ace12a753643 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gpt-4 technical report, 2023
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 644297db-8b94-4964-ad8d-d9d0c68fa519 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Chatgpt: A large language model for natural language processing, 2024
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8742e0ac-8199-418e-b456-4aa0f5a1a862 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learning to predict visual attributes in the wild
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e55e88bc-d4d8-4cd2-9c50-ea5b04d010a0 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Plummer, Liwei Wang, Chris M
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43f3024b-f2bc-4c52-91c7-2f7eddfb5c10 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Jack of all tasks master of many: Designing general-purpose coarse-to-fine vision-language model
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efde47e4-f32c-43f0-8ffa-cd59057f1d59 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444391bf-6e3d-43a6-988c-750be1226fc7 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Learning transferable visual models from natural language supervision
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc710403-a39d-40f4-a2ec-e273c051da41 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d14563a-fa7f-4563-b0d4-e07d0bba71c7 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Faster r-cnn: Towards real-time object detection with region proposal networks.Advances in neural information processing systems, 28, 2015
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40e0e4a3-f065-4fdc-94c8-bfc9dde9bf10 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models A- okvqa: A benchmark for visual question answering using world knowledge
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f446b261-ed47-44a9-b1e3-e3c89900bad2 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79b54ec-36a7-41a4-9f59-194a176e5c8c · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Objects365: A large-scale, high-quality dataset for object detection
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96e4623-5ba0-4a39-a2e4-78bb8261499a · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Woodpecker: Hallucination Correction for Multimodal Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b36a20-3f15-494b-9b7e-85c21bb343e3 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Textcaps: a dataset for image captioning with reading comprehension
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c83ac2-07e3-4182-ac75-6bc7c05496cb · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Towards vqa models that can read
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63324793-9ecd-451d-83e5-20c423547bf7 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Vipergpt: Visual inference via python execution for reasoning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a94df1-76aa-44e5-b482-72557d258da4 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gemini: A Family of Highly Capable Multimodal Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4576e48-1534-43e1-841f-2936f52b26de · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Gemma 2: Improving Open Language Models at a Practical Size
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8556a4e-68e7-4b55-b6a2-e46ddf9619c1 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen2.5: A party of foundation models, 2024
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00b24c17-26cc-4e54-9912-f809d9449f77 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7227ddf-c6ae-4f25-a30b-dc91397c878b · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models V3det: Vast vocabulary visual detection dataset
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ea5dd24-af13-40fd-b158-c0121e0842bd · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6664e084-8326-4525-88ac-8aa5900c3b19 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fb6267-65f7-4ac8-bcbe-8a35e9da365b · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models V?: Guided visual search as a core mechanism in multimodal llms
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d076da-b523-4ae9-95ed-dfada6d68c35 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Universal instance perception as object discovery and retrieval
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ceb3ff64-846a-4a98-bd96-1e607ffe50f1 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bce61a9-aea9-4e95-9caf-c3fdd31c82ca · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Tree of thoughts: Deliberate problem solving with large language models.Advances in Neural Information Processing Systems, 36, 2024
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47d9c578-d58f-4748-af17-7ea6ee6efd43 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ferret: Refer and ground anything anywhere at any granularity
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81556e13-6f8a-4852-8d99-066d8bf27378 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c74538d7-e8a4-4a2b-b73e-4170ddac37f4 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Modeling context in referring expressions
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ee7ea1-47a9-43a7-870b-bd16d1a7f6d9 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad44938-ec86-46b1-9972-491612c46e4b · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Osprey: Pixel understanding with visual instruction tuning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db65240-c2c3-444a-a8f7-6bbfac99a870 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d458aa-0cd3-4836-96de-c49499754562 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Griffon: Spelling out all object locations at any granularity with large language models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 250fc1e9-7210-495a-9cb2-b8a494f739d5 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Automatic Chain of Thought Prompting in Large Language Models
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43994f56-3434-4817-80de-4d69d73512ac · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models yes” or “no
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 420e162c-67b3-4f6f-9d11-b5510053330e · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f776522e-df36-4c05-9dbb-8041809823c6 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fa7385a-6ddc-471a-ab2d-f46df1dbc8bb · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af56e164-1237-4768-94ab-da81c32ec363 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d7846c4-193e-4223-9e09-2d8cbd0b1552 · outbound
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models Unresolved cited work
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d8824a4-823a-4543-b995-e55d68069cfd · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
Reference 179
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b38d76a-528f-4b16-b509-d2292a23dc31 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
Reference 267
Source-reported events for the cited work
Unavailable: canonical work link unavailable.