Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:06.202263Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.07297.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:06.202263Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 01079c24-8587-41b4-9ee6-5a8b73321352 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Llama 3 model card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4adc7ce3-70be-4142-8439-c3db538f8671 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Introducing the next generation of claude
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a287fb6c-90de-4f06-9964-f737307741be · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5658769-623e-418d-8a07-f020d37f21a0 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 847a7c53-0ffc-4be9-b655-69daa69edf73 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning AiR : Attention with reasoning capability
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26524540-6843-4f25-8d64-d6d90b3161f9 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Measuring and improving chain-of-thought reasoning in vision-language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50ac4c8-1124-4a79-b6c8-e32f1c9947e9 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f08cb58b-885c-473a-b805-4942db340f1d · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Spatialrgpt: Grounded spatial reasoning in vision-language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4de4e951-11d1-4b5b-8837-207ab0b1cd12 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Aya-vision model card
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9428915f-571b-493d-8bd4-4e9c9f3684c4 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803f8c99-c8b1-4ea9-aa9e-dd8cd31a484e · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improved Visual Grounding through Self-Consistent Explanations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a94447db-402c-4f07-9daf-f7be8ed79e9b · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bfa830a-d607-4b6d-a050-a712d764dd7a · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning GPT-4o System Card
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c891e4df-803f-47e8-89ec-cdb9b2b70354 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Weakly supervised grounding for vqa in vision-language transformers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89a1cba5-0ab5-4c6a-8fa0-fe426fcdecb7 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc45b3f-a014-47a5-9625-f7b95e3a6b91 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1c45f8-a047-46a9-a055-0e97c45ce61a · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual instruction tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6de4a6e7-42da-4ea2-888b-26422013531a · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d3d5810-42b4-4185-abcb-0427d493b7b5 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Deepseek-vl: Towards real-world vision-language understanding, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3326c16b-02c8-4888-a0f1-2431d76235b4 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de87c41c-c010-4906-bd35-c2f9d1377340 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gpt-4v(ision) system card
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced51df7-21a3-40f8-bc0b-9a2d7e1cd7f0 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Introducing gpt-4.1 in the api
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73f51922-0f6a-4e14-b238-55e7ff329eb9 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qvq: To see the world with wisdom, December 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52598951-ae8e-4590-9c5c-23e5980ed9d9 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Uncovering the full potential of visual grounding methods in VQA
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1e41564-69ac-402c-b641-dd779435f174 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7264deb3-fdc8-413f-8112-00ef15330c3f · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e314ea3-e67d-44e2-9326-95c0ef7349ea · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gemma 3 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 391d53de-c277-45ab-848a-cf4e971f8c19 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ec8827-b6fb-4c75-a226-2429fe40ba5e · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Contrastive region guidance: Improving grounding in vision-language models without training
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa00a24-98fb-41ab-a6ca-c0a1990eabe7 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2430184c-629d-4eec-8bc3-f9736569496e · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62e285fa-da03-4007-b4f9-e48be3f54dcf · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Llava-onevision-chat: Improving chat with preference learning, September 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0d75234-27f1-4d5f-b1e7-762272879e96 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08cec980-e5df-487c-a7a2-2e261c508c7b · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a99c0d05-b26f-4792-a42a-7dc8685f882a · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Mm-vet: evaluating large multimodal models for integrated capabilities
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d4c7102-8b8d-4b1d-b41b-10db977a64f9 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca938052-b05d-4806-b200-a1166d0a1387 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improve Vision Language Model Chain-of-thought Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc0ad3c-9096-45f9-bb46-c620dd382045 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Multimodal chain-of-thought reasoning in language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6a330a3f-c247-42d8-9270-a560a91fd405 · outbound
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.