Pith. sign in

Paper Citation Record · LEDGER

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning

As of 22 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2507.07297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07297 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:51:06.202263Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact4
  • verified fuzzy12
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01079c24-8587-41b4-9ee6-5a8b73321352 · outbound

This paper cites Llama 3 model card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.420164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.420164Z digest=sha256:a4a2a32821356a1f84760df04b0c168e028e589bd73d8d945392f6081a2383f8

Observation 4adc7ce3-70be-4142-8439-c3db538f8671 · outbound

This paper cites Introducing the next generation of claude.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Introducing the next generation of claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.910433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.478021Z digest=sha256:aa303636bdea4ca5dda2385a223848a66c6bbb77a501b71c9b1bffa9e58c093a

Observation a287fb6c-90de-4f06-9964-f737307741be · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.544169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.544169Z digest=sha256:cd3fbe12e52c08b3674fb8e23ed7a38e7ab534e15ccc5e4e1a25a2b17fdb918d

Observation e5658769-623e-418d-8a07-f020d37f21a0 · outbound

This paper cites Qwen2.5-VL Technical Report.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.629099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.629099Z digest=sha256:45c4d3d8dd3fdd0fa2110834aa3d2f880b7adcbbbe8bf8cf489e5dd671d7ebee

Observation 847a7c53-0ffc-4be9-b655-69daa69edf73 · outbound

This paper cites AiR : Attention with reasoning capability.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning AiR : Attention with reasoning capability

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.713385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.686648Z digest=sha256:0a44ae80414d8f5623aef4e0038351ca90b8b18063bdd33472c33d84c3ea886b

Observation 26524540-6843-4f25-8d64-d6d90b3161f9 · outbound

This paper cites Measuring and improving chain-of-thought reasoning in vision-language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Measuring and improving chain-of-thought reasoning in vision-language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.766806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.766806Z digest=sha256:820d0a3c5cc1722fc9f3d0223331764d6033538d7b98f4129a88ab087d8f45f3

Observation f50ac4c8-1124-4a79-b6c8-e32f1c9947e9 · outbound

This paper cites See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:03.841863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:03.841863Z digest=sha256:6734ce7eac25d64894a6bae5f78d131c01f9876cd6679c899e8aa4c2c5adce4c

Observation f08cb58b-885c-473a-b805-4942db340f1d · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Spatialrgpt: Grounded spatial reasoning in vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.565046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.907226Z digest=sha256:f20afd27956006f3315c13443a2340a5bbc488bcca7aabb70edc86f8e8ac7806

Observation 4de4e951-11d1-4b5b-8837-207ab0b1cd12 · outbound

This paper cites Aya-vision model card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Aya-vision model card

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.411182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:03.976013Z digest=sha256:395742443de9c847e7a9f894531a957faf3c9f3bae41535a331a520614e14688

Observation 9428915f-571b-493d-8bd4-4e9c9f3684c4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.037683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.037683Z digest=sha256:f1c257cb110af794a7c816c6fee72b4580a7585337ddca9cf4951c6b46215110

Observation 803f8c99-c8b1-4ea9-aa9e-dd8cd31a484e · outbound

This paper cites Improved Visual Grounding through Self-Consistent Explanations.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improved Visual Grounding through Self-Consistent Explanations

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:51:07.033159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.109878Z digest=sha256:6e426d0c78945d8de66b64dcf5c29142fbb515b8e4ba1562ed6f3481435b6bae

Observation a94447db-402c-4f07-9daf-f7be8ed79e9b · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.229609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.201246Z digest=sha256:9de97a2b8b72eaac63b7463d5866853458217584bd3a4aaf4be5264cbbf5ccd5

Observation 4bfa830a-d607-4b6d-a050-a712d764dd7a · outbound

This paper cites GPT-4o System Card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.289982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.289982Z digest=sha256:0d1269207c976556f41ba9134fc647b3d2ebfce3124582d87f8a4cdfdd28ddc6

Observation c891e4df-803f-47e8-89ec-cdb9b2b70354 · outbound

This paper cites Weakly supervised grounding for vqa in vision-language transformers.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Weakly supervised grounding for vqa in vision-language transformers

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T18:51:06.522495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.365691Z digest=sha256:85fa0f434a8e582b7ec7fba668af3f88374393fe46ef49ebf0e49410ae4480e4

Observation 89a1cba5-0ab5-4c6a-8fa0-fe426fcdecb7 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.420426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.420426Z digest=sha256:ed2633b88a20457cb71ac035ee4bf29f9349d1a7a8d40fc8f23ffa52094f35fa

Observation 2cc45b3f-a014-47a5-9625-f7b95e3a6b91 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.468078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.468078Z digest=sha256:c8a4a6d539ad57a029f4be2fe15b2639bde4288cf4c5c5b0d19759e4f4731051

Observation bc1c45f8-a047-46a9-a055-0e97c45ce61a · outbound

This paper cites Visual instruction tuning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:08.068541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.531795Z digest=sha256:62ff49d7dd59d465557cfd1f6edb80c991cbc7a5c5fa04a99335bd9d68241c86

Observation 6de4a6e7-42da-4ea2-888b-26422013531a · outbound

This paper cites an unresolved cited work.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.586706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.586706Z digest=sha256:92357fc7d6ed99339fb921e35e0a530d6ef8d52eb66a3b430c477e6060553b8f

Observation 7d3d5810-42b4-4185-abcb-0427d493b7b5 · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.654427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.654427Z digest=sha256:0e98e475aeb6e8b3fe05cca3c627e4b4eca360fae223eb0631751464c5635938

Observation 3326c16b-02c8-4888-a0f1-2431d76235b4 · outbound

This paper cites Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.723812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.723812Z digest=sha256:7f82bd39bceef12a9e7a459d9a96b65da44112608b9637243bfa479e5a1df776

Observation de87c41c-c010-4906-bd35-c2f9d1377340 · outbound

This paper cites Gpt-4v(ision) system card.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gpt-4v(ision) system card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:04.807420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:04.807420Z digest=sha256:4301ce42e6632350605751f6148571abc23dc2ae2a4cc5ec7528f7d6faf51b9e

Observation ced51df7-21a3-40f8-bc0b-9a2d7e1cd7f0 · outbound

This paper cites Introducing gpt-4.1 in the api.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Introducing gpt-4.1 in the api

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.933231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.872081Z digest=sha256:02f86b64e67ca90d0486b73c669823b4f032125c73ebd5e6fdec91c87633c925

Observation 73f51922-0f6a-4e14-b238-55e7ff329eb9 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Qvq: To see the world with wisdom, December 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.766881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.903213Z digest=sha256:78f0f48b3bcc4122fec8ff0fed250f228eefa1a7dbf0d6c8d76a193d42380f83

Observation 52598951-ae8e-4590-9c5c-23e5980ed9d9 · outbound

This paper cites Uncovering the full potential of visual grounding methods in VQA.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Uncovering the full potential of visual grounding methods in VQA

Reference 24

Resolution
verified exact
doi, observed 2026-08-06T18:51:06.383863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:04.962971Z digest=sha256:dd6e33d45e22e2f2d1bc4cad8d87cff6bc79b06c3aa32a4ae56be38bc700c1f6

Observation e1e41564-69ac-402c-b641-dd779435f174 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.057773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.057773Z digest=sha256:fee3048fcf8a71aa2ac739d8fc2351af16651b18dfb816b83dca72b7cac18640

Observation 7264deb3-fdc8-413f-8112-00ef15330c3f · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.113854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.113854Z digest=sha256:a43bb761b4e16f97f67c4c2c95dd26e228ee533d1839e1727ef8f0333a5130a2

Observation 3e314ea3-e67d-44e2-9326-95c0ef7349ea · outbound

This paper cites Gemma 3 Technical Report.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Gemma 3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.213221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.213221Z digest=sha256:724c86d5b9b8e0278766cdb8238689d8881ceb2b0bf670eb31dd9b56fc948e4b

Observation 391d53de-c277-45ab-848a-cf4e971f8c19 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.257142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.257142Z digest=sha256:58ed8fdacbb7fd936f25b27719058dfb3b33979c1b85d0d3f47daaec6dd35592

Observation 51ec8827-b6fb-4c75-a226-2429fe40ba5e · outbound

This paper cites Contrastive region guidance: Improving grounding in vision-language models without training.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Contrastive region guidance: Improving grounding in vision-language models without training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.326894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.326894Z digest=sha256:44d5f1bd6b0e9ba4ef5de916d82c19b56c634787d7e4d37c8636642434681487

Observation 2fa00a24-98fb-41ab-a6ca-c0a1990eabe7 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.402887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.402887Z digest=sha256:59d0573b2f724c084bcc1109d2d242adaf6bb599869791b3df86f5a8675503dc

Observation 2430184c-629d-4eec-8bc3-f9736569496e · outbound

This paper cites VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:51:06.738821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.455071Z digest=sha256:d11f43afb43f0ef3e24c771a27ff2517bb8a3ab675705198089bbb070be8b707

Observation 62e285fa-da03-4007-b4f9-e48be3f54dcf · outbound

This paper cites Llava-onevision-chat: Improving chat with preference learning, September 2024.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Llava-onevision-chat: Improving chat with preference learning, September 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.629867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.525846Z digest=sha256:627c74f13b2355d708f5e344871bd99e98fa47a3e0f18479b104b939d04a6f58

Observation b0d75234-27f1-4d5f-b1e7-762272879e96 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.654601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.654601Z digest=sha256:3fca4d9736683a449a633ce608311e70298da68a1eda044c1b646851d386ead5

Observation 08cec980-e5df-487c-a7a2-2e261c508c7b · outbound

This paper cites Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:05.755348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:05.755348Z digest=sha256:113e1858e1299a1058852b4991e11d54d63cea661dd6ea8afadbe472abe1a902

Observation a99c0d05-b26f-4792-a42a-7dc8685f882a · outbound

This paper cites Mm-vet: evaluating large multimodal models for integrated capabilities.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Mm-vet: evaluating large multimodal models for integrated capabilities

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.478808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.824802Z digest=sha256:713760a830f278f79f1bcfd675c1d89da2bc4198c30d7c88a332107d2a6c8426

Observation 3d4c7102-8b8d-4b1d-b41b-10db977a64f9 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.321052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:05.946191Z digest=sha256:3e84f7b60e76f39fdff9e09add741dbc5cb8822c8f137236ea8ecb60bba66d69

Observation ca938052-b05d-4806-b200-a1166d0a1387 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Improve Vision Language Model Chain-of-thought Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:06.064596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:06.064596Z digest=sha256:90f24634bd9df419e064d878fd522fabcd4de6ea92497afdd0e1d6e176b9790a

Observation 5cc0ad3c-9096-45f9-bb46-c620dd382045 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning Multimodal chain-of-thought reasoning in language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:51:07.198965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T18:51:06.133863Z digest=sha256:cfc5b429ec0e44faf47b57d58c31ed22a91d42b7e895bece9b27d4f05717dacc

Observation 6a330a3f-c247-42d8-9270-a560a91fd405 · outbound

This paper cites CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models.

MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:06.202263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:51:06.202263Z digest=sha256:47c47e1182926f19438dd756c6552c03e5cf9b55e168c86e6e5b9575f6e4f7c8

Pith citing papers

No inbound Pith citation observations are available.