Pith. sign in

Paper Citation Record · LEDGER

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

As of 7 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2506.23563.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23563 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:41.103071Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:45.693853Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:04:22.765546Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be9ea60a-6bb1-447f-9f50-bc46dd5b167e · outbound

This paper cites Vqa: Visual question answering.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vqa: Visual question answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:36.964984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:36.964984Z digest=sha256:da26f587fc098e02fdf657766d949badce6918261fb0a12b36170bfb4135e29e

Observation de0012f4-cfd3-446e-bb25-da7f0447338f · outbound

This paper cites Qwen2.5-VL Technical Report.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.018974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.018974Z digest=sha256:e75add6e80387fc10dfdcb324504a13aea1e34daaa5396c2513e5b6361f58a93

Observation f03e35c7-15c6-4412-a186-6ba364de3669 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.060542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.060542Z digest=sha256:8720c206300e128ffe708709c0e09e63d304ec09734b74cda8b65aeb521f173d

Observation 36ffbdb7-f622-4f64-8ec5-c54f06370650 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision- language models with less than $3.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-v: Reinforcing super generalization ability in vision- language models with less than $3

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.107652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.107652Z digest=sha256:7f210da7fa804cd62adbbf29effceea56c400db898a871beb5c666cadd7b1d51

Observation 7730835c-196e-46e4-8e49-79513dffde51 · outbound

This paper cites M 3cot: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI M 3cot: A novel benchmark for multi- domain multi-step multi-modal chain-of-thought

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:44.435523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:37.181858Z digest=sha256:1bb8ecb356b544e63eed7c3615f4d0d8f53aa516ffd1bd881bcb5f4c1e09dd62

Observation 2c7e9642-a714-4ae7-96d4-65bf223d4959 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.266368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.266368Z digest=sha256:d62b9c3ecb2f6d926d6920fdde8c7243caa4c5d330ffe549bdf1865bb2cd1351

Observation a59ed741-d1c7-4846-9d49-55986827a34e · outbound

This paper cites The Llama 3 Herd of Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.363908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.363908Z digest=sha256:1030d8ff76b5857751fc0376f51abc593638960eb5d00ea2fb690e8279f9a6c1

Observation 9ebdc35f-8e27-473e-997c-4698bd2eadc0 · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:44.182233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:37.416896Z digest=sha256:9f1767d5e99b78a55f862c5dd4a2b3614632758b19573accf571cf3bda07770c

Observation 93ab7eea-85ef-4eb5-9497-ead8a6c49b62 · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.956530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:37.474185Z digest=sha256:9a545032275c4e8cc530306b92f5c83db10616cffcf8044c6ce8f379edd5b083

Observation a9715380-7d87-4d83-986f-97ae974160d1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.543852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.543852Z digest=sha256:b75c85bbc066a11dabcd850e4af04433f94272fa82d409a1e3672ab1cf983c84

Observation 77e758fd-1c0b-4b84-ad84-7a408c596d97 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.636240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.636240Z digest=sha256:36ab05d83d2830d15ea400b18a78485feb229fee1cb9ed2bace6bd2ee7d9293e

Observation fb28a5cd-421c-445f-9007-ff80bac185c9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.734988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.734988Z digest=sha256:0db0ad43ef1ed1b4fbede70cfeb6bdb1f219363aa76b6f115253c28861906746

Observation 5b0541e0-2129-48eb-9b5d-c0aaf865050a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.820286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.820286Z digest=sha256:6b318295dcb8c910100d943bfe11cf3254a201a8eb8f3f22940586c5d2612b4b

Observation 0a995200-d266-4af7-b4d1-4c749bb2f0a5 · outbound

This paper cites GPT-4o System Card.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.882424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.882424Z digest=sha256:b8811dfa9a5d416510c5ad3ea8001ae5032861544b2cdf03d3ff9c2c0e258037

Observation a1e9aacb-9249-43fa-a880-9ded340d290c · outbound

This paper cites OpenAI o1 System Card.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:37.947664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:37.947664Z digest=sha256:c5d994c3307d1809ab74cb2edcc020e95826d0f3adcccebb397b309bf7868574

Observation 6d2fc3ba-87a4-408e-b6d2-28629dc2b5da · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.034436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.034436Z digest=sha256:20cff13e7cc293938b367ff577896b071fdffa10ef3c7c0a7f4228eff4338eab

Observation ce92bb88-2148-40fe-9df0-df9730d36ce9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.107966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.107966Z digest=sha256:25ce338d3d8f50dc7f00281bcfabe43bf8fe1cbc8d12c9ca979532f064dac8ce

Observation e47b29ba-02f9-44ef-b388-744b0b1acbe1 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Evaluating Object Hallucination in Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.218393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.218393Z digest=sha256:02891d30fb57dfbea52ac2363436b1095c7c4f50bbe23c2c50f77a7688042cc0

Observation 29fc1abd-5a76-480e-a5ca-aa060e71d691 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.323087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.323087Z digest=sha256:4f5908808d35887504255fe91ea43a435904736fc61c89272c7efd7821c017f8

Observation fc37b8db-903b-4034-bbb2-e83c62082ef2 · outbound

This paper cites DeepSeek-V3 Technical Report.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.393410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.393410Z digest=sha256:e20319289e886465403a86f34dd55f302edc3970bc240b21cd5f836d50251d8a

Observation c142dbff-ded5-415c-9b20-b7d4a1992ca8 · outbound

This paper cites X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:41.520575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:38.463963Z digest=sha256:c10e1df1c06faeee83b1edb03db82433972f61108763671b0129949675067702

Observation 0000f601-0756-4c2f-8a0f-7a2471a0e18f · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vi- sion, pages 216–233.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mmbench: Is your multi-modal model an all-around player? In European conference on computer vi- sion, pages 216–233

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.690894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:38.513018Z digest=sha256:8bd22980478be0a3c042947508fd4fd13cc11e632c3d3826ef0dba4c06871a09

Observation 0e6e344c-68b5-455a-9592-b135010773f9 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.579018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.579018Z digest=sha256:8cbb5e90f414dec083a2cf31c7d20bba218d222d6e32f143b1ec2dbe0beee7e2

Observation cf81a703-ce5f-46bb-910c-e78427fd777b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.631494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.631494Z digest=sha256:99ce0243d16387866931594315bb53c098cdc12b1dcfb7792856b1c60b221fa7

Observation 5144df1c-c94b-4336-8a75-5c547156bb85 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.452752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:38.777914Z digest=sha256:cf54d89b937ee65a61988d76d6dd2f766af89456215952e5d95d2c1e9059c0b6

Observation 0f2f5a40-523a-4187-a9dd-eff88897d922 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.843269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.843269Z digest=sha256:5df02efd604754d56d73e56ed5d3aca9cc4e64ce07ad3dbdbb6102165ce84765

Observation 8f337445-6d0b-4852-9546-0e8596ac8d34 · outbound

This paper cites Claude 3.7 sonnet, 2025.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Claude 3.7 sonnet, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:43.186623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:38.930271Z digest=sha256:782c790f65cc61b8f27fa37d18e4b549e426f0842d7e09fc771afc2e2252f17f

Observation 38f60056-f59a-4e94-941d-da25dc558f65 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.038546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.038546Z digest=sha256:479bf2e9a24d77232bc87fcc2a7ff80a0b27a1bf0097f21dd26ca6137ba115ff

Observation 864c0936-66c9-4841-a69e-43aa9db98b84 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.192570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.192570Z digest=sha256:60b5a42449d25748420105ec252172122c86d849b08d6b05585696fcff53704e

Observation 33880566-25cf-42b6-8ffc-4764fa08ec7b · outbound

This paper cites Qvq: To see the world with wisdom, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qvq: To see the world with wisdom, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.896929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:39.353962Z digest=sha256:3e07080276e21150f9ff5324a64dc38b0c417146538972cd0ac6d8b3fee489dd

Observation b36a3b7a-3683-4148-9a7a-dd9fd47ab33c · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, 2024.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Qwq: Reflect deeply on the boundaries of the unknown, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.657113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:39.503756Z digest=sha256:b28e4767a09b7e9c027231dd48a32519d8aaf4da9fbaa0418efacf9e08c3ea66

Observation 944a66d3-3182-42f5-bbbd-ff2df9985f9f · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.658744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.658744Z digest=sha256:e99d0f2de36e6dfe66dd3460facb340894d63700e9e01d80fb034a12d4a3870f

Observation dc1e6874-59e1-4d60-ac44-28e7c5356027 · outbound

This paper cites Open-r1-video.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Open-r1-video

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.440908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:39.758478Z digest=sha256:e890e4f97aa640a5388cbeab137486229f81f677cd5962ae018b997a0f6b3522

Observation d176d38e-f478-4a60-859b-b22c4b999294 · outbound

This paper cites VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.855190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.855190Z digest=sha256:cfb98e4b37fe8228d8725382298f7782815004b52f8c2e2db7acc3ae72b0cc27

Observation d763cf7c-4589-4b5c-be57-5cf424486580 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.215107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:39.912917Z digest=sha256:c8bb4035e4210151a9b3dfc21b557b98adb96b38f6b2a3a43a089c63f8c80932

Observation 1e22d933-8446-44f4-b425-e1ea0b1ed865 · outbound

This paper cites Boosting mul- timodal reasoning with mcts-automated structured thinking.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Boosting mul- timodal reasoning with mcts-automated structured thinking

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:39.982779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:39.982779Z digest=sha256:4d9a67fc5912bc43e35100708d3cf0f7bc0bd3f74a2b138369c509b29fef37fc

Observation 47746207-6c7c-4ee6-8b2c-07e227d36c45 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.067346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.067346Z digest=sha256:6278417ad7e2a295008712bf7c248269ace07e1dd2c20b5c6cb29facbebc7ba4

Observation 58b2ef09-753d-40cc-9c41-ae15c01fabd7 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.126753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.126753Z digest=sha256:a1c6e44c65ab3ac6318ec4eeb2e0c84f03be4c9adff672e4c4003e4b292b9428

Observation ea48561d-de45-475c-a596-68d921cd0686 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.221780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.221780Z digest=sha256:ba732fcd67bd7ce827734868fedb4efa44b63ae3870c9b88ea8d08bcfcbe2aec

Observation fec71ebb-790b-405a-aa09-33befdc3f34a · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.304579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.304579Z digest=sha256:b4c790b475ec51f7fff934e0f0321212f39dcadf32b17687eebef5a778fc2576

Observation 2aacdaa0-ba2e-4fb9-901f-e0efd17ce28d · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.361835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.361835Z digest=sha256:99fdfe14e4d856561116568c72b33b436ad003b5317cb996685f3397b4870143

Observation 3559689b-0da9-42f6-94d4-496b6fe9deb6 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.429654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.429654Z digest=sha256:ae891a77a87408e3efa9765991ea33795e8837880109240e278d886db0897f6a

Observation 9ea9b439-fd49-40a5-9db9-5e1f2a8f0b91 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.493129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.493129Z digest=sha256:9bc5d2af5e4d4fa89577dd2d4fed79b770159a084f7b5bb2b020e5252d33513d

Observation cf7dae34-fce7-45c6-bffe-2d2f5dfb778b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.565270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.565270Z digest=sha256:57eb2f52e6baedd74d7dab9027378ff3ce73931b1c3c6fc3a0439f5840fe1467

Observation 19eed217-3bfb-42ce-86d9-feddd144e924 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:42.066729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:40.624530Z digest=sha256:ee56d6a46acdd89ae506f592a47253e261321f620d797624c5834f690b3ef699

Observation 4d8d3328-49e0-4ce5-b281-1c03ffe00383 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.702164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.702164Z digest=sha256:154428d8f054478d01723c11be076e63038bfc47a1006cfccc142d60bad4aafd

Observation fe7fe619-ec37-48b2-b76d-529d226dc326 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.775733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.775733Z digest=sha256:5e450b446ae4ab9a21de45b1ef81a85a0f9b4e365216f8391a4924e2a3c42bf1

Observation ee42f373-7c73-4b4d-b7db-39c4af95d126 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.883389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.883389Z digest=sha256:b49e0dabaa11cc8e5f31143bc5900e88ce30ab4d9ed0d650b590b9051135af4b

Observation 1fcd04ec-7e36-4357-9dbc-10dbab8f295e · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:41.899167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:41:40.941360Z digest=sha256:3f2a597c54777ce693153136a4a5864d2e23171fb1a59c290593c9ec78d88cee

Observation 9ca8fd4c-e733-4ed2-ab5d-dc3e9595b0a2 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI Improve Vision Language Model Chain-of-thought Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:41.040789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:41.040789Z digest=sha256:27472d6c8fdfbef7e7e48cd8ee9b258fd7e7a950b9435fd8ea1bd879b56abcf5

Observation 68b418f4-f233-4f31-b187-9a27584609bf · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:41.103071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:41.103071Z digest=sha256:ce1d9b78b2e9b06db62bef2e64eef43d7d0f6570f0d90b49f6243d226dd5a85a

Pith citing papers

Observation 41cb0031-bb58-4119-abee-969cf25cdbb7 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.768219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:e38b2ddfe1e8233976413e199d9738bd0abde469530999d154b109a05d6780c7

Observation 54e58cba-4bf8-433d-a37c-d4188b30aaea · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 224

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:45.693853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:45.693853Z digest=sha256:ffbd3615aceff455d468a8c8fa7f68c14c0e8db542f6f2b491fbd48f1cc2eac1

Observation 8c7ca8d1-37a3-4a88-9c36-c51ee1825ca8 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.540546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:20ac0dd5d23044a2b97c2abaac6bcc94374de77a713703f45ca12a1a7902fdef