Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:52:33.342830Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 32 inbound Pith citation observations for arXiv:2502.09621.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:52:33.342830Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.558808Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
42 of 42 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 5268d74f-0072-4c18-9d41-55e57ef0b1e3 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c66316-3e0b-46d4-98f2-843bcc714195 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d8a2952-ba3d-4120-a0d7-681deac05cc8 · outbound
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 749ba094-bf85-43c2-89b7-cb04ac67f73c · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 832d0332-acf3-4264-a511-6511a964f633 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Direct Evaluation Prompt Answer Extraction Prompt You are an AI assistant who will help me to extract an answer of a question
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1d09d559-de96-461e-9f8a-bcb115ffbfa3 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 241ae9be-adf7-4cd6-b396-f9284dd70b26 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7508e86b-7e80-43ac-9dac-77cc4b31164d · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7342f0ae-4038-414c-a4c8-f818b9232e99 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e591a6b1-a6fc-46be-94cf-34dd91b75fa4 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 746b7039-1864-4aba-8952-4eed97adca3b · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a29817cf-be52-47a9-bda6-8c4f82248fd9 · outbound
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 21d869f2-4f1b-4ad2-a4ed-51e2f37e3a85 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3e9c6a3a-f02e-4225-b0e6-bf7d80d07348 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1032d259-babe-4be4-b0f2-ae96d820a6e9 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency IMPORTANT NOTE: Evaluate relevancy independent of correctness
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 11f8dde7-ef55-4020-aa0d-c3f958bf297b · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation da8972eb-e330-4557-807f-c1afbb1d641b · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a4dad2aa-bddf-4ab1-bb8c-2ed28d09d63f · outbound
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f229736c-cd06-4e7e-a4dd-3cb1bf080ef1 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e884f6-1d79-4d8a-9e0b-bb4a451af5dd · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Invalid reflections include:
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43507bb2-4e81-41b0-b8b2-08a2c2fd8f49 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d429e9d3-0ad0-423d-9e8c-2bfa7c08bdae · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d04a9a0-b288-4756-af10-3de9cad5a1a4 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2abfe026-ca1a-44b2-83d3-562527c2b609 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ea6b87-a217-4af2-92ed-84ad8f2f2634 · outbound
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 658fb650-a3c3-4139-a9a9-be57ba9058f8 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efc26d3b-6a51-4a2a-90cd-51fabbc5dad7 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eecf3a5-432c-43e4-ad94-76153ff480ea · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e986f230-4515-4aff-b65d-5014966819f2 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0884fd-23a3-4fe6-bbcc-285e116b8222 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency You should directly output the choice letter of the answer
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6bb9ab0a-e049-4ff1-b1fd-3479e22a0c0c · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1e297348-8193-4029-9ab6-177516bf3c7f · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency [Non Multiple choice question]
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8796a0ba-ad41-48ae-89d1-40a8783cbe91 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency It could be hidden inside the last step of calculation or inference
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fd13c4a4-d569-4979-a6eb-28eafebd606c · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 48afadfc-f93a-4090-8f79-e0d194c86a7f · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Output Format: Directly output the extracted answer of the response
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bc756eaa-48ea-4aa0-a715-10c6b348493b · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f3237c9e-de57-449e-b153-4bacada0e294 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency [Nan-Multiple-Choice questions]
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a4529c65-ed23-4002-a967-8c6ca0b8cbf9 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f6f4944a-4182-4228-97af-349f37cbc016 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Output Format: 1
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ee09987c-0794-45a9-96ca-92086d10a819 · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency {In Context Examples} Question: {question} [Model Answer]: {extract answer} [Standard Answer]: {gt answer} Your output: 35
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 20d4ca02-ffc5-4564-a8f9-efdfbc6e8cab · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d918680-6fd7-4a71-929a-1afd04f8f45e · outbound
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency LLaMA: Open and Efficient Foundation Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26f2c23-f8ac-4194-a8a7-e5bcb594f061 · inbound
Chimera: Improving Generalist Model with Domain-Specific Experts MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da6b807-10dd-4f9b-ad92-86135e062636 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c05d8c33-2704-4454-b15a-81467a564cc8 · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43395219-3aa2-4dea-8dbb-8e7c711bb315 · inbound
GenCLS++: Pushing the Boundaries of Generative Classification in LLMs Through Comprehensive SFT and RL Studies Across Diverse Datasets MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f638547-27f0-4112-b319-ff9b6036f745 · inbound
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d553872c-3bdc-45f3-84b6-02daa4b3731c · inbound
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c8bc35-93ae-49e7-8dd3-60f36a5cfcb2 · inbound
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64fa8169-87f8-49cf-93a9-1ef6df20847e · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 133
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62227da7-50b8-4328-beae-82ff7276f698 · inbound
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13b0d27-6df3-4901-8c5d-bd1ed2145f42 · inbound
Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd3064e3-fe71-4bab-b0a8-eada1baeb521 · inbound
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a72d196d-62cc-4209-83be-37aec9055dc7 · inbound
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f385f3f-9ae7-46d2-8098-d70880665cf6 · inbound
Reinforcing Video Reasoning with Focused Thinking MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 110bcaec-d601-4d7f-807d-775cf6756cf5 · inbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac6d4c7-64c9-4415-9f84-4ca332bb8d68 · inbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28d5785d-4207-4f33-ae33-38be8ef5b07a · inbound
VGR: Visual Grounded Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6d2fc3ba-87a4-408e-b6d2-28629dc2b5da · inbound
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537b8e27-1d48-45fe-9349-229f5710d9d2 · inbound
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b305162-d8bf-4a8c-b1ec-785991a2c9be · inbound
NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3aaa64e7-2f99-4fcc-81b0-00f248706cd0 · inbound
AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 56a1bbe8-72e1-4ff9-82a5-9707470a5fb8 · inbound
Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f622eb48-a493-4d6c-8d6d-0f1cb57d055e · inbound
Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification? MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7c62b84-fba3-43e2-9958-c5d17e716a56 · inbound
M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22de22f5-fc81-4bfc-863a-512eface1961 · inbound
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67411b12-62ff-4ea4-aa44-03ef635dfc84 · inbound
Diagnosing Pathological Chain-of-Thought in Reasoning Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30637812-e859-4696-a2b8-697ebc23ace7 · inbound
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b2fe1677-9a24-4a1f-a230-d30daf5ec570 · inbound
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 06e5d8b0-8dd7-4a95-84c8-b9f2c1406538 · inbound
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 11673db8-5d71-4ee3-8f23-1c19f5468d58 · inbound
Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7ed23d51-c75d-4921-97f1-fbf624222c88 · inbound
OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 77198966-2992-4550-b956-aed73dcacc83 · inbound
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8a1784-94f6-4d85-9bb5-47133c928073 · inbound
OpenCoF: Learning to Reason Through Video Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.