Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:20.044416Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 5 inbound Pith citation observations for arXiv:2506.04280.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:20.044416Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:54:43.474266Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T00:12:50.323154Z
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5391bfd9-4e4c-4283-b50a-84b1f44de3e1 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c72cb348-886c-46ff-9438-363f60a1b434 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Gemini: A Family of Highly Capable Multimodal Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92bdff2-402b-46bf-92af-c5ffe95e425b · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c657ea92-7c09-48a2-9414-40d49c3df2a0 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449cf1fb-3f45-442d-9e5a-69ee1fe9b17a · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2332d1c-679a-49e9-9264-e55557eb3bbf · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aee44f84-b50e-4f52-8e75-f1d730020eeb · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0b9bfc-e67f-4693-b269-3faff3a84983 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3769c6c6-8602-491c-8a98-64c15c0dde53 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b83ffd3f-dd39-4fec-9365-ec175c24b1c8 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f573ed-96f7-49dc-b444-396ddc047ac6 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a463a9b-c158-4504-b449-7030e94417d5 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7a365eb-80de-45f2-8806-193219db6b68 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark OpenAI o1 System Card
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3783139e-c5f8-4251-832e-b1894c57c50b · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 110bcaec-d601-4d7f-807d-775cf6756cf5 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a7f698c-e082-4f45-9278-6ee28fb1844e · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7622b24-9377-4fb8-bc24-f16a898b13f7 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-OneVision: Easy Visual Task Transfer
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7821c05-6750-4e3b-91b5-d3fde00a313e · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2f30b36-44ac-44b6-a637-bb39ff2dd03a · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1523ac9-2457-485e-81b5-2463cdc7d8fe · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 813a2978-95cb-40c9-b6a3-adacfaeb7e29 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55e9a714-8997-4266-9b20-f54251a95a86 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57dc166e-b0ca-4809-8d63-1189cd44c09b · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 934bdc89-a9a3-4304-b26a-b210f382595f · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 383301d0-48b3-47a5-a9fa-aff5a789d17f · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28d0fb4b-8fb2-40ad-906b-aae4f0d84da1 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d9ba4f4-81b9-478f-84f5-54dd0f3419fd · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MileBench: Benchmarking MLLMs in Long Context
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac23c4c-fbae-43fb-bea4-f92333afc950 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed07209-e2ff-4850-9771-6ba94f4008fc · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a7da2ce-8183-4bf8-b05e-358ed8c1ce6e · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c0f76dd-9160-4dd7-8fd6-571bc8943ae7 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991b90ed-dccd-416d-b568-5e449d05da9f · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072f0cd2-45ba-4135-88d6-0ead62ec9e8c · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark NLVR2 Visual Bias Analysis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6290e526-21b7-4b54-ad8e-4f63cb734c17 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c55dbbe-cfdc-41e9-8a47-d5c02a572196 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a5f715c-91b1-4750-bea6-34f1d1fa0f3d · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c041f0-6fd7-491e-b307-15c86595814e · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daac543d-7b2b-4ab2-a972-1effe3b51c83 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0921162d-cacd-4676-966b-647ed987a204 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7278bfc9-b506-4303-a6a8-b0c457cedc3d · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e2399a7-c682-4cf5-9248-a7ad6c9ad0bf · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen3 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f95da42-0fbe-49f0-b7a7-b9d955e24ffa · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406e5081-45d4-448a-93c0-300a95aa87e7 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-Critic: Learning to Evaluate Multimodal Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 869ae997-430e-4c73-89a3-0a3e0ab35c59 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b207ca52-e644-403c-a340-c47554341d74 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04bcbfa7-9314-4b0c-83ac-9331cbf40313 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 674a23ef-a859-4fa0-abfb-8c3dd55cf2d9 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 374273db-e653-4a63-945d-d745994c824d · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242fbcc3-9161-4688-add2-97f2000fa578 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48a225fb-5b1d-4cf3-8168-4a210622fcce · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal Chain-of-Thought Reasoning in Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7286e3-27d0-49ca-b3d0-4040e9b2d71e · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9afa9931-c062-408b-a212-7cdd30230c4c · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark reasoning step
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c25669a8-8996-44f7-be32-3c9ab4488f32 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09ab0e9d-67e1-47e7-9269-464b8458d6a1 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2cb14538-63db-429e-98ba-48aa6e23b6e2 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c130aa28-c9d8-4f8c-93ac-78e6f1e19275 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7038d16-669d-460c-a5e0-4ef8f49abba8 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56aba37e-d952-4cb6-ad57-00930f296933 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4e85cc4-f8a5-487e-b582-eee8aea80311 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78a0e081-f8d1-4af4-bf4f-192fa85a8888 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Rank higher the responses that least misrepresent these relationships
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88f71158-6ef9-430c-88b4-7f57a2d86032 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should avoid inaccuracies in describing the characteristics of the objects present
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edc4972b-8872-4d0a-b8bc-92ad31259023 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Equally good
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7994fa18-e622-4d8b-b9ee-64d433200258 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3823083d-0afc-468d-bd06-d4c92fc52bff · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c013ae08-676f-4166-8fbe-e7b9fac4652c · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InProceedings of the IEEE/CVF international conference on computer vision
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2760d9a4-0362-4c1d-a434-7a40e9ef31e0 · outbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6864df-b047-4bee-bf80-fedc21632611 · inbound
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d334f9d5-9cff-48f6-91d2-3fba26525cc2 · inbound
Multimodal Language Models Cannot Spot Spatial Inconsistencies Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92fa0042-066d-4efb-aa58-7d57f2915315 · inbound
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8425e803-0d62-4c85-bd60-4fbd56364dac · inbound
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1c2534f-54a4-4485-8371-831cfd49b2d7 · inbound
StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.