Pith. sign in

Paper Citation Record · LEDGER

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 5 inbound Pith citation observations for arXiv:2506.04280.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04280 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:20.044416Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:54:43.474266Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T00:12:50.323154Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5391bfd9-4e4c-4283-b50a-84b1f44de3e1 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.093166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.093166Z digest=sha256:5f2caf4597a7dcafde279e8872c814f2f2be8a6e6980706da06dc1af4b1e2193

Observation c72cb348-886c-46ff-9438-363f60a1b434 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.125992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.125992Z digest=sha256:80676d64c3a57ed126351692cc2930344b572734632ee54cc5464adeedc3a7ae

Observation c92bdff2-402b-46bf-92af-c5ffe95e425b · outbound

This paper cites Qwen2.5-VL Technical Report.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.149385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.149385Z digest=sha256:26fde2e33d41350a5bba7fd12d24149dfaf954d0058fb05bc4510be44025e101

Observation c657ea92-7c09-48a2-9414-40d49c3df2a0 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.168229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.168229Z digest=sha256:d60b134e3af6f64d2c0a619c3348ec9002ef4edf0678e8e41b6eaed09e2deb0b

Observation 449cf1fb-3f45-442d-9e5a-69ee1fe9b17a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.206037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.206037Z digest=sha256:2f455a6a5f5ed0997d01919f98a4791fd1af9c243ee3d8d6340d3364373abb61

Observation b2332d1c-679a-49e9-9264-e55557eb3bbf · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.314570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.221144Z digest=sha256:c0a100733a317822fce36989e1686f8bb9b123fb397443eabc9f650e25f51b04

Observation aee44f84-b50e-4f52-8e75-f1d730020eeb · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.232426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.232426Z digest=sha256:f27568249133d6d9fe2be16c433161bc9393d527a5bffbd7bd6398b8f4390251

Observation 4f0b9bfc-e67f-4693-b269-3faff3a84983 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.246975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.246975Z digest=sha256:77fafa4f5f6ee8af8a06ee18fca887c4746037a6202577acb4f5e82dac0b54ad

Observation 3769c6c6-8602-491c-8a98-64c15c0dde53 · outbound

This paper cites Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:05:21.666096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.279701Z digest=sha256:d12e39c8fc96da0225d7053f7123acab6a3b7a47c34d626a1ed67c17dbc68ab4

Observation b83ffd3f-dd39-4fec-9365-ec175c24b1c8 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.286471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.286471Z digest=sha256:cefea4983756b46c9208d9adc9139601da1354198564a08a676491ed29546273

Observation 87f573ed-96f7-49dc-b444-396ddc047ac6 · outbound

This paper cites Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.291846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.291846Z digest=sha256:40944f8d8052e7602009a4458a173d445986fdab259531e5a420a67c762eb027

Observation 2a463a9b-c158-4504-b449-7030e94417d5 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.067313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.302473Z digest=sha256:7fa8c4ff97feb216b9517a155d02a7ef47645eac167dbf8c105412da097dbc37

Observation b7a365eb-80de-45f2-8806-193219db6b68 · outbound

This paper cites OpenAI o1 System Card.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark OpenAI o1 System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.318506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.318506Z digest=sha256:d41fbed7f844ad0a834334ef9cd16fc33784f264fe08f4419dc94e4ae417e4fb

Observation 3783139e-c5f8-4251-832e-b1894c57c50b · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.352234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.352234Z digest=sha256:89ba878e56857ee3d81eb5dc8c8ce1bb83f1b2a95287e835cec9e5553066c2c8

Observation 110bcaec-d601-4d7f-807d-775cf6756cf5 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.357259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.357259Z digest=sha256:8592147a7ec89b77b67316ae1580b0a0b499b20d8d1805970833b6e91bb7928a

Observation 7a7f698c-e082-4f45-9278-6ee28fb1844e · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.033572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.373052Z digest=sha256:0c3dcb5a7b3f260665074b00bd6be23c9c2cde7e037daac17f3c152d3018076f

Observation b7622b24-9377-4fb8-bc24-f16a898b13f7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.396582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.396582Z digest=sha256:4fd43aa48d0dc56b9c6de6fac3091f87c34aac045314e539d4c4b7cf0fdce40b

Observation e7821c05-6750-4e3b-91b5-d3fde00a313e · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.405502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.405502Z digest=sha256:f8c6cf711133476cbc6aba71f04ac334a5d063b75c25a0c7c8b38bbbe3ba54f2

Observation e2f30b36-44ac-44b6-a637-bb39ff2dd03a · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.423016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.423016Z digest=sha256:bc87dd0b66cbf960cfdefc40f1582a663c9829a79f75ff7ebdaa939b6c226fd8

Observation d1523ac9-2457-485e-81b5-2463cdc7d8fe · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.986394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.386373Z digest=sha256:2f31acda3421bf6bfbd99e671da5cd00366f46d19763a7557aab6d05ae5ece0f

Observation 813a2978-95cb-40c9-b6a3-adacfaeb7e29 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.889599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.445145Z digest=sha256:99f87c091dd92b2c43d63f588dbc10b18ce42a741d36d5ba9fa6cdeb9fe194f6

Observation 55e9a714-8997-4266-9b20-f54251a95a86 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.462282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.462282Z digest=sha256:90e94c9e1162817fe484a23f4960065e98ea6fa87c6e2e795ef8eefd78342c6f

Observation 57dc166e-b0ca-4809-8d63-1189cd44c09b · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.472918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.472918Z digest=sha256:a9f374ab897d62465824199e7b515856c496b2f989a3e8dbdec7329e776fbb46

Observation 934bdc89-a9a3-4304-b26a-b210f382595f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.927157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.432019Z digest=sha256:cd2b94a87dc7df22b044c558bf2c4faa51727f21a1a00082f510c3d2559371ca

Observation 383301d0-48b3-47a5-a9fa-aff5a789d17f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.810498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.491198Z digest=sha256:8f9328bb6d1534eaaaf058dc88f2685c2ac0cb70caa474c493258cdc6a067326

Observation 28d0fb4b-8fb2-40ad-906b-aae4f0d84da1 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.499218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.499218Z digest=sha256:40ff57a46e5d130ffcbe42b4b099d323304adcf40ebca4e37ee522395c4041ab

Observation 0d9ba4f4-81b9-478f-84f5-54dd0f3419fd · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MileBench: Benchmarking MLLMs in Long Context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.523326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.523326Z digest=sha256:188590c3a0704d3aff1bde7af0410318c137380366073b34140b0bf6f78c3f9f

Observation 7ac23c4c-fbae-43fb-bea4-f92333afc950 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.483720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.483720Z digest=sha256:86635fe51ec0769c7b604900881c70f7e84adea5fe889dd1350bedf46ad5dbd9

Observation bed07209-e2ff-4850-9771-6ba94f4008fc · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.744082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.546625Z digest=sha256:351c03265c6e385876859ddb0c1264a77c194d3981ed1f771221248c7dba8b7f

Observation 2a7da2ce-8183-4bf8-b05e-358ed8c1ce6e · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.555184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.555184Z digest=sha256:3df2ed2787effa36f325c18228911a4a1aecbcd30e3a20b761ffe68a35c73f47

Observation 2c0f76dd-9160-4dd7-8fd6-571bc8943ae7 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.565467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.565467Z digest=sha256:6c8e28e22aeabd945ad95976a08d675846d2e948902fcdda4e2561bb8fed76ad

Observation 991b90ed-dccd-416d-b568-5e449d05da9f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.580744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.580744Z digest=sha256:e2a1f8ad6e3c01e6506d34572b09f122b02ac274d9f63a3f7f755f93d06e73bb

Observation 072f0cd2-45ba-4135-88d6-0ead62ec9e8c · outbound

This paper cites NLVR2 Visual Bias Analysis.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark NLVR2 Visual Bias Analysis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.537088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.537088Z digest=sha256:323106a7dbe63552a4008dda6da8448d620e227f0d688d3bf6c5e6dfcbaf1dff

Observation 6290e526-21b7-4b54-ad8e-4f63cb734c17 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.613480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.613480Z digest=sha256:03e149ffbd55ee1b480d0b20a91feac383d5b1372b53d7febcabe318c2672f11

Observation 0c55dbbe-cfdc-41e9-8a47-d5c02a572196 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.676752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.622195Z digest=sha256:8831f8fca55e73e27779ff428b6ee46c321ee7a16353886488203bd5b82b4869

Observation 4a5f715c-91b1-4750-bea6-34f1d1fa0f3d · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.643578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.643578Z digest=sha256:1b69522aa6e7f2374f810d6127292e090804c80a147bac40df0c75f30750635b

Observation c7c041f0-6fd7-491e-b307-15c86595814e · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.654435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.654435Z digest=sha256:365a04cafdfd6325c47644ba3dd71848c90057a1ca37a0ff02f80837a46e46a4

Observation daac543d-7b2b-4ab2-a972-1effe3b51c83 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.601376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.601376Z digest=sha256:8882e170da765cc8f1c588f65d3d614af737059705d159a1862fec1e3abf1ea6

Observation 0921162d-cacd-4676-966b-647ed987a204 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.683399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.683399Z digest=sha256:618e7619a9a0c7bd460eb57ede0d5c93a3bffb5cec7628bc0daa995f58456bae

Observation 7278bfc9-b506-4303-a6a8-b0c457cedc3d · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.704763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.704763Z digest=sha256:49a76cbbc21193aadf4ba1d72ae452a05a40ac9d9e8fd53510e996779f101981

Observation 7e2399a7-c682-4cf5-9248-a7ad6c9ad0bf · outbound

This paper cites Qwen3 Technical Report.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.713959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.713959Z digest=sha256:090676aa4a78b111067d91a94c6273939b3d9ce333b941ff56c4863b9dcfe43e

Observation 4f95da42-0fbe-49f0-b7a7-b9d955e24ffa · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.719944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.719944Z digest=sha256:1e85b18e15615323df727528f1af8b9a53858039aaae0a223dc3a51d907ed7d7

Observation 406e5081-45d4-448a-93c0-300a95aa87e7 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.667340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.667340Z digest=sha256:1db5dc778436e38c3f878fd01e00e94eec617e60d85aedde4ac16b9e61d6cc5f

Observation 869ae997-430e-4c73-89a3-0a3e0ab35c59 · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.748441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.748441Z digest=sha256:9bb31676aca53f2925d486b9456bc10926059ffc81ce95dabebe29d66213e4d5

Observation b207ca52-e644-403c-a340-c47554341d74 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.759069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.759069Z digest=sha256:c69f505d019def4cce8173b9060cf65a7f3625048c65d94f022d203d4b7ce77c

Observation 04bcbfa7-9314-4b0c-83ac-9331cbf40313 · outbound

This paper cites WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.768115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.768115Z digest=sha256:ea3299f71c0d5db8c75177aa7eefd2d7c60464731b4bf4687abae1c2ad7aa2ba

Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.779262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.779262Z digest=sha256:ac434a0ec8612d77ee879a21e11731ba7ece3bb6a3bae67ca77da315fc154605

Observation 674a23ef-a859-4fa0-abfb-8c3dd55cf2d9 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.737053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.737053Z digest=sha256:4de89a17f50daf72a59a6b470d5c5c70e204d40c8a2f1310c11a7ce205c9a4fe

Observation 374273db-e653-4a63-945d-d745994c824d · outbound

This paper cites MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.808950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.808950Z digest=sha256:5a3ece1eee4f369e19709a63be2b71006f07f15915709c9c50e8777121c86a20

Observation 242fbcc3-9161-4688-add2-97f2000fa578 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.821194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.821194Z digest=sha256:5e0ad159c8cabf8ef9001139c45a7ba2db40f7eef0e1d25cb56181eb7382ccc7

Observation 48a225fb-5b1d-4cf3-8168-4a210622fcce · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal Chain-of-Thought Reasoning in Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.790349Z digest=sha256:1b71d961b2b7895092294212aff69035b3731c4e0c25baa7a3c5df6735407a71

Observation 1e7286e3-27d0-49ca-b3d0-4040e9b2d71e · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.541576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.834814Z digest=sha256:f9d833593738fbd5f42da7aa609e278c8254b88aee9fcbd8fb88f6bcfeb50198

Observation 9afa9931-c062-408b-a212-7cdd30230c4c · outbound

This paper cites reasoning step.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark reasoning step

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.504739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.856925Z digest=sha256:6d8329ccd64e2eb1207f7a21b0cd1bebe8a4cb1faccf166e67f9ef9ac29819fb

Observation c25669a8-8996-44f7-be32-3c9ab4488f32 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.311098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.900673Z digest=sha256:7fac81bc14b4e49f8800f981d573b7438e1dd30f0736eea3c3717f590f88d32b

Observation 09ab0e9d-67e1-47e7-9269-464b8458d6a1 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.259112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.946132Z digest=sha256:e97469f3ee78f8e17582eff263b4a854a13ef4ef647389cb54ef5eaf2e3197ea

Observation 2cb14538-63db-429e-98ba-48aa6e23b6e2 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.465449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.961296Z digest=sha256:6e9dfdb0286bb94faf2cfa68a949194e15e7a0e319c9db458eab357b21f8edfb

Observation c130aa28-c9d8-4f8c-93ac-78e6f1e19275 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.395288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.971376Z digest=sha256:f8598324717fcf1dae3c90e0eefc5366a75dfd965dfa0b6b1bac20fe971abab0

Observation b7038d16-669d-460c-a5e0-4ef8f49abba8 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.363949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.980113Z digest=sha256:f35adbb6cc9f8b7fd53a4cca5beb3c846156742851cadc0738e1f26928a85c03

Observation 56aba37e-d952-4cb6-ad57-00930f296933 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.199841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.988591Z digest=sha256:b8fc91de415207438c83995e1396c946fa07dafbe836ae8ce8aad24894f1f2a4

Observation d4e85cc4-f8a5-487e-b582-eee8aea80311 · outbound

This paper cites Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.116840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:20.013797Z digest=sha256:e4629521dd224b4c74995a22105f8ef868dda950e936e3d5b5bc9a46bf1cab32

Observation 78a0e081-f8d1-4af4-bf4f-192fa85a8888 · outbound

This paper cites Rank higher the responses that least misrepresent these relationships.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Rank higher the responses that least misrepresent these relationships

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.066928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:20.026221Z digest=sha256:4d6c24f9c47ffbe7f6f3e8a395b456084db7c41451926cf2aee1ec1ca7487795

Observation 88f71158-6ef9-430c-88b4-7f57a2d86032 · outbound

This paper cites Responses should avoid inaccuracies in describing the characteristics of the objects present.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should avoid inaccuracies in describing the characteristics of the objects present

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.020746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:20.034083Z digest=sha256:b8cb53fe54f9d24776ba055a3b0ee407a126ead1075fc42c83a89ace53c88859

Observation edc4972b-8872-4d0a-b8bc-92ad31259023 · outbound

This paper cites Equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Equally good

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:21.962642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:20.044416Z digest=sha256:a9534be564f2df0cb1855ba8159a65f9a0fd6f07baf46b73a26c24ead9a448f2

Observation 7994fa18-e622-4d8b-b9ee-64d433200258 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.509807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.509807Z digest=sha256:868650b1a5ec0a053902baac01e4086096c97c522557a6287931481e3526a23a

Observation 3823083d-0afc-468d-bd06-d4c92fc52bff · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.114146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.114146Z digest=sha256:473aca545f6ce803b91a2051ab198bd0c7faac8ba849a992a154c61415b01978

Observation c013ae08-676f-4166-8fbe-e7b9fac4652c · outbound

This paper cites InProceedings of the IEEE/CVF international conference on computer vision.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InProceedings of the IEEE/CVF international conference on computer vision

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:23.155612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:05:19.262204Z digest=sha256:83b0278f6be6e0a96010e186f6fec10e1b1b4cdb442febec09a6a7ac955987a1

Observation 2760d9a4-0362-4c1d-a434-7a40e9ef31e0 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.193447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.193447Z digest=sha256:204d3c8e237dff4b4d517cf6d312509eeb71f0259b2f935c5639a570fb72b65d

Pith citing papers

Observation 5c6864df-b047-4bee-bf80-fedc21632611 · inbound

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation cites this paper.

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:43.474266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:43.474266Z digest=sha256:6616d4dedb87935aadd489d2ff0580a2f6d6cc8b3610c185e539836d57dad52c

Observation d334f9d5-9cff-48f6-91d2-3fba26525cc2 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:13:24.779717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:96ddca7663088f91a96b676031c4710887e65eac1442828e467b2e2bb5996ba9

Observation 92fa0042-066d-4efb-aa58-7d57f2915315 · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.830119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:b351b60bdfe612ac6319f2e49ae1f82dbba147b2e64d142288a39b8c4d2f5002

Observation 8425e803-0d62-4c85-bd60-4fbd56364dac · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:02.750336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:dc9299b5dd2d37e6390fc9d3be52a7bdc18a541b9deeb74fcbec75d6add0aea3

Observation a1c2534f-54a4-4485-8371-831cfd49b2d7 · inbound

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning cites this paper.

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.324575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:15:56.598968Z digest=sha256:44c3d463e6e0a0c6d86ac668c8cb75ad43ea2de9d1404dce767aad74a1f04c19