Pith. sign in

Paper Citation Record · LEDGER

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

As of 23 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 5 inbound Pith citation observations for arXiv:2506.04280.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04280 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:20.044416Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:54:43.474266Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T00:12:50.323154Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5391bfd9-4e4c-4283-b50a-84b1f44de3e1 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.093166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.093166Z digest=sha256:edb173ff77548292fa73fe9a691bed67a6ecb5fcb718e099f3798c524106cc0d

Observation c72cb348-886c-46ff-9438-363f60a1b434 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.125992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.125992Z digest=sha256:fe08e5c1960209f01077bfd47158ad5393d9ba6e370886a870853a1d12f53a6c

Observation c92bdff2-402b-46bf-92af-c5ffe95e425b · outbound

This paper cites Qwen2.5-VL Technical Report.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.149385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.149385Z digest=sha256:9212ff2301eb7e13c285e9533bd601ae12f6c34ea8a910a2109d4f85bb2e305a

Observation c657ea92-7c09-48a2-9414-40d49c3df2a0 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.168229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.168229Z digest=sha256:a2b66a702343718ff186a38abfcb171eefde98dc3075af0bd19b1e0d57a4c979

Observation 449cf1fb-3f45-442d-9e5a-69ee1fe9b17a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.206037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.206037Z digest=sha256:9b18fc3d4f238bf1090e0d7e5fe3332f8f66af08ef23e92398ed3b15f8b573e5

Observation b2332d1c-679a-49e9-9264-e55557eb3bbf · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.314570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.221144Z digest=sha256:30cf10e52ad8d852945d04de0fd09fee6a71f08bf690e209c94965228497d0f0

Observation aee44f84-b50e-4f52-8e75-f1d730020eeb · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.232426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.232426Z digest=sha256:261a4766b2c7e34ffd4040549cdaa23aa368e2fcd10387a2b82d815810571056

Observation 4f0b9bfc-e67f-4693-b269-3faff3a84983 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.246975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.246975Z digest=sha256:2ca14d46e78d7dd83e8d2563636c9c929dfb40f7845220b6b369efa5093407f3

Observation 3769c6c6-8602-491c-8a98-64c15c0dde53 · outbound

This paper cites Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:05:21.666096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.279701Z digest=sha256:bb815d30c05e6e5cc85e859601997cab70c9804998ba037c25b43871d19ed57f

Observation b83ffd3f-dd39-4fec-9365-ec175c24b1c8 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.286471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.286471Z digest=sha256:696bc002ab2e81cb81e565ffd4ed975511a62e18c565445be7b65686cf0c5c2c

Observation 87f573ed-96f7-49dc-b444-396ddc047ac6 · outbound

This paper cites Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.291846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.291846Z digest=sha256:f6e2a6d469d4865d1e422ce95d6ded1957e662eadf6ed819dd7e198399e6bfe5

Observation 2a463a9b-c158-4504-b449-7030e94417d5 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.067313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.302473Z digest=sha256:45f88359ea03adcb81b3bd123478cc1b7c02ff8434a31d93094fbfec96412994

Observation b7a365eb-80de-45f2-8806-193219db6b68 · outbound

This paper cites OpenAI o1 System Card.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark OpenAI o1 System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.318506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.318506Z digest=sha256:2470883b18449e96eebc21a91d285dd38a71049f809a0567dc220089df76818a

Observation 3783139e-c5f8-4251-832e-b1894c57c50b · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.352234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.352234Z digest=sha256:caba8942c8febccbd61f5ecf1a8b8fa5bc04cd2605d9fc296cafba1f75e90ce3

Observation 110bcaec-d601-4d7f-807d-775cf6756cf5 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.357259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.357259Z digest=sha256:ba3772fe5b4b21f8dbcff994226c9af6fe0838b7224354464db3ecec60c4a409

Observation 7a7f698c-e082-4f45-9278-6ee28fb1844e · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.033572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.373052Z digest=sha256:aeab5c0f07236966be189760b2a38f8b31234797805cab877c4711356eea8994

Observation b7622b24-9377-4fb8-bc24-f16a898b13f7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.396582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.396582Z digest=sha256:d3db066aaf7b546335f850270dc572c2eca4dbe3aa714e7046ee4f8cf38118b7

Observation e7821c05-6750-4e3b-91b5-d3fde00a313e · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.405502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.405502Z digest=sha256:fcc37e9ac886757eb791733113b4dd218afc463ecef802d12440207bc7235a29

Observation e2f30b36-44ac-44b6-a637-bb39ff2dd03a · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.423016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.423016Z digest=sha256:820dc56383152f7d3dab105a7bb1456865e2152ef587c592354920d61c19cd44

Observation d1523ac9-2457-485e-81b5-2463cdc7d8fe · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.986394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.386373Z digest=sha256:355e1a47e4b3b838a03bae2ead51332f7a4e63a819acd96a6531ccd8bb0f1811

Observation 813a2978-95cb-40c9-b6a3-adacfaeb7e29 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.889599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.445145Z digest=sha256:907e181fc7831f69d6018921bc7c65ba94894264611bf2dcc87701fc603eb594

Observation 55e9a714-8997-4266-9b20-f54251a95a86 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.462282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.462282Z digest=sha256:49881ed1a1a690375027b407d4c947f57e67bcf37fb871c93baf8afc7191174b

Observation 57dc166e-b0ca-4809-8d63-1189cd44c09b · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.472918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.472918Z digest=sha256:fc1ccf7005137f0383a763bfe92277553c2945f1e20e6359a8be75e4e1ee8773

Observation 934bdc89-a9a3-4304-b26a-b210f382595f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.927157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.432019Z digest=sha256:c440840ffc0548455a9cc40ecda6acb2c01ed46cbfd570b55f11b9e7ed5e9976

Observation 383301d0-48b3-47a5-a9fa-aff5a789d17f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.810498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.491198Z digest=sha256:1a8c427e1795ac666fe674329c715dd10ca2986b4d6a586d10924a97d92d9dfc

Observation 28d0fb4b-8fb2-40ad-906b-aae4f0d84da1 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.499218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.499218Z digest=sha256:ff21596c06816418a702fbcf823bc304c4bfe7fae604e46297ab6bcc7af073e2

Observation 0d9ba4f4-81b9-478f-84f5-54dd0f3419fd · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MileBench: Benchmarking MLLMs in Long Context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.523326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.523326Z digest=sha256:3164c3e38ae493e190ddb69d95f8ef7e3cc58a17119acd763e6fb339a056ac05

Observation 7ac23c4c-fbae-43fb-bea4-f92333afc950 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.483720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.483720Z digest=sha256:26b9c3a74f1182bd48148c903f17326872d628bcb11fa67a1f812003b7d7edd2

Observation bed07209-e2ff-4850-9771-6ba94f4008fc · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.744082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.546625Z digest=sha256:3d53fd3f97a8e77f7030b0619c85fe5c2b633dd867681a83d3f2fdaeb78fd970

Observation 2a7da2ce-8183-4bf8-b05e-358ed8c1ce6e · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.555184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.555184Z digest=sha256:cf2359331bb24325b296309a7a9b9b1ac6cc8c9d96fc397401ccea40317c3ac7

Observation 2c0f76dd-9160-4dd7-8fd6-571bc8943ae7 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.565467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.565467Z digest=sha256:8530814a3ee9f9f0aec3fa9356063ba88fa5a1d69e67cf97d15b168bdcf8de81

Observation 991b90ed-dccd-416d-b568-5e449d05da9f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.580744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.580744Z digest=sha256:698726b333037094964d55cdc956afe014ce780a861b2e8ca64ab39c8f633acc

Observation 072f0cd2-45ba-4135-88d6-0ead62ec9e8c · outbound

This paper cites NLVR2 Visual Bias Analysis.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark NLVR2 Visual Bias Analysis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.537088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.537088Z digest=sha256:0ab467b066c23aa653df4fa746a67c27884c32665a42d0a2ebc4e77d1301801b

Observation 6290e526-21b7-4b54-ad8e-4f63cb734c17 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.613480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.613480Z digest=sha256:95252599d6c267756800217f2a8af2442ee256e8aebe7650ed1d74732f7211e9

Observation 0c55dbbe-cfdc-41e9-8a47-d5c02a572196 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.676752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.622195Z digest=sha256:f444d35a48afed330946a1fb401c03d4270a33316007666aa3f35f856cfd888e

Observation 4a5f715c-91b1-4750-bea6-34f1d1fa0f3d · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.643578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.643578Z digest=sha256:150559304579b15eea89448137c38833eeef15004e1c225bba40ca8258e34233

Observation c7c041f0-6fd7-491e-b307-15c86595814e · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.654435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.654435Z digest=sha256:ab0c56d67b22593f5984da79240dfc03307dc7488401e97263c9ef5c477c59e7

Observation daac543d-7b2b-4ab2-a972-1effe3b51c83 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.601376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.601376Z digest=sha256:a231602b1eaf560515ca93dad54673890562ebb5bd6ce3ac1e7ce90fb9acd679

Observation 0921162d-cacd-4676-966b-647ed987a204 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.683399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.683399Z digest=sha256:8992d74805c497ce0a724a5b326f929f58e331d4badd6fac5ae9d82606a754a1

Observation 7278bfc9-b506-4303-a6a8-b0c457cedc3d · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.704763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.704763Z digest=sha256:b59a7f1be4acb23d1901b8b7717a3d8f044acaa1baa8cdb5bf178add4860295b

Observation 7e2399a7-c682-4cf5-9248-a7ad6c9ad0bf · outbound

This paper cites Qwen3 Technical Report.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.713959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.713959Z digest=sha256:86a7e75eae05e40342e23bbdd152d45dcf53bb3c058bd450bec014bc7f97dd1b

Observation 4f95da42-0fbe-49f0-b7a7-b9d955e24ffa · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.719944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.719944Z digest=sha256:918aee8880304b56bccb557d3bbc6c6e60effcf37c9b7b41ead9cc77de4af918

Observation 406e5081-45d4-448a-93c0-300a95aa87e7 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.667340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.667340Z digest=sha256:fc0d0a8874c318ad75f1c13dde857d8528f9794fcb5e0ebfb5d8f0beae4f15a5

Observation 869ae997-430e-4c73-89a3-0a3e0ab35c59 · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.748441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.748441Z digest=sha256:5b46bc61e42ba54ea1aa5a8263cf4cae402cb006518ea7127709d11f214c5521

Observation b207ca52-e644-403c-a340-c47554341d74 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.759069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.759069Z digest=sha256:745cd8e20b87bcff6259707ac1723a07e9664c4072b495637b39c8aad7d8b0e1

Observation 04bcbfa7-9314-4b0c-83ac-9331cbf40313 · outbound

This paper cites WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.768115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.768115Z digest=sha256:c042b534516625f2d4306a259438aed0693c59d4e41cbf09c465688d57e50fc5

Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.779262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.779262Z digest=sha256:cbd5df8291c4645a2170f381b8d99b16bcd07e3aa906b138ff07c8565fc96afd

Observation 674a23ef-a859-4fa0-abfb-8c3dd55cf2d9 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.737053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.737053Z digest=sha256:fe3ff7fdb6e9608d0e22a543e1c62b09ecc5bcfc405a8d8a41499f789347081f

Observation 374273db-e653-4a63-945d-d745994c824d · outbound

This paper cites MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.808950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.808950Z digest=sha256:e8b91d004c5d6f761060732a137aa03a62467feb0d75f0beaa1c0a80b2a2969b

Observation 242fbcc3-9161-4688-add2-97f2000fa578 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.821194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.821194Z digest=sha256:2b74bb40212c077e3264a5b2df378577118eeabccc41575c2fcce70a00e833cf

Observation 48a225fb-5b1d-4cf3-8168-4a210622fcce · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal Chain-of-Thought Reasoning in Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.790349Z digest=sha256:31b780b5fc47427a8548779bbd409cd628639e0863cc6b3b9cb7d0fc55add52b

Observation 1e7286e3-27d0-49ca-b3d0-4040e9b2d71e · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.541576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.834814Z digest=sha256:b1dd6e0fc04f5e7702f9666088e0d3d75042eb8faac58ef50fb229035ccc4051

Observation 9afa9931-c062-408b-a212-7cdd30230c4c · outbound

This paper cites reasoning step.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark reasoning step

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.504739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.856925Z digest=sha256:417c8216ebe65abdd840debd536146c493f2f4a8971949fb8a7fea432e4e4958

Observation c25669a8-8996-44f7-be32-3c9ab4488f32 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.311098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.900673Z digest=sha256:b56a9f466a62f392f690f30059a68104218e23d2d06b671b09572c3a5aac352d

Observation 09ab0e9d-67e1-47e7-9269-464b8458d6a1 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.259112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.946132Z digest=sha256:66b7f45a7d631b3d440d77a67b7de1bbcc0c8ef885ed446d1f0f9b03da36a438

Observation 2cb14538-63db-429e-98ba-48aa6e23b6e2 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.465449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.961296Z digest=sha256:b3b32d8055bfe78d06c5cb945024c366b71dfb403441b41e05792abc5a691084

Observation c130aa28-c9d8-4f8c-93ac-78e6f1e19275 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.395288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.971376Z digest=sha256:8589411e784168df07798433d2154277c2b5c552aae48bf331bcc33d2e0dc565

Observation b7038d16-669d-460c-a5e0-4ef8f49abba8 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.363949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.980113Z digest=sha256:a0973f93aad0f70a9216c662c52f2145b34233da3e68dc23f5fa92f0470717b3

Observation 56aba37e-d952-4cb6-ad57-00930f296933 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.199841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.988591Z digest=sha256:dacb51372ce193d24c615c7166f394c6cbd6721796e4b7eaf20558d07b838f52

Observation d4e85cc4-f8a5-487e-b582-eee8aea80311 · outbound

This paper cites Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.116840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:20.013797Z digest=sha256:7862f43ac4305cf00b079dd700cfdffd818f839883a086662fe25e279c019454

Observation 78a0e081-f8d1-4af4-bf4f-192fa85a8888 · outbound

This paper cites Rank higher the responses that least misrepresent these relationships.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Rank higher the responses that least misrepresent these relationships

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.066928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:20.026221Z digest=sha256:1261f291aa3e1cdbce0cc0b806ccec907671e32a7ef630b93cbbdb7c9dd61da4

Observation 88f71158-6ef9-430c-88b4-7f57a2d86032 · outbound

This paper cites Responses should avoid inaccuracies in describing the characteristics of the objects present.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should avoid inaccuracies in describing the characteristics of the objects present

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.020746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:20.034083Z digest=sha256:a9056b7933da2ea38ffefc24ab93116bbd9885c627be71d532f5342e307cca22

Observation edc4972b-8872-4d0a-b8bc-92ad31259023 · outbound

This paper cites Equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Equally good

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:21.962642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:20.044416Z digest=sha256:8006428050a8829f3d8ded572dd201b41b442a85ab4f3f6fb6fd213be71cd1b1

Observation 7994fa18-e622-4d8b-b9ee-64d433200258 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.509807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.509807Z digest=sha256:d2735745f6330f954e3750eb35a89f7c42409c9c3de77d415a66887a3b4384a5

Observation 3823083d-0afc-468d-bd06-d4c92fc52bff · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.114146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.114146Z digest=sha256:43af02a565dfb6462ee54c579fc85688cdf46ea0fcfcbc0330ca657c9bdd3ac4

Observation c013ae08-676f-4166-8fbe-e7b9fac4652c · outbound

This paper cites InProceedings of the IEEE/CVF international conference on computer vision.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InProceedings of the IEEE/CVF international conference on computer vision

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:23.155612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:05:19.262204Z digest=sha256:be85786f5a5d0e021e919b5ad00a4007f0584416c4ad08e2f4af085f46529619

Observation 2760d9a4-0362-4c1d-a434-7a40e9ef31e0 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.193447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.193447Z digest=sha256:f0d44d8d58a0ee5eafcb55e362b6bd1fce8e627ee722eabec5d9466fd6d01fab

Pith citing papers

Observation 5c6864df-b047-4bee-bf80-fedc21632611 · inbound

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation cites this paper.

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:43.474266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:43.474266Z digest=sha256:1c575ea3b55e3db050f83ab195a937c9c53b4fc11817e31e3c538e88790444a1

Observation d334f9d5-9cff-48f6-91d2-3fba26525cc2 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:13:24.779717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:9923b7bf8f6d0fb134218b8466e1ec1598ca24345adab3dc53fba3d472e0617e

Observation 92fa0042-066d-4efb-aa58-7d57f2915315 · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.830119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:b0df72bb3881a19e3fd58ea39744571501bd9d95d2d07597064e3ed053216a63

Observation 8425e803-0d62-4c85-bd60-4fbd56364dac · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:02.750336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:3225f51868a4a481a233f7b9156ecf32659ee6d03f49a0ba868a2883f548439d

Observation a1c2534f-54a4-4485-8371-831cfd49b2d7 · inbound

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning cites this paper.

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.324575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:15:56.598968Z digest=sha256:de1bd944e4211dbe97eeffc1aae590eabcc76486e962cebec129288f3c97ac2f