Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.12742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12742 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:39:41.318268Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T00:02:50.478421Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d1ca39bc-2e83-4c11-a2d1-f441b52e5a69 · inbound

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens cites this paper.

Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:03:24.558647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:03:24.558647Z digest=sha256:cc8a72e7f65781b06dd6a9f0a16c9cc32ae40beb23b7fa39c012a3d5fa749551

Observation 16774ed2-75b7-446a-b504-a67a3e08feb2 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.005301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.005301Z digest=sha256:f2d3c958d3f0e197599b2694839cb00a2431ff99836d41f677cece4adf180dfc

Observation c6a0b118-e0be-4c3d-9bfa-3a2839f44dd0 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.202763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:cedb330f8e1f7ab92984a3a971c68fc3d2b8ed31f2b8612dca2aeb116ccf1840

Observation e83060c9-88e1-44e1-b50c-ecc175129d88 · inbound

Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains cites this paper.

Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:39:41.318268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:39:41.318268Z digest=sha256:d1360454a80de6ae419004de1d41ce02a8b3a57af68da092d268a8136070aaf7

Observation c998bba7-4b75-4325-9045-4da573f3b1b5 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:09.693716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:09.693716Z digest=sha256:0d34ff132daa4308195c24abdc9234c75debc47b18843fbd91aecaa60eef70c8

Observation 00d1e608-90c7-4f33-8808-8973a85ceb53 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:23.728013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:23.728013Z digest=sha256:90a099ce7744f8499341e7a07d61526db029a3f9fa5a944e7e6cc83d0937bd70

Observation 8ba2266b-47ea-4432-87f9-5ebfd913950d · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:23.068331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:23.068331Z digest=sha256:4fb6685f06b2ae76ea7a7c0c4825580ec883defe481a14de7ceb0f5c3e47a2c7

Observation 0980ce6d-691a-4e29-ba97-c736a7cd9863 · inbound

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models cites this paper.

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:52:14.967780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:52:14.967780Z digest=sha256:ef1ec534ead7ae2b498fe8d0fed3511cd410fc5ee51fbb15b2e581b1bbc1fde5

Observation 92f13c2a-8aa9-4fef-8940-59b03860dc44 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 182

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.978903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:d1bfa92d43bdf8a2419220d03c2030c5ffb43fdfd53f13b0db24b2256e24e610

Observation aec1ebc3-760e-41fc-b7d4-a9a163d4f9d7 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.978503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.978503Z digest=sha256:a78f511a5b12c82a10cb78b1fe2e19e221f7d9d01c7545b6ea672f1d32566be4

Observation b2425791-05f2-48b0-b1f6-befefdef3029 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.146864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:36707d786bc9f4ac5ee29108fff2edd763d38e43d8264fcbb49aa5be0f48cdab

Observation bd96174b-9006-4678-a4c6-49ff1f505736 · inbound

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning cites this paper.

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.480515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T23:15:56.598968Z digest=sha256:3598e138342c8f66fcacbda66211dc7b3cea1c5bc4c6c3b5cea7b1323213c163