Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.12742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12742 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:23.728013Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T00:02:50.478421Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6a0b118-e0be-4c3d-9bfa-3a2839f44dd0 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.202763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:3f1e2696ab63e8d40d485007a564f5693fff8975ebbb43f92d39ec79a1843ba4

Observation 00d1e608-90c7-4f33-8808-8973a85ceb53 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:23.728013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:23.728013Z digest=sha256:81cb5ec84fa4639a934d642b4c31d0e5b8b3ea78635f1730b1081611063e4f6c

Observation 8ba2266b-47ea-4432-87f9-5ebfd913950d · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:23.068331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:23.068331Z digest=sha256:d4a0c034df65e2b1edcd178f60d568f36c2bcb7cf83d3635f3c13c53f7f8fa7c

Observation 0980ce6d-691a-4e29-ba97-c736a7cd9863 · inbound

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models cites this paper.

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:52:14.967780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:52:14.967780Z digest=sha256:3280a2fdd31f32b42dbc630781cf5453e67afb61f218df21c07652b030c37ccf

Observation 92f13c2a-8aa9-4fef-8940-59b03860dc44 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 182

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.978903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:cc0b85882ed208103a1f78987a00b66f19534b860a83f2f8317a6e4e126bafe6

Observation aec1ebc3-760e-41fc-b7d4-a9a163d4f9d7 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.978503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.978503Z digest=sha256:10270a752787978ae8ab9ef3831824af5a02775078ff79c45f3bb13e18cfa512

Observation b2425791-05f2-48b0-b1f6-befefdef3029 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.146864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:a1b68b4b6bae75febba040fb14c16caf4d3778c9e435acdfa169c7db8745f9af

Observation bd96174b-9006-4678-a4c6-49ff1f505736 · inbound

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning cites this paper.

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.480515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:15:56.598968Z digest=sha256:6ae2d9066b78e1aa706ff811f87f5e5420a09eb777b2581840a6b957b8e00fec