Pith. sign in

Paper Citation Record · LEDGER

MIBench: Evaluating Multimodal Large Language Models over Multiple Images

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2407.15272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.15272 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:34.466045Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:20:36.463031Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bddcacd3-b19a-4486-bce2-0921f67279f1 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 233

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.464952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:bdc50493c5674fe2b0d1fd2c4a38a26ad5fad024dd21050dfa382717065e8992

Observation 3148e369-a2de-4a8f-aab3-1ace73c11c2d · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.466045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.466045Z digest=sha256:6abb9bd337e7ebdaca36434f2d98c34bdeba96845408bd7a76d2a1f7a3c9b0d1

Observation 90dc752f-6c24-4de5-932f-02fdb89d9861 · inbound

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding cites this paper.

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:29.679885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:29.679885Z digest=sha256:0506fb6f6a90599da1b4179edea3bbe243c9b84465f11ff21063a2a7320852f7

Observation ca0dc0d1-1b8c-4076-b211-6d6fc480ed29 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.176871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:6859cb838b06bae2b109980bb66928fa63a9c954ab4a5e71b3909d5af5a6592d

Observation 4543bf77-6c6c-41a9-b0f3-ac94b6292f88 · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.526702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:f307f57a850024a19b123844fe0d3748f807aec21ac4ec4ab7b754f3173eba4a

Observation 8afc926f-646d-4ecc-867d-b7266e2a1970 · inbound

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration cites this paper.

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:01.703230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:26:45.182597Z digest=sha256:ac4ed253d1560bb8d16c3389a6b40aef415b7931229553ffa1665c43de23d160