Pith. sign in

Paper Citation Record · LEDGER

MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.11833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11833 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:11.997396Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.009666Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation abd895d6-bc80-4a86-baba-d8f1cd9a72bc · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.787020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:ac5e998a3c211b171c9d1fdeed0999bb15268cd3ea763c002d3ea0d3f1831abc

Observation abff02ee-3b32-4204-bb6a-1b944bf634f0 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.472394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:9adc21e594ffeb7c5f7b7cef7f77e10525b0cc3ed4ef9f65492ef0e003ea1923

Observation 2b48b8d1-9ec8-44bc-98dc-14d5758f7f76 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.741943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:d59866fa7127bbff9116ea191620a43754d15defaa6278dd12725884d4a46155

Observation 1ab510f6-ba8f-4d3a-9464-a203b0137c7a · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.349112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:aecca1e3bb0ccbc2f6e5799c1593b1311768619325dd740977bd4586971122f7

Observation 09a076c8-e187-4eee-a58b-a509c02c1004 · inbound

Medical Large Vision Language Models with Multi-Image Visual Ability cites this paper.

Medical Large Vision Language Models with Multi-Image Visual Ability MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:11.997396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:11.997396Z digest=sha256:5b3119d89a92ceed74145af34fb45946f3fe9f9464b91c5f58e481650aa57bf6

Observation 39bb3852-249a-4e9c-83a5-81c5995c78b5 · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.308787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:4f24a2a494c7ef419b1db6830ca2d0672caf02737777dcb421e294276d31c807

Observation 09821de1-a06d-4d42-8d64-0d39800c32f7 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:06.523108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:06.523108Z digest=sha256:bf1aaea5954c4e08c1faaa1b8395ccafed3562a7705194ac7f69c31483790fd6

Observation 5695cc77-06fc-4aca-a66e-e9a988521821 · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.346839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.346839Z digest=sha256:008a44cb6a989097e83fea0ed13fece3484e0b5a289540dae345a3d4abaa45e0

Observation 8a8da736-404d-457d-8ec5-ad7898f24989 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.873449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.873449Z digest=sha256:97490c9b0fd3f36bd93a3406bc6712a618b72d7e337fdccaec86b98d083d8108

Observation 699acb2b-98e2-4393-8f77-c1e5e8648c94 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:23.165410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:23.165410Z digest=sha256:cd047931221b6793d9830534f9323dec48ddcac2b5f18af8567fe13f7f964f3d

Observation 5e10f521-ca79-48b0-a676-b458bc36ebd3 · inbound

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents cites this paper.

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.373649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:31:15.525331Z digest=sha256:ebba14d2202ea42440594d7b20ac392a121e0e230d71907a3ef880a82b56460b

Observation be3073c6-1515-4b5c-8202-190bdd5e2c93 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.269872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:2f0b579d4114dbe8d8fbc9bf73de650e96018c9d91baeda07164a43cbc7ce8b0

Observation 16216f9a-7d11-4345-8490-85ecae3c69a4 · inbound

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory cites this paper.

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.012795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:02:58.640918Z digest=sha256:777b3f1385a805169a4b7cef129ac300445b9d29fab4c9cd6c26abd96f519838