Pith. sign in

Paper Citation Record · LEDGER

MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2406.11833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11833 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:35:45.546581Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.009666Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation abd895d6-bc80-4a86-baba-d8f1cd9a72bc · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.787020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:a672b275ef5655af745efaecba1c5863141a3d19dc8e0223ec43f0939ac410aa

Observation abff02ee-3b32-4204-bb6a-1b944bf634f0 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.472394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:d969079800a0b5f9297b5145d299159d83bd8f4c6384780508fad475242ab333

Observation dc0fb77b-b3fe-40c1-8da4-d8508bb2c756 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:36.987942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:36.987942Z digest=sha256:d7fdcd15d684ef95ff540f8460dd176758c366f2582fd216dd3ec5405c8be669

Observation 2b48b8d1-9ec8-44bc-98dc-14d5758f7f76 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.741943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:c426c4eadc350a23a11e7f9e001b504dd28819d27857ad8c87f96dc08f68cffa

Observation a2f7b6dc-e188-4e9b-a2e5-41d75ff62a5b · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.524846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.524846Z digest=sha256:eed3519bd09847fbfca8f41b942552ab692328f8bef644c438993f1029c5edaf

Observation 98dfef0f-a47d-4f1e-9919-aafa5b46d6cd · inbound

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models cites this paper.

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:57.130337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:52:57.130337Z digest=sha256:311b8c451215c70c5db6ef85171f6902592b146547c09057c9476d72ce88fbe0

Observation 9c9b2023-4620-4bd9-958a-e7b2eaacdf17 · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.199448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.199448Z digest=sha256:bd53476fde5539dc442559cec12c28179a0b23bf27406ce1025cae9a129ed98f

Observation ed1e4452-2194-4c8a-b3c1-851cd6ed0d6a · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.608140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.608140Z digest=sha256:e86899b2bc117f985026fc87fd2724776036ef7bc4e7cea005b69e69d680b930

Observation 1ab510f6-ba8f-4d3a-9464-a203b0137c7a · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.349112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:734f1af5ad8f0c3e7baa1049527f35537c22c98375e4a6905747904c23e05ee8

Observation cc460974-4449-48d9-81ff-b4efd56d5420 · inbound

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput cites this paper.

Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:35:45.546581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:35:45.546581Z digest=sha256:4b9048c3c438afa74fc04874581cb0789e43ce67e87d3002b2dcd225ad74ea87

Observation a838b0bf-022e-4f1e-a280-44925a1c28ed · inbound

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis cites this paper.

Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:13:37.587409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:13:37.587409Z digest=sha256:04cefc0c54a72295802e826bbf028b1da65041f1fbf19ef9109b288ba79668e2

Observation 09a076c8-e187-4eee-a58b-a509c02c1004 · inbound

Medical Large Vision Language Models with Multi-Image Visual Ability cites this paper.

Medical Large Vision Language Models with Multi-Image Visual Ability MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:11.997396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:11.997396Z digest=sha256:36de233fcf0c592c146b7bb1065b9d3cdf26ac249c6860f9c3929dbcfad78123

Observation 39bb3852-249a-4e9c-83a5-81c5995c78b5 · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.308787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:b45145aba2a6cdc786b631e4b2ce107d3b8231f63145d3125163da8af0134595

Observation 09821de1-a06d-4d42-8d64-0d39800c32f7 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:06.523108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:06.523108Z digest=sha256:fb3eae67de3f5f1c0d28767c0387373237b07c01ddb07a39b0f3847ef05644de

Observation 5695cc77-06fc-4aca-a66e-e9a988521821 · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.346839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.346839Z digest=sha256:1c6234ffded56bbe3497a9df0ec78419ec67807bf5c4694b0d44dd535f8f6700

Observation 8983ad3d-ffe6-423a-8e4c-056a0479e28c · inbound

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing cites this paper.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.496411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.496411Z digest=sha256:dc7ffcbe7ba85f488a37d88b0b42daad73637968708ced157e5deb6cffd30f3c

Observation 8a8da736-404d-457d-8ec5-ad7898f24989 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.873449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.873449Z digest=sha256:5c9a599a7105a2614ffc265d114d0c8d465f8a377d3fe220f38a34e8c2957459

Observation 699acb2b-98e2-4393-8f77-c1e5e8648c94 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:23.165410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:23.165410Z digest=sha256:286e5597ded685448f715c16a1198e14869909d0be870d5b5695f8a797b4fa29

Observation 5e10f521-ca79-48b0-a676-b458bc36ebd3 · inbound

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents cites this paper.

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.373649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:31:15.525331Z digest=sha256:891d3dfa4ddced1c8a9c7ad8c4b90c67f6e6562440d88c64be7f9e378271689e

Observation be3073c6-1515-4b5c-8202-190bdd5e2c93 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.269872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:6780bd7b2bca193d41e8db370e152c8db0fd4a9508ac1a8a356150409aee1e37

Observation 16216f9a-7d11-4345-8490-85ecae3c69a4 · inbound

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory cites this paper.

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.012795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T21:02:58.640918Z digest=sha256:f15026cbf388ca7d3d83c050f85c58b0b5a2deb34f37550517f58bdd18d73a73