Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2406.11833.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:35:45.546581Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T21:05:04.009666Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation abd895d6-bc80-4a86-baba-d8f1cd9a72bc · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation abff02ee-3b32-4204-bb6a-1b944bf634f0 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 235
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc0fb77b-b3fe-40c1-8da4-d8508bb2c756 · inbound
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b48b8d1-9ec8-44bc-98dc-14d5758f7f76 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 159
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a2f7b6dc-e188-4e9b-a2e5-41d75ff62a5b · inbound
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98dfef0f-a47d-4f1e-9919-aafa5b46d6cd · inbound
GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9b2023-4620-4bd9-958a-e7b2eaacdf17 · inbound
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed1e4452-2194-4c8a-b3c1-851cd6ed0d6a · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 146
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ab510f6-ba8f-4d3a-9464-a203b0137c7a · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cc460974-4449-48d9-81ff-b4efd56d5420 · inbound
Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a838b0bf-022e-4f1e-a280-44925a1c28ed · inbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09a076c8-e187-4eee-a58b-a509c02c1004 · inbound
Medical Large Vision Language Models with Multi-Image Visual Ability MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39bb3852-249a-4e9c-83a5-81c5995c78b5 · inbound
ImgEdit: A Unified Image Editing Dataset and Benchmark MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 09821de1-a06d-4d42-8d64-0d39800c32f7 · inbound
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5695cc77-06fc-4aca-a66e-e9a988521821 · inbound
CoMemo: LVLMs Need Image Context with Image Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8983ad3d-ffe6-423a-8e4c-056a0479e28c · inbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8da736-404d-457d-8ec5-ad7898f24989 · inbound
Docopilot: Improving Multimodal Models for Document-Level Understanding MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699acb2b-98e2-4393-8f77-c1e5e8648c94 · inbound
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e10f521-ca79-48b0-a676-b458bc36ebd3 · inbound
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation be3073c6-1515-4b5c-8202-190bdd5e2c93 · inbound
S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 16216f9a-7d11-4345-8490-85ecae3c69a4 · inbound
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.