Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2404.16006.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:00.710731Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:50:10.187396Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f07e53ec-0581-49c0-b1d4-c97c9b444c5e · inbound
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2203c694-e3c4-49ff-9ae5-f353746bb08b · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8cee43a7-b33a-4385-a8e3-8b2429966526 · inbound
Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1cf8737c-5e06-4ea3-a4af-69cf81b378a6 · inbound
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 64c8d833-788c-4e0b-8f5d-9d19661b3933 · inbound
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb0e1f1-6f62-4fc8-846b-6f1d09200c36 · inbound
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1ed913-0ec8-46fb-bc60-0014c828bf2a · inbound
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41f9a92b-b410-4268-9701-a2246ee958eb · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 278
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3f4071c6-3d95-48d8-91fc-6fcd2733ea34 · inbound
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b89acf0-6d5f-4b6d-9e8a-432c6991a3c8 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 94ee971f-2195-4d7b-ab91-d332cb320c83 · inbound
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c41d2a-8e4a-4132-bf31-981baf459a15 · inbound
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e128430-72d1-4109-8572-04960d51c3ba · inbound
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 122975f6-eed8-4776-8518-f9f0845635f9 · inbound
When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61ff7e0a-45e4-481a-840e-ac8901063e05 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3b940da0-3cc7-4cf0-a4d7-ebc30e705242 · inbound
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 156
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67fd94ee-9da5-4dbd-b54e-bf48ab174532 · inbound
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05d5a28d-78ef-46e4-87e7-9f579b2fc602 · inbound
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd1247fd-ca59-4c6e-8d84-ab8666766be7 · inbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0552af4a-974a-47da-b9ed-a3540cec5b83 · inbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 359cf8d7-ad72-45e5-8231-774b361010e1 · inbound
CoMemo: LVLMs Need Image Context with Image Memory MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b975bd57-1ac1-4c2c-9a4b-a220e9bdc4d1 · inbound
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e302ac2a-9625-464c-9a83-d3a02c2d75d8 · inbound
Taming Vision-Language Models for Medical Image Analysis: A Comprehensive Review MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 231
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea9b439-fd49-40a5-9db9-5e1f2a8f0b91 · inbound
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 194404aa-8c2c-4111-bd6d-595658e045ed · inbound
Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dfb38a4-4956-47d4-9875-0bea4f982ff9 · inbound
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b1b18b-a872-47ff-8655-2e2f8a63c8c1 · inbound
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a3a3dc-d9ba-4060-ba7c-8bd3e211a70c · inbound
MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623233a9-efc2-4d9b-826d-7f857dfd8e81 · inbound
WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7b3eee89-68d3-4e4b-8a3b-26b972ad55da · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cfdf7c3-1f09-4e39-8537-ae6415cfb2a5 · inbound
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d72ab86-6f58-411f-83cb-727847e57973 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d08a92f7-3106-4809-bb68-07ac28c4ed8c · inbound
Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bde5717-df73-4b65-8f18-d240ca650786 · inbound
SPHINX: A Synthetic Environment for Visual Perception and Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a7e25596-8751-4631-ad05-a1ae6b969843 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1e326ccb-4340-43a7-99de-9e432e5af207 · inbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e9f9dc04-1acd-416f-ac6a-8f94a592e8ba · inbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a854ec4-9137-4869-94f1-f5370a6ef96f · inbound
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2f087083-3736-4a58-96f2-e94383d27b47 · inbound
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62ba793-183d-4028-87e8-a70d8dab483c · inbound
HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 17f99df8-7dbb-49af-bc06-2fcabbc55bbd · inbound
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 739c674d-a30c-48f2-8f91-e0179c88281d · inbound
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 575ae2a5-bd24-453a-abba-a6d0cea6c8a8 · inbound
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dac1914c-acbc-439b-b6b9-538dd6b00fbd · inbound
DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ccd1fb11-eedc-46fc-a1d7-f1fed4e22ceb · inbound
TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 03d97ab5-429c-4341-9b82-2c8f44b93dad · inbound
C3-Bench: A Context-Aware Change Captioning Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 479473d3-2b3b-4790-aa44-1eceee4481a3 · inbound
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b5d014b-bf73-4425-b7b5-cf1bcdd2f25a · inbound
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a36d6b0f-1480-48ef-bb41-49ab19ee5561 · inbound
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51cafaa1-e006-4bd0-a2f0-c01b90c8e58b · inbound
Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs? MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.