Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.11985.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:58.185565Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
5
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 2c0cffd4-9796-468c-9105-a91c41822bbc · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0bc929f0-7cd8-4d2c-a166-8b4f02b75fa7 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 228
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38269d09-3bca-4cbd-8bb8-73503b83363f · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61cd1887-8c50-4ebe-85e3-3ea3c9b7988c · inbound
Qwen2.5-VL Technical Report MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b33f1460-4f7d-43ab-8b0d-7e224e99433d · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03f0d185-6e7d-4775-a7bf-4f27119e219c · inbound
Multilingual Multimodal Software Developer for Code Generation MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddae192-bd41-49eb-aed9-21f4a513d996 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fa77b7a4-1134-4c0f-9ab0-63b9d36b2162 · inbound
LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 34fb4af9-5580-4886-ad1e-d1a63464aaad · inbound
Multilingual Vision-Language Models, A Survey MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f3b2e45-ec2a-407a-b6f0-eaddc4eb220f · inbound
LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111207c7-2583-48df-907a-babeb134fe9c · inbound
Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf8efb61-fa51-4edd-884f-b7ca1d56c4e0 · inbound
Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cfb4d057-5666-4f78-a646-a46b32205dc3 · inbound
Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d132e84b-6c28-4475-8e41-eb619ea4a887 · inbound
Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed168971-597e-4e1d-b4c1-736851e055d9 · inbound
Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f45e8b3e-a9b3-4e73-9033-66b690564a28 · inbound
DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.