Pith. sign in

Paper Citation Record · LEDGER

MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.11985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.11985 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:58.185565Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c0cffd4-9796-468c-9105-a91c41822bbc · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.703331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:812863fb18477591fe4fa7ce8e89d2c45d26d21229e764a1eb83dcd59a25d0c4

Observation 0bc929f0-7cd8-4d2c-a166-8b4f02b75fa7 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 228

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.036152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:40bbf1bf51ec3edc0977ffbd63f7214a64df099121ecbbd86ad0df8d13c237a8

Observation 38269d09-3bca-4cbd-8bb8-73503b83363f · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.122658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:97d7987a9350e7cad39b7aa9c45b262f36a326a8b0e98fd9c5a71d48779bdf0b

Observation 61cd1887-8c50-4ebe-85e3-3ea3c9b7988c · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.101037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:fcfb3cab8ccd6acb7e6e13738c62efdba55f74b7e3d927681c99cadf8f9c2e4b

Observation b33f1460-4f7d-43ab-8b0d-7e224e99433d · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.096968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:f91552b86cfea3b1e65107d83be6a2b424cbc28327531744f19dcbe822362fbd

Observation 03f0d185-6e7d-4775-a7bf-4f27119e219c · inbound

Multilingual Multimodal Software Developer for Code Generation cites this paper.

Multilingual Multimodal Software Developer for Code Generation MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.185565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.185565Z digest=sha256:31d99b1e72184fd0931f8634f089aaef523838b9484730de86a29a895ba7733c

Observation 4ddae192-bd41-49eb-aed9-21f4a513d996 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.775085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:829ed61d5fdd0693f58603c47e7c3e3ff33885ed5503e2408ca40fb944f5bb25

Observation fa77b7a4-1134-4c0f-9ab0-63b9d36b2162 · inbound

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA cites this paper.

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:56:42.310528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T17:53:08.136677Z digest=sha256:24aa84644d73f0ad5eded7acdd920917ba34fa770779b77f40849cdc16659148

Observation 34fb4af9-5580-4886-ad1e-d1a63464aaad · inbound

Multilingual Vision-Language Models, A Survey cites this paper.

Multilingual Vision-Language Models, A Survey MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:02:37.275213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T13:02:08.000814Z digest=sha256:448b8e0c706897ba384dece3f5e30384d8925d6376d0c4f50beb9a1044aea168

Observation 9f3b2e45-ec2a-407a-b6f0-eaddc4eb220f · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.130815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.130815Z digest=sha256:e5735ad56b927ba74c35b071c260224543efafd75e39d9ad05e0fc8af57fbd64

Observation 111207c7-2583-48df-907a-babeb134fe9c · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.114265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.114265Z digest=sha256:eeaad8a5cef4845a97b073cf5a4243bba814b2f8cc1529f04d674769c22e979f

Observation cf8efb61-fa51-4edd-884f-b7ca1d56c4e0 · inbound

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction cites this paper.

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:03:12.256459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:01:57.029143Z digest=sha256:c90e61ad74139dfa311c54169290178f33827475711e61b3c81461e45f1aa8fc

Observation cfb4d057-5666-4f78-a646-a46b32205dc3 · inbound

Multilingual Training and Evaluation Resources for Vision-Language Models cites this paper.

Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:38:43.165665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T05:11:00.461167Z digest=sha256:7e01257eb4e6356d486b4d6aef59214fc36f8b7d5d8a05937bc6c36cddc4dc29

Observation d132e84b-6c28-4475-8e41-eb619ea4a887 · inbound

Multilingual Training and Evaluation Resources for Vision-Language Models cites this paper.

Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.869841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:8008ec0d5a0d15492e3f2c61efcb2fc8acc60a92d8f352f058be99ba249d324a

Observation ed168971-597e-4e1d-b4c1-736851e055d9 · inbound

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting cites this paper.

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.576013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:28:51.711892Z digest=sha256:df47e8fd8e7508ea0f342dd894e3676f9880ebf8e5407000560cd2bc339ffb4a

Observation f45e8b3e-a9b3-4e73-9033-66b690564a28 · inbound

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning cites this paper.

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.359361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T06:01:20.803078Z digest=sha256:9b35c9065aa262b39840147d8bbc380a71a5a4d13f29f948f06c48691d62170d