Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:41:03.076855Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2505.24120.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:41:03.076855Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:03.671614Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-09T05:55:31.163509Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a7584910-2c37-4e9a-b185-661879891cde · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gpt-4 technical report, 2023
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f311c7-8fad-4b1b-8402-395389ae472a · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Llama 2: Open foundation and fine-tuned chat models, 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89ecddcb-d994-480e-8a5b-2eabb2e6abc4 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94fc0825-f2b2-4750-b2e2-1a5f9fdb828f · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30fdd865-079d-4720-a636-45aff8245f09 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d9aefb-fa90-4664-b2b6-5d030929f230 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Kimi k1.5: Scaling reinforcement learning with llms, 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 161e185d-2897-4d8b-a607-9c89c2732e88 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gemini: A family of highly capable multimodal models, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 324b1f69-6911-4cc9-b3fa-04e331ba0a3d · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Hello gpt-4o, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b54aee7-b2fb-4eac-b9a5-a85a64adee3d · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b68a956-1007-41cf-af05-323095274103 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 691bd1e8-04a8-46f7-b709-c11d53c7ca57 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs V Jawahar
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ab11a5d-6aa8-4048-8acc-13bb2133af01 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mmbench: Is your multi-modal model an all-around player?, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c39f7e8-b013-4221-8c2f-f992ec631666 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Lxmert: Learning cross-modality encoder representations from transformers, 2019
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6834180-8224-4b54-951f-a16af1736e98 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Uniter: Universal image-text representation learning, 2020
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edf10535-f2d9-4a8d-af74-9a019feafb9a · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Learning transferable visual models from natural language supervision, 2021
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e307d979-e9d1-4bfc-a8de-6c5589e7d12e · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Le, Yunhsuan Sung, Zhen Li, and Tom Duerig
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0165464b-00b6-4e8e-a48f-96a31ff68147 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Evev2: Improved baselines for encoder-free vision-language models, 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2ff3864-8d3d-4561-aa19-d46c5f3a310c · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95b64913-819f-4659-ad65-d74794211a08 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Introducing Gemini 2.0: Our New AI Model for the Agentic Era
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bdf30c32-9d80-4129-8d5f-a9364fd3e3b0 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gpt-4o system card, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab5f2dc8-7eb7-4aff-80dd-9076505a8771 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs A diagram is worth a dozen images, 2016
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d901f0a3-432b-4f3b-aebf-19d4b261882e · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Ocr-vqa: Visual question answering by reading text in images
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f998ecc-864c-4531-a960-a465f0fd37ed · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08c6473d-8e3c-4986-b1da-3893616464c5 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf849b6d-5d5e-4054-a099-03bda40f6d56 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df31dd7b-4de3-468f-8f15-7c3087c86cca · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df63926-5e1e-41ba-b9f6-9a9c46e755e6 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68aa9c35-5de9-4e43-9aee-38970d0c5ab7 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Measuring multimodal mathematical reasoning with math-vision dataset, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b2e37e9-dc90-4229-b1fd-7aba4120cf33 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Claude-3.7, 2025
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 510321dc-36d3-4e2a-9f78-94bb6aa88056 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qwen2.5 technical report, 2025
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 170d1a4f-c702-4868-bd5f-03b805dbf84e · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mineru: An open-source solution for precise document content extraction, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c60abfe-3b64-409d-8ce5-e5c6e6f7704e · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-v3 technical report
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fef5838e-27b0-4f96-a1fc-0e97b6ef2156 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gonzalez, Hao Zhang, and Ion Stoica
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85c4e7d7-c828-42a9-822b-530d57871178 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Introducing our multimodal models, 2023
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1515deb-5dbb-41e9-aa5b-4c50b36a65f6 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 288caf5c-3887-43dc-a5c0-798a40f0870f · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132f83f7-0961-4137-8bbb-9dbd4ccda514 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3752dbd-0ef9-45a2-9bbd-02c598537319 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0533e597-c71e-438c-af5f-70a61bf93df6 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Building and better understanding vision-language models: insights and future directions, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dba9b979-683c-4a3f-baa4-df085e00bf06 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Improved baselines with visual instruction tuning, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34071225-9ce5-4881-8e4e-cdaae9a96577 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8920688-1ab2-4d0b-a3c3-4934be1addb0 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qvq: To see the world with wisdom, December 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb4b8d66-110c-4d78-81af-ba890d089ee8 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs So the final answer is \boxed
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1724ab91-8c2a-4e98-b23e-3b69f20ed704 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fd73f73-ea3a-4c2e-8e8b-760cf06a0ba5 · outbound
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs No," please identify the main unreasonable aspects or obvious flaws in the solution; if
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 135949bd-3881-4ccf-9c25-8f809311f889 · inbound
Skywork-R1V3 Technical Report CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 429530aa-4830-4ffc-a5c8-fdadb6c92953 · inbound
Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.