Pith. sign in

Paper Citation Record · LEDGER

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2505.24120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24120 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:41:03.076855Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:03.671614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T05:55:31.163509Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7584910-2c37-4e9a-b185-661879891cde · outbound

This paper cites Gpt-4 technical report, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gpt-4 technical report, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.070452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.070452Z digest=sha256:9c74af5c5838a5bc24461e5f7ededffcc87870d7e0fd3acaf94d8f080470b28f

Observation c7f311c7-8fad-4b1b-8402-395389ae472a · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.813861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:40:59.118262Z digest=sha256:696cdf25503803d024b20e103acd4ae34e4651cc9cae0af75c7da249fedf8f97

Observation 89ecddcb-d994-480e-8a5b-2eabb2e6abc4 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.703189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:40:59.202715Z digest=sha256:487c5538fb6d9add06f01178d6375f713ba68b1aafe62e74380defaa886d0fd2

Observation 94fc0825-f2b2-4750-b2e2-1a5f9fdb828f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.273403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.273403Z digest=sha256:33e4f5d997a2b2a29ca784e8d39819d6ec174de2f474579a05fed4268bfc81f4

Observation 30fdd865-079d-4720-a636-45aff8245f09 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:59.348853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:59.348853Z digest=sha256:aaf7ca14adb7c639985fede3ef3ac2d869fc280d0d29e252c18c34d56a209dc3

Observation c7d9aefb-fa90-4664-b2b6-5d030929f230 · outbound

This paper cites Kimi k1.5: Scaling reinforcement learning with llms, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Kimi k1.5: Scaling reinforcement learning with llms, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.582667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:40:59.418298Z digest=sha256:d5143399b04a6c041f243b5f4ccc5104fa7924558ce5f24a5275b229bb8a76ea

Observation 161e185d-2897-4d8b-a607-9c89c2732e88 · outbound

This paper cites Gemini: A family of highly capable multimodal models, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gemini: A family of highly capable multimodal models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.441409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:40:59.496657Z digest=sha256:237d52dd0fa9675c6192a75dbe6c352b2fe673eda6906d7610981f2d7d2913da

Observation 324b1f69-6911-4cc9-b3fa-04e331ba0a3d · outbound

This paper cites Hello gpt-4o, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Hello gpt-4o, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.301657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:40:59.588916Z digest=sha256:dcc6d5f295bee73c3458cb53020adb6e19a7e25c10e0b387d89bcdc70cb492a7

Observation 4b54aee7-b2fb-4eac-b9a5-a85a64adee3d · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:09.154599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:40:59.657293Z digest=sha256:274bf8ab807460c09d69f66aada1ee7d927228931ea8252c56e330a86a074160

Observation 3b68a956-1007-41cf-af05-323095274103 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:09.028912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:40:59.737185Z digest=sha256:b59104a45d7376b18d9c3e9a0d87e6561c8464a8457ccdf5645149a085eafe7b

Observation 691bd1e8-04a8-46f7-b709-c11d53c7ca57 · outbound

This paper cites V Jawahar.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs V Jawahar

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.888367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:40:59.885061Z digest=sha256:54b522e660f78071a98639ce10a0adc23b842dd08767059370e2353f8df0ed78

Observation 5ab11a5d-6aa8-4048-8acc-13bb2133af01 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.731263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.011399Z digest=sha256:dcbac805434edeb43fcd6a57bc1d407a14bcf869a801c880fb7d261f3f0246b5

Observation 3c39f7e8-b013-4221-8c2f-f992ec631666 · outbound

This paper cites Lxmert: Learning cross-modality encoder representations from transformers, 2019.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Lxmert: Learning cross-modality encoder representations from transformers, 2019

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.599235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.089071Z digest=sha256:72dc6a279a7bf839813bdc8455db3f0f1bde500189c094830210e02ca6f22fb1

Observation e6834180-8224-4b54-951f-a16af1736e98 · outbound

This paper cites Uniter: Universal image-text representation learning, 2020.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Uniter: Universal image-text representation learning, 2020

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.441957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.145387Z digest=sha256:245ff2f8c427c61b87def589210adf2e3e86f528d5ea182ee271efd89c3a257e

Observation edf10535-f2d9-4a8d-af74-9a019feafb9a · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Learning transferable visual models from natural language supervision, 2021

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:00.211295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:00.211295Z digest=sha256:4fbff92b787652729728a46dfd097c42300ca08774cbc30fe7d759d09c7fe5bf

Observation e307d979-e9d1-4bfc-a8de-6c5589e7d12e · outbound

This paper cites Le, Yunhsuan Sung, Zhen Li, and Tom Duerig.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Le, Yunhsuan Sung, Zhen Li, and Tom Duerig

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.262751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.304151Z digest=sha256:565cf5790d64f34e5a7fde3aa0076098f58f8adbd535864710324a88d867dd8a

Observation 0165464b-00b6-4e8e-a48f-96a31ff68147 · outbound

This paper cites Evev2: Improved baselines for encoder-free vision-language models, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Evev2: Improved baselines for encoder-free vision-language models, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:08.072117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.398803Z digest=sha256:fcf7ccf9719bc745753e1281c8e2b920131ef51ff9fcc6d10033c0a85a39ea9c

Observation e2ff3864-8d3d-4561-aa19-d46c5f3a310c · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.866112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.520998Z digest=sha256:ab3c80ebd111ab510e26e7f55136cb83c59c942523a71d3a8f435b65635e5945

Observation 95b64913-819f-4659-ad65-d74794211a08 · outbound

This paper cites Introducing Gemini 2.0: Our New AI Model for the Agentic Era.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Introducing Gemini 2.0: Our New AI Model for the Agentic Era

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.721855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.644759Z digest=sha256:7ceac52d3a519a274bac586ba8b9ca9de8b32aa1eba3697fc158db179dc9ceb5

Observation bdf30c32-9d80-4129-8d5f-a9364fd3e3b0 · outbound

This paper cites Gpt-4o system card, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gpt-4o system card, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.594072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.733515Z digest=sha256:4b96aa709a8d96b2eb3aa6f9e6123a864e0cdd046fae224b86743f57161aa628

Observation ab5f2dc8-7eb7-4aff-80dd-9076505a8771 · outbound

This paper cites A diagram is worth a dozen images, 2016.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs A diagram is worth a dozen images, 2016

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.426138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.823620Z digest=sha256:de3c3b96b19fa5b37e6e4d70f1f0755c12618c7bc35e99d7eebad23e8189e6c9

Observation d901f0a3-432b-4f3b-aebf-19d4b261882e · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Ocr-vqa: Visual question answering by reading text in images

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.206904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:00.934581Z digest=sha256:d3c37ebbdf53482db92c6bb61b6d73d25443a5cc912f569236688851508c95ed

Observation 3f998ecc-864c-4531-a960-a465f0fd37ed · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:07.060454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.029385Z digest=sha256:0f1ade8f09c5900e704c4c70001c29880a05779f9a7a94bd0f9f44c6f5d35018

Observation 08c6473d-8e3c-4986-b1da-3893616464c5 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.897367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.115297Z digest=sha256:d35a3fec14e17b86da9fbf03dda2bb7240abab6b87875aa8f2d5ed48f7630876

Observation bf849b6d-5d5e-4054-a099-03bda40f6d56 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.764386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.212066Z digest=sha256:49ca4d2ddbd311b30bc1f07f5512c41936f0043c5531d795ba270c16e0a726a0

Observation df31dd7b-4de3-468f-8f15-7c3087c86cca · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:01.334554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:01.334554Z digest=sha256:bb4c0de273d769994fb739fd5880a7ccc5713e0e97aeec720fc3d358ed5013fd

Observation 7df63926-5e1e-41ba-b9f6-9a9c46e755e6 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.610441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.400126Z digest=sha256:b1c2acd083b82bd4ff6935625f99709a70b57964ddfc2384cdf900891fe0d878

Observation 68aa9c35-5de9-4e43-9aee-38970d0c5ab7 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Measuring multimodal mathematical reasoning with math-vision dataset, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.468702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.467755Z digest=sha256:0429fcf8aabfaed3a20f3f0adcc1bba3d0a5a1c9813601f8bcc65f90d53a97a1

Observation 6b2e37e9-dc90-4229-b1fd-7aba4120cf33 · outbound

This paper cites Claude-3.7, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Claude-3.7, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.326033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.540750Z digest=sha256:c313a964d24dc7e239412ae59fc1a51f46fd10691ec5400cceb19ac6aae16322

Observation 510321dc-36d3-4e2a-9f78-94bb6aa88056 · outbound

This paper cites Qwen2.5 technical report, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qwen2.5 technical report, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.178864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.641120Z digest=sha256:609b48ff11fa33a9b97ac6c5d108e9ca30461db6d7ce7fbae7aba168a61ce466

Observation 170d1a4f-c702-4868-bd5f-03b805dbf84e · outbound

This paper cites Mineru: An open-source solution for precise document content extraction, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Mineru: An open-source solution for precise document content extraction, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:06.014990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.733700Z digest=sha256:657133fa2f396a089cf86998c201fae9ac77044708e3fcc28f194125e5639203

Observation 4c60abfe-3b64-409d-8ce5-e5c6e6f7704e · outbound

This paper cites Deepseek-v3 technical report.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-v3 technical report

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.843258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.833051Z digest=sha256:4b0288f8a9bbfea40db9c37fe42ee9f8b10962b44e96e8a904802b3f379333df

Observation fef5838e-27b0-4f96-a1fc-0e97b6ef2156 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.706057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.909165Z digest=sha256:08d94ad7773af3f5fe2aadbedc2a4243134b67147c4a598cd69ab3ab214d93c2

Observation 85c4e7d7-c828-42a9-822b-530d57871178 · outbound

This paper cites Introducing our multimodal models, 2023.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Introducing our multimodal models, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.449469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:01.994475Z digest=sha256:130027ebe6342abcc205e587fa7a3408bb6ed5c4096f7c904ef318d0710257ae

Observation b1515deb-5dbb-41e9-aa5b-4c50b36a65f6 · outbound

This paper cites Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:05.254171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.078331Z digest=sha256:66aba36fc5389934b777c164ac46bc913de02dfc0780654342237669ad1670c7

Observation 288caf5c-3887-43dc-a5c0-798a40f0870f · outbound

This paper cites Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:02.241977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:02.241977Z digest=sha256:e6ca4782cad797798080f552f3595c54f057c3db441ae86008c68571a956274a

Observation 132f83f7-0961-4137-8bbb-9dbd4ccda514 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:05.072111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.315018Z digest=sha256:9731cd8397bd296cc06d666dbab586b4ab14cde9ae1cc549e4dacc6052ef8c66

Observation e3752dbd-0ef9-45a2-9bbd-02c598537319 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.921260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.379786Z digest=sha256:bb0252d35d2da6e0a77baf7c07cec3b3a150ac314cf795dd6f45fdda8d30d3f3

Observation 0533e597-c71e-438c-af5f-70a61bf93df6 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Building and better understanding vision-language models: insights and future directions, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.758949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.482049Z digest=sha256:acb8b0781462f2b8f6e610c7fe16258f6671d66962593c27d76cba1cb4267457

Observation dba9b979-683c-4a3f-baa4-df085e00bf06 · outbound

This paper cites Improved baselines with visual instruction tuning, 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Improved baselines with visual instruction tuning, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.535501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.594606Z digest=sha256:828391170189b2921193a7950ba7159951580d41648378cf979801ed0c68c3a4

Observation 34071225-9ce5-4881-8e4e-cdaae9a96577 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:04.279722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.668534Z digest=sha256:cb004f597b2683ae2499cb9d4f27e2e43c21c2de552e751cab7df12c1df6fa88

Observation f8920688-1ab2-4d0b-a3c3-4934be1addb0 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Qvq: To see the world with wisdom, December 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:04.124055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.742793Z digest=sha256:546e123bc564cc4e6a9419bf9e467fa4ed4966624d956be5aea4ff5e84f39a0b

Observation eb4b8d66-110c-4d78-81af-ba890d089ee8 · outbound

This paper cites So the final answer is \boxed.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs So the final answer is \boxed

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:41:03.865748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.858561Z digest=sha256:3ce5262ac11059e1a76f1791247f055e8e95ac902dc861e19d95e599f456b055

Observation 1724ab91-8c2a-4e98-b23e-3b69f20ed704 · outbound

This paper cites an unresolved cited work.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:41:03.646708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:02.974741Z digest=sha256:e5d8e20cc9c83b082ff669d3bb7c91b29544d1f69b206c8f335338028f51b0b1

Observation 8fd73f73-ea3a-4c2e-8e8b-760cf06a0ba5 · outbound

This paper cites No," please identify the main unreasonable aspects or obvious flaws in the solution; if.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs No," please identify the main unreasonable aspects or obvious flaws in the solution; if

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:41:03.400314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:41:03.076855Z digest=sha256:f4a0baee5c6e38b454dd4330707276c9aba1fcba2aea6ca7c338ff0c0a2b8a7f

Pith citing papers

Observation 135949bd-3881-4ccf-9c25-8f809311f889 · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:03.671614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:03.671614Z digest=sha256:189d8f70cd196535ca0484616531a391595b47f5212588d0fe09f4a283c25cdc

Observation 429530aa-4830-4ffc-a5c8-fdadb6c92953 · inbound

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning cites this paper.

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.166191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:21:30.235583Z digest=sha256:9abe8bfe06bb9fe0b7524af67374e0a942dc0ee307f2945547dfb09d3663726a