Pith. sign in

Paper Citation Record · LEDGER

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

As of 17 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 4 inbound Pith citation observations for arXiv:2508.12680.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12680 v1

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:24:08.529010Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:09:33.968186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:06:10.238711Z

Reference resolution

100 of 113 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfd5f7f5-cedb-4235-b24c-3b2b79e6a404 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.025253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.025253Z digest=sha256:6a43d825abae177de7d8c4422caef8ee028c9235b18d40e58d2d50a6f4d78192

Observation ce8eff9e-56fc-4499-ad9d-9ffb88505115 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.031605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.031605Z digest=sha256:0891f88f3c91e438891f7976d739fd396cc224fcc408400b919b326da1593f20

Observation 6aa7cf1a-2207-493e-b21f-e67f780a5272 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Measuring mathematical problem solving with the math dataset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.037031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.037031Z digest=sha256:9e3d74856047589705fb556bc6989bf9f49df497b398b29e9e66dd07cb3a589c

Observation 4a2cf671-3e03-4208-86e5-e1bf28c41528 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.042129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.042129Z digest=sha256:56336823d1869881c966d395b30c8478774ea46d28ed69231ea39eaa1b6d6be7

Observation 5aea5da8-1175-4b5c-a206-e8dae25642b7 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.047280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.047280Z digest=sha256:97d61f05e6069c27b71424c792fdc5e512a7f4fb1890ae20891f3533a3b45d97

Observation c00739b9-2aff-4843-bf97-9b9bc98461f3 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.052248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.052248Z digest=sha256:a3844419bd463a05d1e32b8022993e4613153b66e4ab8da02bd25209ffc9680c

Observation 464d6984-42b7-43e7-980b-ed5dd2d088d2 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.057500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.057500Z digest=sha256:53348d3f19106c3a3de374aa2b0284f2d3d0ace3f3d8ca8d6de64cb24f2430b6

Observation ade186c9-82a5-49a8-be9f-a10ccce1f593 · outbound

This paper cites Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.063540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.063540Z digest=sha256:04db5fc9c30542620e978a0b6a623f43127945ee1b288916125daa93740cc21e

Observation fe4eab4d-e348-4dd7-85bd-740000ab12f2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Qwen2.5-VL Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.069223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.069223Z digest=sha256:78071089cbfb8d7535a82b6ee8ffa5a134d0307a6044d5493c6593dafd9e8ecf

Observation ff11dbdd-89e4-4039-b08e-30fd1a15485f · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.075537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.075537Z digest=sha256:7c23ba7f745b58d3b49d318410c250065716b3e163293edd3ce8cad059f1db8a

Observation b335f2bc-d348-4448-9e50-7861a3efc2c7 · outbound

This paper cites Visual instruction tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Visual instruction tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.080483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.080483Z digest=sha256:731feba41e19ae96184acd2e5c9c59ab029bf0b5df7c9a6dc71eff3b6f9aaec7

Observation 58da3d48-c780-4ecd-ab43-50f69b2a6eb4 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.085320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.085320Z digest=sha256:d3385bd9bc6fc406c383bfd3908c88a28bdc8d0ac7c2f86dd648cd0c43ebbf2a

Observation bbcb08f5-158d-40db-98bc-4747335ee78a · outbound

This paper cites Vlm-r1: A stable and generalizable r1-style large vision-language model, 2025.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Vlm-r1: A stable and generalizable r1-style large vision-language model, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.090540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.090540Z digest=sha256:dd07ed01fa8aae683c5f70fac30f93bd5ae828b0380145ea7840f9f1f27ef022

Observation 3bde4210-ac89-43db-a1de-bf6ee3940371 · outbound

This paper cites Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl, 2025.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.095874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.095874Z digest=sha256:29b1aabd019b4904da41bdf996ff22c3e863eaca110eb543404e3d493467c107

Observation 7622440a-edb0-41fd-b314-135e30cbaa0c · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.105506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.105506Z digest=sha256:a7d145be7a74ca4faed3e6b33df1dcf8c06b8484627c8dba803be029f6c01717

Observation 112b9614-e7a1-4a83-8c93-a72286806a15 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.110787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.110787Z digest=sha256:81b749629c356ec3070f1c80b88633e0cd364e94584f97c5f9d05510c3733065

Observation e23391ff-c8d4-463f-a79c-78f2522cadd0 · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.116175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.116175Z digest=sha256:b417dc3b4fea6958c878fe261376fe40b95db02a18e7f9207c237663b3bac389

Observation 6e8cb2c6-8893-4d19-ae6a-e7dba4b1234a · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.121024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.121024Z digest=sha256:d5d34ba7083bd760d055fe096904ba7644797a5ee72575bb16025d216fa2d956

Observation 642c9b11-9ea2-4095-8624-faaa2e2aafde · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.125709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.125709Z digest=sha256:f82efbefdd10a0b8883e06096407014aa673a06145fba69377183dd661128bd4

Observation 90239c86-c51a-4bc9-a116-367c475a62c1 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.130682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.130682Z digest=sha256:c9973f8f7eeb2bafbe1521ed3589ccbd6665d6c84326221ccefdb898ba6ab07e

Observation 0544f2a6-a603-43a8-804b-37d51e1ee67e · outbound

This paper cites Kimi-VL Technical Report.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Kimi-VL Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.135546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.135546Z digest=sha256:07c19f808552bcebcd4f44bf16f5b203186b2b95135e1988d28d794b9ec63078

Observation 2af91df2-18d5-4b30-88f3-717e80af052b · outbound

This paper cites Improve vision language model chain-of-thought reasoning, 2024.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Improve vision language model chain-of-thought reasoning, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.140526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.140526Z digest=sha256:560f1b3d94bcbbb28813ea2b62317af53b829e6d88808bea1f2e162151d4e614

Observation 643b81b0-db47-446e-84f0-a9d29287b31c · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, 2025.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Llava-cot: Let vision language models reason step-by-step, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.145289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.145289Z digest=sha256:246d64797f21f403bec11edb10c6b550a1f0a6f86e3379b300277b3e817e58ca

Observation d2394a16-e466-42f0-a81e-1ad9d082294d · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.150221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.150221Z digest=sha256:194b1869f755e645e83c9f29f3065afc83e25395e862578d2c4bc08ce8e333d1

Observation d2cec0f4-2c7c-45ac-8c25-2662a84eb841 · outbound

This paper cites Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.154904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.154904Z digest=sha256:a25bdd0edbbb0f44515d926adde79493adc30c4681c9cdab8413d96ebd2ebf48

Observation da7e05ac-18da-4499-9480-2d1f810a8482 · outbound

This paper cites MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.160749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.160749Z digest=sha256:195f3e31f914a97abba35ab53e43f5b40c43f0fb923c69260210dd1a004968e8

Observation 17e8a213-d68d-4a21-b5bd-71b06b1fe91a · outbound

This paper cites LESS: Selecting Influential Data for Targeted Instruction Tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LESS: Selecting Influential Data for Targeted Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.165737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.165737Z digest=sha256:8966d5cbe50da3cd1e5b34c95bdd1c880da63d1928ff626e0556d10d04994f56

Observation 183598ef-8a3e-482c-8f50-d146f36bd236 · outbound

This paper cites Estimating Training Data Influence by Tracing Gradient Descent.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Estimating Training Data Influence by Tracing Gradient Descent

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.170718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.170718Z digest=sha256:1dc6ed456b81e6b213a427f384b1643bed94158f3b9a91b475ee6683da022622

Observation fb1a6e3f-3679-4b9f-a601-dedf248447a8 · outbound

This paper cites GPT-4o System Card.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation GPT-4o System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.175673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.175673Z digest=sha256:65b4128d82ae4f2ea53ff75ac61e77fe2dff0abd6987d36bd17a9fd913dca9ce

Observation 606f5ab8-6bd8-415b-aa7a-9850ede75773 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.180238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.180238Z digest=sha256:42262e265896db8e49175d20ae8abb21e311250d74c7d15ba37cb5e56635e792

Observation 6251beb0-6008-4404-8d7a-7d8c47d1a376 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.185246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.185246Z digest=sha256:9a67c38ca15e2ee66744fb27ddf25730d942662bc9391ce1692f7f735467b992

Observation a7eb8a0b-c7b5-4ad2-950d-304f8a234132 · outbound

This paper cites Qvq: To see the world with wisdom, December 2024.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Qvq: To see the world with wisdom, December 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.190465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.190465Z digest=sha256:385fb255cdb607ff1666201e08f98357a700fe86d32cdde71dda5f571c1c1d36

Observation 68b71451-4a7a-4856-86ba-b191f7adc41a · outbound

This paper cites OpenAI o1 System Card.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation OpenAI o1 System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.195348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.195348Z digest=sha256:8eb94b466351523cd59070f210f36ee5c1596533cb38e9dd98fe461b17d99a5b

Observation f24ef1ec-9434-4f03-a185-96858763d27b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.200230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.200230Z digest=sha256:5eae58fc3c616b03f5774d8eae0ddda6a98f9e6f47c4f8aa725e40aa46707b0f

Observation 88f412a1-b5a5-45f7-a2bf-78a6b741fa94 · outbound

This paper cites an unresolved cited work.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.205225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.205225Z digest=sha256:94f0a22541618345b95a7dd44a60afe0f262be7bd1f95ffbffe7a08d4d7c0ef5

Observation 5d4b8d6e-260a-49aa-9a78-6e1471136246 · outbound

This paper cites R1-v: Reinforcing super gen- eralization ability in vision-language models with less than $3.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation R1-v: Reinforcing super gen- eralization ability in vision-language models with less than $3

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.210995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.210995Z digest=sha256:7908227ca95ea26b7087522a88113bf07d29928a5f34164b29cf61fb6b348805

Observation a3c2899c-091c-4675-aa56-5eb2bec22a5b · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.215793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.215793Z digest=sha256:8ee7746acff5b6595cb38fcff91e6fde23aaeab55f4917e6bb70ce7b8cd70c25

Observation fbd8b0af-21e1-4a8c-9841-09f1b4e70411 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.220934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.220934Z digest=sha256:be470b855cf970d4cd42b1d09aaea59f60e5b441ea6b83529de69cabed7e729e

Observation cc2396e8-cd50-45ba-8cd9-3706eac66bb3 · outbound

This paper cites Lima: Less is more for alignment.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Lima: Less is more for alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.225981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.225981Z digest=sha256:fb9b3f1260eec397902cbaad4d99b8a8506ef0b23317fe3c2d2ec64db0e1e9d7

Observation f35c73f3-39d1-45fb-807c-6ba3f365470f · outbound

This paper cites Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.231109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.231109Z digest=sha256:addfb77977a3b4896192e99afef26da8da622cae7e2c1b963da72483818cea36

Observation 12946d17-c9d3-47d6-b083-e47acee5baef · outbound

This paper cites LLM-Assisted Code Cleaning For Training Accurate Code Generators.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LLM-Assisted Code Cleaning For Training Accurate Code Generators

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.236832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.236832Z digest=sha256:bd730bf6750a740f95fd3f2bf4fb57b190bdf28f9f01e83343dac763732fecba

Observation c6e8cc79-7e3c-4088-a071-b849fad1b69e · outbound

This paper cites What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.242672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.242672Z digest=sha256:8d75241f641820f813b0af38ff59c4803c6961b5cd42663a165bc2df2875f374

Observation c73c9da2-a9c0-46e3-a726-cbb2d4663d9a · outbound

This paper cites Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.247730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.247730Z digest=sha256:a24827e78c235cc9d79d8f3b0da437d2010e31e18ca5eacae5f970b9153cbfd2

Observation 0cb1384e-8139-4298-8508-cca5bcef0348 · outbound

This paper cites OctoPack: Instruction Tuning Code Large Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation OctoPack: Instruction Tuning Code Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.252677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.252677Z digest=sha256:80b616b3d551d346225739f1b59f8b89438ae649995146933b856dbc9852cce1

Observation 30a91729-8420-44bf-8e9b-39f6759bcbb0 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Lora: Low-rank adaptation of large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.257664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.257664Z digest=sha256:23826e9d64c7feb1088c7eafcd9f67b1cb1ef97ab2caeae211719c88f96626d9

Observation acb0228c-59ed-4336-bbd5-7f6fc4c76ada · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.262872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.262872Z digest=sha256:0853196b7d3eddab068fc4cf404c5ba5b0a8456e758123df27fbf0b391d4af24

Observation d4d2e983-fcd9-4d5b-808c-e86268208079 · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.267817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.267817Z digest=sha256:749c47a9e1157ab0d46e4373d003f3f6e4bd93a35b72614dd6bdf3b96de8886f

Observation 46cf7362-e5d1-46a6-bb99-4f1ba45e377c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.272810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.272810Z digest=sha256:7fb47a86ce60617a99c1c850a667cb174bbe28358f362c3812190a2c9a5eab81

Observation c5f39aea-717f-4e37-9215-6e84e5962235 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.277999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.277999Z digest=sha256:860b541f382ede122a482730dab46b98c6db2fa23ee08ed6704144d10ab89c57

Observation 75c08491-55e3-47da-af83-9fb8f069cbc3 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.283136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.283136Z digest=sha256:982c7f380105343dc1068622bd370cb018934c5bed1399497fea87a425068342

Observation b9656623-4767-4186-94f5-b2ae1fea1b86 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.288087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.288087Z digest=sha256:7245fcbb72883f51e135a31cc16f7fafa14b9fb837256e3f59a86e264db1bcab

Observation 697d2f21-c6f5-45bb-b7ca-7fd28bfbefa6 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Measuring multimodal mathematical reasoning with math-vision dataset

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.292710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.292710Z digest=sha256:7eb792c6d4da9dd614fb8b13d63896ccfcf16e071ef86707424f6be1fa9083a8

Observation 065bef9b-408c-473e-ac24-3904ed4f939e · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.297641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.297641Z digest=sha256:7881bcfd58e9d66edd161e9004eba3cb1c31ca86b71f13259dd182072f4d9b4a

Observation 16a605b8-022a-4945-9fd4-0bc8dbb33625 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.302455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.302455Z digest=sha256:11f088b57a7cbf86ba68ced9854312b83aeb280e52d41d1766e07b3ef3ca4344

Observation 659a2e93-4df8-4d7d-9d13-5bcad45354f1 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.307694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.307694Z digest=sha256:0c0ff6ebea23593f2cb32696d0a1d5554972059270bdafa474433c17d9ab9618

Observation ab350932-1d83-4672-a1b2-6435f145c084 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.313024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.313024Z digest=sha256:d1365bf9d05b13c047cfa8bf4a5def64ec13da3d6af840aa5025b71b0c58b18b

Observation 2a673b41-5ccf-4fb4-8a99-b9d678df3c72 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.318268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.318268Z digest=sha256:5e9e7c6978352de3c27f51334df23b1ec3220e5d5e2f870ce00904415b8ddc8d

Observation bd9394ee-27c1-4d83-ad73-20f2d6ae886d · outbound

This paper cites ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.322949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.322949Z digest=sha256:f32ef969f528a2bc1c7da30fb538eb5d3267091db53c1de3d1ca54597f75a273

Observation 72bfa829-d95b-4eb4-abd1-7d58dce836d3 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation A dataset of clinically generated visual questions and answers about radiology images

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.327988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.327988Z digest=sha256:1967addc0abb1ceb35bb43960bf03c438b2ea335a4d71e5f4234a1ff767e3431

Observation f841676b-ba39-40c1-b25e-5ffa31b25379 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.332746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.332746Z digest=sha256:0de789fe53bb2b07ddb59cf35c75cd4305dc7efa973169904c132dc7350437c9

Observation 4a3f783a-95aa-4417-97e9-298d15cbc6f2 · outbound

This paper cites Slake: A semantically- labeled knowledge-enhanced dataset for medical visual question answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Slake: A semantically- labeled knowledge-enhanced dataset for medical visual question answering

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.337699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.337699Z digest=sha256:4c781f80d338ac4f1ea5b1629f2653cd6cbd83f19438062f09635b706d685a3d

Observation ff7cb12c-657c-4e22-8879-62ac8b89aca1 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.342648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.342648Z digest=sha256:39e62bc4dbab920928added59b4ff69d7e50e17c8fc8a4a4ccf1a001f7747b12

Observation 72f64ce5-8e37-4f3a-b84e-f3d780d1ae00 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.347677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.347677Z digest=sha256:792d58f84f27043d4a9497e121b4bf9c46b04769dc8a8706c1c7fd8f4e464448

Observation da29709f-a27e-4bc2-b702-7e3df9903af3 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.352870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.352870Z digest=sha256:cb31cc4debfcdaacb08cde05a7d16f83ce443075603d9b6ab103397b08c36899

Observation 34ae2762-e1ed-41ae-8ea5-71819f1e94ae · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation LLaVA-OneVision: Easy Visual Task Transfer

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.357697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.357697Z digest=sha256:c5744f019a61948fbe0a8521a92e392f9b9730538b4aa52db54ef3ff403b5634

Observation f10e8a88-ecf0-443e-9358-9d7289054b12 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.362835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.362835Z digest=sha256:440ff86f047647b5b793330aa1d65def859062aaf73461b52b6af50012ba1056

Observation 65c6103f-3e1b-40b1-8d0f-fd6070afde1c · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.368349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.368349Z digest=sha256:cfad54403525d7f93b5637193cb9ad4695d6d07c6883823919445e296741da06

Observation 0e724969-432f-4218-8cb3-d931bb78a8dd · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.373195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.373195Z digest=sha256:4b746332d2b4fd2eebde67640d98cbf317a76ec5207205d3faaeb9e42fa5dd48

Observation c6bf0777-e8da-4546-96d7-41539bd2234a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Proximal Policy Optimization Algorithms

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.378185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.378185Z digest=sha256:1f025dc93cc318c2c41c37513d3fbc4380a18c44272c2add3fbf3d8d6d876750

Observation 376bacae-793c-42dd-9047-57ab352cf7a8 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.382930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.382930Z digest=sha256:b7760c3830691e827db40f64104bfcc13f8c5c3b0a25e81b625440969d62ae33

Observation 63c1af2b-ad80-420f-97f4-2035f9b86daa · outbound

This paper cites Dvqa: Understanding data visualizations via question answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Dvqa: Understanding data visualizations via question answering

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.387919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.387919Z digest=sha256:5feea1165aea3cc4bd1e81e15a136e024fae81ddd3de549aa01ad9457816c743

Observation 4ed0157a-9167-4bc7-b5e0-d26bd703336f · outbound

This paper cites Plotqa: Reasoning over scientific plots.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Plotqa: Reasoning over scientific plots

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.392660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.392660Z digest=sha256:243cb9cf6f909dc1a71647d2086e4821f8193fe3c6f6606521839552eed7b8be

Observation 96109899-f9ff-431b-b58f-a0ed11b218ed · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.397621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.397621Z digest=sha256:2a9139255972e6ce92c5197069caabe472197755ddb98cd35fa7a42c5219ff8f

Observation 48732e57-7c6b-425b-8266-7a93b8ea40f1 · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.402852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.402852Z digest=sha256:9171171a810f4fbf72e5d0103dee44125dd594c024caa35cf16a66c02eb4486e

Observation 2ae55cf7-6324-4e57-b2c2-2e1e7fa2bfe2 · outbound

This paper cites ChartBench: A Benchmark for Complex Visual Reasoning in Charts.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation ChartBench: A Benchmark for Complex Visual Reasoning in Charts

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.407822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.407822Z digest=sha256:6782dbfeeef094790e7bb1df07d1b62312d6ab4e57c7b58e731b6e0171d45ea0

Observation 23c48653-37e2-4852-b360-82b30dbae3e2 · outbound

This paper cites Unichart: A universal vision-language pretrained model for chart comprehension and reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Unichart: A universal vision-language pretrained model for chart comprehension and reasoning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.412542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.412542Z digest=sha256:80a0991e4c6a368e074278ebb7dbbcccee4b4ede1978a5e38b412466ce490213

Observation 6132d4de-611d-43a2-9dc3-0ffc6e463a72 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Docvqa: A dataset for vqa on document images

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.417054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.417054Z digest=sha256:dd139d133477e3f8a6ea864921e528cb0075f7ad2d28176eaa0f563c3c8c467b

Observation 30ce4f23-7c12-45e3-b24e-577d02c4cd73 · outbound

This paper cites Harnessing Webpage UIs for Text-Rich Visual Understanding.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Harnessing Webpage UIs for Text-Rich Visual Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.422038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.422038Z digest=sha256:4f8277a2b4ed9a8cdc3fa661d84b1992d2aef8735ed1a4c244d59a57fe4e79c9

Observation 9418bd85-c530-41be-aeeb-1d4b5681b22b · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.427313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.427313Z digest=sha256:6e1f9896e8dc8a67be26608507ca0aabc7ed1e645e381c2f52e62ef9aaab59d1

Observation 769eb7e2-f057-4600-a319-7c77768b5187 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.432638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.432638Z digest=sha256:2b3420e988d45a9b6bf8f0bd1dc75886e18dcf4453a1f6e6dcae3d0fa0c126a5

Observation d08d20f2-47c6-4c1f-a5ad-588660b69e8b · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.437400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.437400Z digest=sha256:1234a4091bb251cf7075e07cb213f8f59db0509dd853ea60e8f192c47f8b73b2

Observation 3b5cebb5-9ff0-4f8d-88d3-4b0a296a8cf4 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.442245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.442245Z digest=sha256:bb835d42dfa966756664e20c97260b9f437668aa2d9fa4dc96868007ef868958

Observation 84cebb8a-e84c-4ded-942a-a9e0dab82fb6 · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Solving geometry problems: Combining text and diagram interpretation

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.447210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.447210Z digest=sha256:eb6af65ee8c60b5eb42e56b93462491e37746989b3ad96d85fdb21330ce5377a

Observation 69d8adda-2460-48f0-be5b-5764e70775e3 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.452090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.452090Z digest=sha256:f76d681b5f518fce6f073db01a68ace0f85de99fca3d6dc9a0dbedbd9945c375

Observation 1ae9e6f8-e958-4a16-bd67-35a411a8f7fa · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.457298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.457298Z digest=sha256:35bbc991f05142bb430918c472e8b45f6d6c5a17f15dbe39330ce3e10bb5717f

Observation 3229e0cd-4c93-42b0-a050-a02eba852c68 · outbound

This paper cites A Corpus for Reasoning About Natural Language Grounded in Photographs.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation A Corpus for Reasoning About Natural Language Grounded in Photographs

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.462326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.462326Z digest=sha256:f315f1048957dc29896751484bb9d3b95fa505f6010b03052b6a55151b66df18

Observation 4e6016ae-9dd9-47a4-b9a8-4917f9ef2ff6 · outbound

This paper cites Image Retrieval from Contextual Descriptions.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Image Retrieval from Contextual Descriptions

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:24:08.770841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:24:08.467219Z digest=sha256:22b6cb12b4a065f432f18c2d242890e961c1470751882ab95788cb2332302d18

Observation 6a872cae-9d7c-4986-9b99-f5f33a6cec46 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Lawrence Zitnick, and Devi Parikh

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.472416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.472416Z digest=sha256:862b80e746bb14199815a9ea7d1fc7380caef84026d6d83a608d35a45c4ec92f

Observation eb282081-0b1c-42d3-8e7d-198b15c2a976 · outbound

This paper cites Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:24:10.024942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:24:08.477274Z digest=sha256:6cf28e20178767517ab9ee18646826c6471b393669ad43c67aafbb5b4f658b04

Observation de48a35c-7bd1-4689-8db8-a768aece6e2e · outbound

This paper cites A diagram is worth a dozen images.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation A diagram is worth a dozen images

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.481966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.481966Z digest=sha256:aa237131634e36cead24e1cf8a14f85713d817f60a65ef78515f846eaacd10b2

Observation d15722cf-d116-4081-a0d7-a43e7e8b4b34 · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:24:09.996101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:24:08.486479Z digest=sha256:e1bf3fd6e2fe2b05803872a93eaa713a37dca38fbaef0ede827696ff00bfdb4e

Observation cd6f8cb2-0f48-4a38-a74d-5288742c67d2 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.490975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.490975Z digest=sha256:0ecfb0b0abcccf90ac6d78f002cbf65e822df75f21974ce1bb7182a336097213

Observation ab713047-994d-4443-8b21-b87b0005d038 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Vizwiz grand challenge: Answering visual questions from blind people

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.495486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.495486Z digest=sha256:4f71eaf4da54e32b3675cb732b722a39166ee1f90d041175259e6e4a78a58274

Observation bfe0133e-91c3-4c81-85a7-0729729d3aa8 · outbound

This paper cites Towards vqa models that can read.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Towards vqa models that can read

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.500513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.500513Z digest=sha256:5656a7b9e8c25294d8949bc27a68b7ac633c5e473576c11c4a18be062866bd01

Observation b17bf1cf-2bfd-4532-bedf-bc28197e34dd · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation A-okvqa: A benchmark for visual question answering using world knowledge

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.505250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.505250Z digest=sha256:0fb1b1fbc5dd7d76f38b76c1bd4921160584de75a308df6527e664dbdc0b7a6d

Observation e24b1632-9004-4fb0-b05f-7da7199d71b8 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.510153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.510153Z digest=sha256:ffccb40aabcf357c97069fd8a597aeb8e2d2797c32a16613b8f34b35c969ae08

Observation 15866e9d-4e44-40d8-9ab7-a6b5d64790b7 · outbound

This paper cites PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.514831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.514831Z digest=sha256:0c3216ca68e5236b226e97da4d09dde58665fb1d93f4a90a01ad46b0b4c2d5fc

Observation 0d26467e-9283-4c20-9eb0-d4c876cad2fb · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.519580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.519580Z digest=sha256:f6f3bfdb6c4bef176b4745b0ff26a1105a73abed36750abbb70d38f265f89067

Observation f701a439-9b26-42c4-9d6f-f0b003162a47 · outbound

This paper cites Chart-r1: Chain-of-thought supervision and reinforcement for advanced chart reasoner.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Chart-r1: Chain-of-thought supervision and reinforcement for advanced chart reasoner

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.524451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.524451Z digest=sha256:f56f34372c75878c3e3cff53d2add5f23a2ecd62edac9420ec6892cf7bc6c62d

Observation 76ea1cc2-7929-4419-bcf2-0c36ff67336e · outbound

This paper cites You are a QUESTION-TYPE classifier (do **NOT** answer the question itself).

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation You are a QUESTION-TYPE classifier (do **NOT** answer the question itself)

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:24:09.922336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:24:08.529010Z digest=sha256:6e3eb08e9d32b39de197921b8898b5a165f0ff047752909d27f030872d62ef32

Pith citing papers

Observation 6bf98c01-7ddf-4da9-888a-95077f7dad34 · inbound

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks cites this paper.

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:30:58.744656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:09:33.968186Z digest=sha256:1b846b03ba62488cbd2b9d33df6b5231e6a3b7c79f3445232e661d37d996a8aa

Observation 27a7c529-19a8-4113-909d-20fb608e0500 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.296681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:f9440b4f3147866091c0a041c6f234c6df6835598ca184412be9d0e75a0d8293

Observation f3d80314-5b0f-4b0b-bdf7-8e8705610c62 · inbound

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution cites this paper.

CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.242064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T12:40:34.424341Z digest=sha256:687d399ad06938d75aab3003a8c39d467ad353947a5b827bfd07ed5695779f99

Observation f4626655-0973-4564-9de3-00cec701f71f · inbound

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning cites this paper.

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:09.398205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T14:02:29.480442Z digest=sha256:7f4710e590e1dcdde0a2b84bde3a9043859c1a09978732b16da7c90a36c385a9