Pith. sign in

Paper Citation Record · LEDGER

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

As of 20 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 6 inbound Pith citation observations for arXiv:2506.10128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10128 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:40:13.423620Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:39:18.945567Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:23:28.446166Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 529e2bc9-2351-42ea-a4c2-476f671be824 · outbound

This paper cites Vqa: Visual question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:18.085418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:06.644920Z digest=sha256:a20ece484557650e690618efba7f367515342d81b8f109b3ead3135392e6c844

Observation b187ba96-3bfb-41b1-907c-2d4870c92b3a · outbound

This paper cites BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.691389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.691389Z digest=sha256:0914c0bd794ba492388ecc446b25beb04ab50abda2be18c891b354c90b65d7f2

Observation 295ab65c-4f3b-437d-a1cc-28902ee60831 · outbound

This paper cites Qwen2.5-VL Technical Report.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.769409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.769409Z digest=sha256:619df7a164a4b932854c28c3c1e37fcaaa0cb03eaa2cebe022bf45bb184a9225

Observation 0b6f18c6-6c7c-4c55-8a25-776297f6b43e · outbound

This paper cites Improving image generation with better captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.852093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.852093Z digest=sha256:66dcedcca5fef290be0ff2775a5ca20abf544150f777e668f0320a5b31936bc7

Observation 61c4cd51-1079-4709-8a57-8f561729cab0 · outbound

This paper cites Language models are few-shot learners.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.950928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.950928Z digest=sha256:784b8aaa033040461a873340a89feb39c0d98294639f576afc38405f1ce05322

Observation e148fb5c-6010-4a9d-ac4b-2882ddc5bda8 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision language models with less than \ 3.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs R1-v: Reinforcing super generalization ability in vision language models with less than \ 3

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.814290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.055773Z digest=sha256:d2cba9b7a70dcb84ccc8c8e4e6ecd87b80c05d8158d19528698122b50266382a

Observation b165f18e-9ec9-4319-b4ae-01800a07b4a2 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Sharegpt4v: Improving large multi-modal models with better captions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.479967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.167181Z digest=sha256:002a44eb195a0fed1a791b7f9743a0191c6a9ae51ad235730de1de859c862dd5

Observation 16c87061-4b08-42b8-9650-ddbf4263dfbf · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.282039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.282039Z digest=sha256:305bdc12f38bc77d4d9a0850b54ac752e036c8adeb3de4142ed36a2433029dd7

Observation 4278674d-e3d4-47ce-97c4-9aa33a023f0a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.352289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.352289Z digest=sha256:fa9815cf49818833521618be1b8b60fca4a61d6f057f5e5a63b44648329447ad

Observation 0f547de1-4259-4664-ae43-04fc56b9f4fd · outbound

This paper cites Palm: Scaling language modeling with pathways.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Palm: Scaling language modeling with pathways

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.423678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.423678Z digest=sha256:9f64c019064f52d74a7daffb5c612b3b73dd66271fa073df2f8d8eaf04ca4106

Observation ecfc56bd-d7c3-441d-89cc-0496c50d2888 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.530509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.530509Z digest=sha256:b8a4cc26b17bc9357f87d694425958bd1186f40eed74fce2251de2bb248c168d

Observation dbfc9be6-f678-4ffe-b456-45023e87f0ef · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.610397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.610397Z digest=sha256:17cd52eb3029b4e10a6fcba1f3cf729540fc6224a4911e63b229ab42c77b895d

Observation 6d3f5767-fb1f-4614-8730-bdcd5fdd560e · outbound

This paper cites Enhancing large vision language models with self-training on image comprehension.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Enhancing large vision language models with self-training on image comprehension

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.196649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.695543Z digest=sha256:4c27189a455f4903fb468c8fb5e4ea81ed45b44553ef86d2b46ea8f4fbe3dd56

Observation c9ae7c9d-2336-4ce8-a60a-05d82a6a63b4 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.764980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.764980Z digest=sha256:a359a2c949b0937e6d2329613f5408dbc2b5f92a4264437d7e1bc1258dbe4510

Observation 0af995ae-2b43-48e2-804f-57696d41c56e · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Virtex: Learning visual representations from textual annotations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.922173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.899067Z digest=sha256:7689c75abe59c172f3278d4d8d88a38e63ec50eb2154ec13d2b2db728c1bde10

Observation 22bcc49e-abe1-4239-a509-8b6c2363d5ac · outbound

This paper cites ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.980626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.980626Z digest=sha256:78cbe1e108b0e31c3b537b9f00d1bd5ad51be291e9c0549adac7eccfff2d4a00

Observation 57b69c66-5c8e-451d-b509-23db3c4f7fb6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.094721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.094721Z digest=sha256:8418cdbe393c074d991a0e1f94d1bc0f9344fae02e4520642e97c8aab911e095

Observation 9024f7a6-09d5-44b8-a2c3-b8d1e33ac7af · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.211304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.211304Z digest=sha256:adfdb0c840821451585898a6de56e6e8d6c9bed07492d7dd11d8a7b14db27b18

Observation 4a3b41b7-8ec1-4758-af92-85adf205d4fd · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Measuring mathematical problem solving with the math dataset, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.696254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:08.309188Z digest=sha256:3eb065511004c734216d1c910a45fbe1a7e2256b11f133b6c7d587a8dc754115

Observation a280af47-a3c9-45c7-aff0-7c06e02d297f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.388160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.388160Z digest=sha256:38783e9f028c48560bc43c5c6f15b5147f68f37a75cf0c16d2a4b53b207e989e

Observation d93473b7-9c98-41ff-b8f5-0e2bc50c2d90 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.503288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.503288Z digest=sha256:e8f0ca9769d222287724b3875e86271cac7b889b67694683d90752c6ba6460d3

Observation ff924b2c-8786-46bf-8965-ee35e28c2680 · outbound

This paper cites GPT-4o System Card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.599239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.599239Z digest=sha256:08b436a2abf4c58aeb0eba88d491b285060ae2a42310109b9a3976f906671bde

Observation 55ce0b4f-f7f3-49ef-b31a-191a7cb46f55 · outbound

This paper cites OpenAI o1 System Card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OpenAI o1 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.696397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.696397Z digest=sha256:265667b9876fb73543c43c4cf8fe083ba2e3a3c1251be9eafd408c09ff415aa8

Observation 01e4d8dd-532b-4c51-b697-43c1f4837be4 · outbound

This paper cites Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.808188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.808188Z digest=sha256:a274bade51dd8534133f1c2538ec6e54872d7f2f4f95f03489d5d4aaa1181248

Observation 4fcdadb3-48e7-4df2-a402-d1309441b60a · outbound

This paper cites Large language models are zero-shot reasoners.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Large language models are zero-shot reasoners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.912895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.912895Z digest=sha256:1ac436b56e33b3791d8fa7886a902a6911df0162377d24dd7ae66a169ecadacd

Observation e728da3d-5c57-40c5-aa2b-ff39f97320fe · outbound

This paper cites Mawps: A math word problem repository.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mawps: A math word problem repository

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.994633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.994633Z digest=sha256:62058aaa4ee4e853657b60b3cd7c6335ca29b6f918f4848772c1bc5185c416c1

Observation 30fd708f-6fe5-4f35-a70a-d1d0c37805a5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.073387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.073387Z digest=sha256:07f03b53106d0e25c193af253bd12adcadfbcd3f3cb270b1f6a7d43649e26da7

Observation 38288d0a-69ec-4b66-b564-7c40f6460735 · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose assistants.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Multimodal foundation models: From specialists to general-purpose assistants

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.437363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.168271Z digest=sha256:f869f299c81690995e9a71122050682c7274133cd8c35517ba5eb7de43d753ec

Observation 9f99e18b-c7a3-42f5-8ecf-e7e4d906f080 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Describe Anything: Detailed Localized Image and Video Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.244875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.244875Z digest=sha256:c2d278ab5bbf7195e3e11ff86579453b285a09054c65f3aea53490f3cc1665d9

Observation 95ca5810-2016-4a10-94ba-bc9b988f8388 · outbound

This paper cites Let's Verify Step by Step.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Let's Verify Step by Step

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.313417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.313417Z digest=sha256:d265ec52debed1e0e3c7cca5afcdeb40b8b390414aebd6b37ac22bacf62de74c

Observation 41609576-8914-4c4a-8c88-98c69927a580 · outbound

This paper cites Visual instruction tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Visual instruction tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.436777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.436777Z digest=sha256:b3e9cff5ed3aa1457184fc454b0c45a70ac483c0779a67aab5c8d2bbde81c29c

Observation 290e02e6-bfd3-4282-b57a-f67afa3e8367 · outbound

This paper cites Improved baselines with visual instruction tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.179830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.523923Z digest=sha256:4078564cfe478f508ad965457da201a36ccf5dec9599e028d4fa3ef42149dc03

Observation f4f3460a-ac9d-45af-b995-8b5b62531ca1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.949313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.582883Z digest=sha256:c7f83b172aa9caa33a47fe6499b7a592b7dead70cb59a760c29caa5598b5b5b1

Observation a16ab06f-44fb-40d8-819f-5b49a4350f77 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.661306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.661306Z digest=sha256:1109d9e865cc14e59e0bf0c5b633fd2234ef8089a475f8035ffe5ab76dbe3fbd

Observation f3d4d0d7-fb17-48da-bc52-bc06a535396e · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.721963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.721963Z digest=sha256:fb61e2defd51ca1e4c3e879a9a4237ef491f75598f748ae91fbad5cb4005513a

Observation 6bb0b974-fd24-4e7b-bed7-a382064bb313 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.920757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.920757Z digest=sha256:0599765dd83b28b612d3e13c97b382cec5e0933bbd7693dd64b21ce3050c467e

Observation d39a454f-ee62-445b-b7b6-4257d893cb8e · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.021002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.021002Z digest=sha256:0d209d83979ad7d68e8d3f73ab611ff8a58c054d95b4985ede9d1d721197bdf0

Observation 722d67f4-93c7-460a-a6be-ff80c6b107ed · outbound

This paper cites s1: Simple test-time scaling.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs s1: Simple test-time scaling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.113399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.113399Z digest=sha256:f004841a5179e2a5210a14e3f695b855cf37bf5a71e282ded46bb6794b8a4fa4

Observation 25071772-0c3a-415c-822f-7b225daad8ec · outbound

This paper cites Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.179618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.179618Z digest=sha256:64a7a6fbe96862008a10f546c49115eabf6a04f4a1d3c826fb3001fbb5002e2b

Observation 2adb2fad-fb06-4f7e-8be5-9b1f66733eff · outbound

This paper cites American invitational mathematics examination (aime) 2024: Competition problems.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs American invitational mathematics examination (aime) 2024: Competition problems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.691239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.250543Z digest=sha256:24075e6a2f681225b511526bb268e213f35424692e974233d2d6bacf0591e9aa

Observation 37a766c3-8881-48b5-811e-16d0fdcefe63 · outbound

This paper cites Gpt-4v(ision) system card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Gpt-4v(ision) system card

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.331609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.331609Z digest=sha256:0fc1589388853f25c2c746a315e277c108a471954d5dc4bf999acd996d7b4136

Observation 1caae194-1900-484a-bbe1-021d5c4fa937 · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Are NLP Models really able to Solve Simple Math Word Problems?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.409957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.409957Z digest=sha256:bdf8fe4ccc44cdbc7e232e241fb0ced3b6fdc56eed45f80e3cb8667816867453

Observation daf7b12b-f92d-41f2-a761-5b29de60949c · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.484752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.484752Z digest=sha256:255c629b306a4d83fc42f9f4fabcb94666af38192e9d29cc5d833cb1611ca612

Observation ff5c66e0-b5f7-4028-9bfd-dc9b592fb697 · outbound

This paper cites Connecting vision and language with localized narratives.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Connecting vision and language with localized narratives

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.460773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.569209Z digest=sha256:9b6cc9ae7eaa5e82f39c2028658ed97ab0761a0417f8db41a5a434c29988c5b8

Observation b8f3a763-69aa-4b6a-a1fc-332c2d1c65c5 · outbound

This paper cites Vision language models are blind.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vision language models are blind

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.655345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.655345Z digest=sha256:8e9920744357ea0a20ad31b78366c265a909f95f3519e6a57f693e2e4656906a

Observation 6e65327d-6dad-4ae9-a7e2-549795a65948 · outbound

This paper cites Object Hallucination in Image Captioning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Object Hallucination in Image Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.721223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.721223Z digest=sha256:d375b863ae5cb6f82d2382f354c9dcd066e459cf0e213e3dc56074e3ef28c163

Observation d6973e5d-3d08-4214-bbe4-a5764c3ead37 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.790536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.790536Z digest=sha256:277ecca830ed8592ec3ca57f508580e6bea346ca285485e02837c5123c3bb590

Observation 5e318aa9-e996-4b3e-b39d-12ec4690f758 · outbound

This paper cites Learning visual representations with caption annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Learning visual representations with caption annotations

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.227903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.860273Z digest=sha256:28e8978f52224f6b19715ddfadd21d27f587df134c89982d142a47ce693c516e

Observation 5d4ac2c6-bece-4c01-bdd7-b0b8913054f9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.915730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.915730Z digest=sha256:8b5c96a58084c699d394343ab86e71e864ca2fea3c616719d7c5e902c7c7a2f9

Observation 7de2822a-dbea-45c7-869b-77e795584691 · outbound

This paper cites Towards vqa models that can read.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Towards vqa models that can read

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.020825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.020825Z digest=sha256:375152f75fe8c3fc8325236dd2fc75da200c9eb5d0b3346f7285a358d6e8bb3b

Observation 269f59d6-2c04-4850-a77f-cac1edfdf181 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.110047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.110047Z digest=sha256:dfcb90dad775bc74b30e34f5e528f4c9bc3f50cbf3dc9a1f73e921c54e30e12d

Observation 60076ad4-8eca-49d0-9755-2c0751f93e70 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.186415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.186415Z digest=sha256:4149df53dcdc6319b063c2389da9fee950c150e830ffc59cc298a280b7f1da16

Observation 4d1d6408-ca4f-4847-91fc-9ff115899846 · outbound

This paper cites Image captioners are scalable vision learners too.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Image captioners are scalable vision learners too

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.018460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:11.255836Z digest=sha256:9e901439750e672b249d3b08b4179448308029da4e226c43ec0ccf7df19b9076

Observation b64bb195-56a1-4555-9596-aaea9b217669 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Solving math word problems with process- and outcome-based feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.367240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.367240Z digest=sha256:0c0ff2fa0db39c41e14ebfe3cf9ec02490004297ee54f9a7e5d5626088f70925

Observation a3c4f7f1-599f-46c6-8676-c515b928c0cb · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.443100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.443100Z digest=sha256:340f8ab061a8dd6f6bdd583528afc85d33cc6d2777a55c127398674d55ee39e1

Observation 8964877a-7d06-4be7-a438-7545a15f4629 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Measuring multimodal mathematical reasoning with math-vision dataset

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.511675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.511675Z digest=sha256:a0661ba2dd8444d29aeb9617134e35816e1c5b6753c833fe0b4009b9998ba7a7

Observation 6b811447-8545-46ff-98c8-3544f21aa077 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.584688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.584688Z digest=sha256:e5330d426ec6fd32f1eda017bdb452e571f4074dba2aa6a9287362c927906843

Observation 677d734b-80d1-4bf2-b4ce-31d3dbac3d35 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.660360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.660360Z digest=sha256:f9464069a3634d725df39cc0bb519d262cf120162c0a72cf4ae368cdbc36ffc8

Observation d0da67b2-01d5-4d05-a158-b71c70f17dfa · outbound

This paper cites Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.741673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.741673Z digest=sha256:3390db0bfaecce92c846716a5558fac0257c32221169abf8c2033bf896cacd8f

Observation a9559091-c7ca-459c-8da7-146301ceae8d · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.821726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.821726Z digest=sha256:cc8553dc721bed38db6932b87fb155d4a3f377de8b13444399edd38f3385f5b0

Observation 2110b634-0937-40a8-9cef-88c00df6502c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.940289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.940289Z digest=sha256:2ec57d1c04e2f02b27cc47b30d3f6631ae953db874d4ff466c0e6843ff6b694e

Observation 3b43b8ad-bc20-443f-a553-8b2bf2522a54 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.800291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:11.995421Z digest=sha256:a8a35c0caa3c14e14f91f9fd58928d0e57ed042a5a72f9cb2cfe3594ad793613

Observation 3b1d654c-9cee-49bc-a702-9e6a175e7a8f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.078191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.078191Z digest=sha256:239f10e634e8e6a87d95a1c8e88d13c041e9a87943f4e7c862b7fb5698b3f568

Observation 54962dd3-54f6-4f1b-b9d8-8513fe06f20b · outbound

This paper cites Grit: A generative region-to-text transformer for object understanding.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Grit: A generative region-to-text transformer for object understanding

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.545092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:12.170101Z digest=sha256:f18fa8cabbf801d2a1b34fbe9349f3edd2dbad0efc8b704d0633134a77103199

Observation ecbdf432-dae1-4225-80a3-8e3fafb23c0e · outbound

This paper cites Mind's eye of llms: Visualization-of-thought elicits spatial reasoning in large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mind's eye of llms: Visualization-of-thought elicits spatial reasoning in large language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.273757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:12.229035Z digest=sha256:597f3437c6e341208bb736653180f8ef12661e95318060a84b9e8d14434d4718

Observation 52378c17-6432-4034-b1cb-652c1758abaf · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.319293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.319293Z digest=sha256:fbd61df20b9096a72f44bd3e643785d5b011dc13403c610b66203f58bef23ee6

Observation b15cf5c1-0f16-4079-9e14-187d98653191 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.386906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.386906Z digest=sha256:604e80aa600901c1194ce054f8bfba23bdac4f18eb9c4603c8ef22b9bc7a6f61

Observation 8006c40d-0ba6-455c-9cb8-4c07c34460d3 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.468494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.468494Z digest=sha256:70fb2888d46ba516e29a94c5a07cf8fdfb6886f965cdbd3e425cd414339861cd

Observation e67b9949-c7f6-4feb-aa0c-0f818da0e6db · outbound

This paper cites Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.565375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.565375Z digest=sha256:a67bb46786aa45e9c58849a514ca9b465e36bf65a38453de75747d21258cb762

Observation 8d3a5b4f-6365-49df-ad0d-e2fd06af8d16 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Tree of thoughts: Deliberate problem solving with large language models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.641376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.641376Z digest=sha256:85f47e23dd7d2c9daf94c4641ada84e649f1e47850c2e888450828dce4d17b57

Observation 8e004ac8-a322-4de4-a87a-5b0d3bacc28b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.722299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.722299Z digest=sha256:3d92a62090f80c5c572246e83abddc945be77f85e246ea208713c7f95363bf33

Observation 1764ca9d-63de-47ee-a84e-1487f08812a7 · outbound

This paper cites MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.782437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.782437Z digest=sha256:824590e5373ca0e967124ae40a5b22287042a60a5153b02af91a221b65523660

Observation 547fe278-652e-451d-879f-2af3e13b2ae0 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.843716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.843716Z digest=sha256:b31217b3d84337f385eb206a4d3cdb8e7c4a8687a5f03b5b32bb321ba7893126

Observation 0e7a1d09-8551-47e3-bb06-3a9c18d20f65 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.908645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.908645Z digest=sha256:709ed61eb7f899bdba31c72d614e752daec0a5b353d7f91e01bdb08acd1d5386

Observation da366a1c-b9f7-4218-9d91-626a390921c3 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improve Vision Language Model Chain-of-thought Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.982412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.982412Z digest=sha256:d0c8f2b4cf3570e39c60e020c31ac40e517e41f8009800b84a514df14fa28ea5

Observation 825be651-6293-4306-aa15-83976889b032 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Multimodal chain-of-thought reasoning in language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.087716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:40:13.071416Z digest=sha256:d82e2001d4a239ab61754ea8687b45857a9c03bf5ad9f6b45d5153eb8e062fde

Observation 69ab859d-d281-4114-ba82-84a7bf3038d7 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.125288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.125288Z digest=sha256:553e4885141eed5cfdb74ea6545dcf0d38a677413590488b8f30c35bd6dbb778

Observation 1c4c4e5f-7d75-452b-8a3a-e4fc03ffd881 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.207587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.207587Z digest=sha256:1bfb8919363cbfaca606740d199e2d655fb12fb20aa9f528a98f3aed15e81e5b

Observation addd42a4-9d0e-4b7d-883a-740d7ad3fc5d · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.311339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.311339Z digest=sha256:8377261f50dd782518d447980b5976ede033199b3a9a80c2d0ee0289d51aa235

Observation 833b404f-7f94-40e2-aa27-34de5a8f59f4 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Calibrated Self-Rewarding Vision Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.423620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.423620Z digest=sha256:6049729a78970baac848fa753a14f902fcf1feb5a6db440560466ff9b714811b

Pith citing papers

Observation fd00f03b-0e85-4a46-a5f8-0b937a4a4f66 · inbound

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning cites this paper.

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T18:39:18.945567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:39:18.945567Z digest=sha256:adcab1ce2e3ddda272ce74421f3d4540ca9fab02990fae08f9e53cce7e32e277

Observation 9baea096-df66-4209-84a4-cc84c740d227 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.864757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.864757Z digest=sha256:408967954c79f9824e132670257f68700a7c477358418a294be4925f7047c8ce

Observation c1007e17-2e9d-42c4-b6b4-27490a6c1045 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:42.673501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:42.673501Z digest=sha256:1a1ee63b7e2881311e97a7552e628ea573522a05b6e6177b7f66fefbcd07e490

Observation 3997fe41-a744-4773-a272-8bbe15188270 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.447802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:3c26b75cd343f33622840ca077d3dd72b8e7872ea663e4ac1af42ec1b906a0ac

Observation f3d4dad6-963e-43b0-a461-ffe4a4862636 · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:42543fbae8643dafeb660fc0162cf1d0671430ddea9b957873b187701651e3cc

Observation 5d28c6f6-f5b3-443a-99e4-67578dd105ef · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.425487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.425487Z digest=sha256:8e06f1f5efbb2aa30b5c8a22e88e59ce89c20d1c7c43a71f761904df2ab3358f