Pith. sign in

Paper Citation Record · LEDGER

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

As of 7 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 5 inbound Pith citation observations for arXiv:2506.10128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10128 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:40:13.423620Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:24:39.864757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:23:28.446166Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 529e2bc9-2351-42ea-a4c2-476f671be824 · outbound

This paper cites Vqa: Visual question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:18.085418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:06.644920Z digest=sha256:1ec386780a2bd7c0b8923340a41f0b5ce02dd90d6ceaa049f32265be62dea949

Observation b187ba96-3bfb-41b1-907c-2d4870c92b3a · outbound

This paper cites BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.691389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.691389Z digest=sha256:441b142f3a2f12c526d7ea133d91b69efd86e4594222d8c8c8edfdcf1b5088f6

Observation 295ab65c-4f3b-437d-a1cc-28902ee60831 · outbound

This paper cites Qwen2.5-VL Technical Report.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.769409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.769409Z digest=sha256:fb84a82d9a49e1a5ec7dc70472a360a729ed881f9da16ba79bbe0dfae87dbb1b

Observation 0b6f18c6-6c7c-4c55-8a25-776297f6b43e · outbound

This paper cites Improving image generation with better captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.852093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.852093Z digest=sha256:fc928931d6ab8c29854c1c03155d91b88724173deedc229ac55daeac47019c2f

Observation 61c4cd51-1079-4709-8a57-8f561729cab0 · outbound

This paper cites Language models are few-shot learners.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:06.950928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:06.950928Z digest=sha256:271b1ab73d2be24fef899291aebfc29832fb899f86c48245b169b0b798848ace

Observation e148fb5c-6010-4a9d-ac4b-2882ddc5bda8 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision language models with less than \ 3.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs R1-v: Reinforcing super generalization ability in vision language models with less than \ 3

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.814290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.055773Z digest=sha256:fc51ca599e81d01b319b66dffebbc1e5e98151159490f1aadfbc5df627507748

Observation b165f18e-9ec9-4319-b4ae-01800a07b4a2 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Sharegpt4v: Improving large multi-modal models with better captions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.479967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.167181Z digest=sha256:cd2302c48d9d7db234bb63a4f3c62e26e69bcffb46566609830963dae6d697d2

Observation 16c87061-4b08-42b8-9650-ddbf4263dfbf · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.282039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.282039Z digest=sha256:af6e6ca864001b672e01cfc815cb6105ba823e5902682bba60e9964856a9a6d5

Observation 4278674d-e3d4-47ce-97c4-9aa33a023f0a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.352289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.352289Z digest=sha256:92fb7cc6c91d103619125fded9759abb3346478d6dce3e9159a25d241d8c80fe

Observation 0f547de1-4259-4664-ae43-04fc56b9f4fd · outbound

This paper cites Palm: Scaling language modeling with pathways.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Palm: Scaling language modeling with pathways

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.423678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.423678Z digest=sha256:69f0a6ac2145774b1aa2844c5f93f2c91d6bae6343e55b37ef9f80a67a2e61ec

Observation ecfc56bd-d7c3-441d-89cc-0496c50d2888 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.530509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.530509Z digest=sha256:718b37c6edbae25ff459a2bc6479e16f278c2183e9c4cceb8c0a8307b2da191a

Observation dbfc9be6-f678-4ffe-b456-45023e87f0ef · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.610397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.610397Z digest=sha256:e41db62c7fff062723b22c8778d180818d740f402d7f8faeafc094a02dbf3fe1

Observation 6d3f5767-fb1f-4614-8730-bdcd5fdd560e · outbound

This paper cites Enhancing large vision language models with self-training on image comprehension.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Enhancing large vision language models with self-training on image comprehension

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:17.196649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.695543Z digest=sha256:41eba342edf8d851cdbd397c2cfead13fd8602b7a9b0f3efbdaab02d99e0842c

Observation c9ae7c9d-2336-4ce8-a60a-05d82a6a63b4 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.764980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.764980Z digest=sha256:9ab6b19b1d5aaff85a731aa7c981209858231e08bb602a7ecda4b8ff184d8c7c

Observation 0af995ae-2b43-48e2-804f-57696d41c56e · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Virtex: Learning visual representations from textual annotations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.922173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:07.899067Z digest=sha256:395e30755538b96b97e8357ee9b0b3ae2b3775f120fdde50efa5d6a424230e68

Observation 22bcc49e-abe1-4239-a509-8b6c2363d5ac · outbound

This paper cites ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:07.980626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:07.980626Z digest=sha256:324c6a46373e8c07d53497195b1b6bd3a3a98f65b29c46f961aecf5ecd0eb26d

Observation 57b69c66-5c8e-451d-b509-23db3c4f7fb6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.094721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.094721Z digest=sha256:381b99af6cd7a77b70cca1b9a359604d4b307630161ce03704cbafc31822e01b

Observation 9024f7a6-09d5-44b8-a2c3-b8d1e33ac7af · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.211304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.211304Z digest=sha256:4e0cdb828db7682d539a7fc974ca28e4b2920ecff1b60ce5b384bda14be7aff0

Observation 4a3b41b7-8ec1-4758-af92-85adf205d4fd · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Measuring mathematical problem solving with the math dataset, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.696254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:08.309188Z digest=sha256:3032f5c2bd45602196ea9d225dfe1a4f2a0a4b0c51c61232efbf36b8d57a7758

Observation a280af47-a3c9-45c7-aff0-7c06e02d297f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.388160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.388160Z digest=sha256:0d7b9bda73b4dff09f643051c54feed7ff80c53d13dfcb5a21a937b74ac4b096

Observation d93473b7-9c98-41ff-b8f5-0e2bc50c2d90 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.503288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.503288Z digest=sha256:aa6aec712ad767fb903016dc40c730a6995f3386717e4039b34ff8833abc4df1

Observation ff924b2c-8786-46bf-8965-ee35e28c2680 · outbound

This paper cites GPT-4o System Card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.599239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.599239Z digest=sha256:3cd5618bd2f02579bf2892201e43423ab3ccc0869a96e6916f06eb061a43407e

Observation 55ce0b4f-f7f3-49ef-b31a-191a7cb46f55 · outbound

This paper cites OpenAI o1 System Card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OpenAI o1 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.696397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.696397Z digest=sha256:29de4685abdcdc0e190522beee399abe5c8c14cf122ce4f766fdd5a3b5930471

Observation 01e4d8dd-532b-4c51-b697-43c1f4837be4 · outbound

This paper cites Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.808188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.808188Z digest=sha256:0a163a501a8c00daabd20d55979fdf6af60e5537cd61caa1dfd32b54fd122a2a

Observation 4fcdadb3-48e7-4df2-a402-d1309441b60a · outbound

This paper cites Large language models are zero-shot reasoners.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Large language models are zero-shot reasoners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.912895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.912895Z digest=sha256:14241288fcd8cd94fffc2eb48232e76a8ce9adad5f95953af31e0885aa98161b

Observation e728da3d-5c57-40c5-aa2b-ff39f97320fe · outbound

This paper cites Mawps: A math word problem repository.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mawps: A math word problem repository

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.994633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.994633Z digest=sha256:c0e8a1150f344f7ae9989202f393bf343c24f725a8ce52dc2d5c0e004854ac6e

Observation 30fd708f-6fe5-4f35-a70a-d1d0c37805a5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.073387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.073387Z digest=sha256:356e1b753a7883f97c98708911db253146220873c6f6a34f292f59e52ff18eca

Observation 38288d0a-69ec-4b66-b564-7c40f6460735 · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose assistants.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Multimodal foundation models: From specialists to general-purpose assistants

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.437363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.168271Z digest=sha256:e63bd9567c79d2865565bd5a948c9586ca55da7737696f4f3e59295674ef606d

Observation 9f99e18b-c7a3-42f5-8ecf-e7e4d906f080 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Describe Anything: Detailed Localized Image and Video Captioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.244875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.244875Z digest=sha256:ef6a7abd64dd0b8f123367c6ae175664986176b3ab30f4e90c688822540be374

Observation 95ca5810-2016-4a10-94ba-bc9b988f8388 · outbound

This paper cites Let's Verify Step by Step.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Let's Verify Step by Step

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.313417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.313417Z digest=sha256:762b43c89dcdc0013d0d2c7a483b46643bfb3dd8ff2ef835651713890de0fff1

Observation 41609576-8914-4c4a-8c88-98c69927a580 · outbound

This paper cites Visual instruction tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Visual instruction tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.436777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.436777Z digest=sha256:b6f8592b4b77bdee5ec8ca6d8ab279919520873b5d4716fa032d671d6cbfa02b

Observation 290e02e6-bfd3-4282-b57a-f67afa3e8367 · outbound

This paper cites Improved baselines with visual instruction tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:16.179830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.523923Z digest=sha256:c62af38d9ee29b2b83e576a45eb36ab3fc093f0fc5cd8e2838d3e251a36dd63f

Observation f4f3460a-ac9d-45af-b995-8b5b62531ca1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.949313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:09.582883Z digest=sha256:7a409e68942365ea3faf95facd117b9c3e38754eaf1e6c6dce91f73656131eb5

Observation a16ab06f-44fb-40d8-819f-5b49a4350f77 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.661306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.661306Z digest=sha256:8e499d7af8aa7318d6c3ac9993b1b6f2d63187ac95d67aeb2bb3faae8c44ade6

Observation f3d4d0d7-fb17-48da-bc52-bc06a535396e · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.721963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.721963Z digest=sha256:678944d877bbab708e8e2ab8614217ef4d199dbe4c3b967ed31bfb34cbb91acb

Observation 6bb0b974-fd24-4e7b-bed7-a382064bb313 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:09.920757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:09.920757Z digest=sha256:b862ec11cb13f7f977de5a270701d4e9a70ff341243178a008c365371c08210b

Observation d39a454f-ee62-445b-b7b6-4257d893cb8e · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.021002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.021002Z digest=sha256:673b9942c5a85b44003c5648f5f7224d6a50ae15c47c7c812fe7927b174d451e

Observation 722d67f4-93c7-460a-a6be-ff80c6b107ed · outbound

This paper cites s1: Simple test-time scaling.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs s1: Simple test-time scaling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.113399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.113399Z digest=sha256:7e101613349479e8a043b83d02e6332ab641f76253c1e5d9a4ba548546dfc9e4

Observation 25071772-0c3a-415c-822f-7b225daad8ec · outbound

This paper cites Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.179618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.179618Z digest=sha256:8568382588b691a793946803a2c4ab68b73f6d1a5b0e1126bd121a8dd295f3c3

Observation 2adb2fad-fb06-4f7e-8be5-9b1f66733eff · outbound

This paper cites American invitational mathematics examination (aime) 2024: Competition problems.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs American invitational mathematics examination (aime) 2024: Competition problems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.691239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.250543Z digest=sha256:92ffaea92d84ec33c73da977a8840c99e6afb9cc206018be1d14462b8eafce4b

Observation 37a766c3-8881-48b5-811e-16d0fdcefe63 · outbound

This paper cites Gpt-4v(ision) system card.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Gpt-4v(ision) system card

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.331609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.331609Z digest=sha256:0b83188c79beed66208df41af610910800a35220417ff9065ff1ba721acc0eba

Observation 1caae194-1900-484a-bbe1-021d5c4fa937 · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Are NLP Models really able to Solve Simple Math Word Problems?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.409957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.409957Z digest=sha256:7cd1c5a3d9cfa7c4cc7b43665736893de2af5f2e6de62dd249e5288985599938

Observation daf7b12b-f92d-41f2-a761-5b29de60949c · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.484752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.484752Z digest=sha256:2d691a1e88479abe6c7efec07f266ed644313bd68697a952ee24d82c14b1b2d9

Observation ff5c66e0-b5f7-4028-9bfd-dc9b592fb697 · outbound

This paper cites Connecting vision and language with localized narratives.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Connecting vision and language with localized narratives

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.460773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.569209Z digest=sha256:4f8199973eba7d4b2e54e6415a96975548ccd20cf2237b8fb7053929994cf9ed

Observation b8f3a763-69aa-4b6a-a1fc-332c2d1c65c5 · outbound

This paper cites Vision language models are blind.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Vision language models are blind

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.655345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.655345Z digest=sha256:5f9fe48974f75734c4a5adb4af4fce04a4dbff7412c6a9cfd46273eeadc74c8f

Observation 6e65327d-6dad-4ae9-a7e2-549795a65948 · outbound

This paper cites Object Hallucination in Image Captioning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Object Hallucination in Image Captioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.721223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.721223Z digest=sha256:2786d3feaf5fb2d43a1e427dc6c1e6da6410a5aca60d50ec31e91cc6b88223b7

Observation d6973e5d-3d08-4214-bbe4-a5764c3ead37 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.790536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.790536Z digest=sha256:38f3e61c7788ff06c919f2cc46dbb7c591a28b6aeb40e4b40e56ef431b2ce571

Observation 5e318aa9-e996-4b3e-b39d-12ec4690f758 · outbound

This paper cites Learning visual representations with caption annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Learning visual representations with caption annotations

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.227903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:10.860273Z digest=sha256:66f3eb0df2c50005d32e2c2f906354b04c089a2e481d152b61c2c2277784027a

Observation 5d4ac2c6-bece-4c01-bdd7-b0b8913054f9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:10.915730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:10.915730Z digest=sha256:7f2b241ce6da8f47d01820b8221fc6311ff97d0a526491d635502f77296b8d13

Observation 7de2822a-dbea-45c7-869b-77e795584691 · outbound

This paper cites Towards vqa models that can read.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Towards vqa models that can read

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.020825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.020825Z digest=sha256:364ddacdf9b9bde73594d5aebee322da5f6e8f3d493e3248f9aaca99bbbc692d

Observation 269f59d6-2c04-4850-a77f-cac1edfdf181 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.110047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.110047Z digest=sha256:62007993c3e3fb61879cdba36b415df65c149605bd573d27a405f1889f55f71e

Observation 60076ad4-8eca-49d0-9755-2c0751f93e70 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.186415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.186415Z digest=sha256:9bdd578879cbe5a39ff5837d3cf3b7be9c62110d075ce801e2a33c782d48f1d9

Observation 4d1d6408-ca4f-4847-91fc-9ff115899846 · outbound

This paper cites Image captioners are scalable vision learners too.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Image captioners are scalable vision learners too

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:15.018460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:11.255836Z digest=sha256:182de2ca9da175cce94d3632f291e541ec0251281dd282e5efe261b3b72e9a2b

Observation b64bb195-56a1-4555-9596-aaea9b217669 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Solving math word problems with process- and outcome-based feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.367240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.367240Z digest=sha256:69bbc2b81763300152af797cd36657235cd8ff61c035276bf407d707f1ceb2d5

Observation a3c4f7f1-599f-46c6-8676-c515b928c0cb · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.443100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.443100Z digest=sha256:d20784b293f2c9cd84853b96fc46851a0bcdfe4958a7b7b49a428adf84bebfea

Observation 8964877a-7d06-4be7-a438-7545a15f4629 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Measuring multimodal mathematical reasoning with math-vision dataset

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.511675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.511675Z digest=sha256:6c8226dfa189dea05e9aa620ae264a5e1803588ff44d465186d2aa370ea18135

Observation 6b811447-8545-46ff-98c8-3544f21aa077 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.584688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.584688Z digest=sha256:bc377f002d25e97ca033adbf8e2fd503af3bd02749fe8350b12e6a065b662072

Observation 677d734b-80d1-4bf2-b4ce-31d3dbac3d35 · outbound

This paper cites Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.660360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.660360Z digest=sha256:4936f183da24fe395689194cd7e3fa762a826bb48369f58ff11ce2dd89630f35

Observation d0da67b2-01d5-4d05-a158-b71c70f17dfa · outbound

This paper cites Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.741673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.741673Z digest=sha256:9bf35c22838532a1d8fe461d9a1fd6bfa75861cf629f576d26ee8b661d9dd658

Observation a9559091-c7ca-459c-8da7-146301ceae8d · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.821726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.821726Z digest=sha256:d84603ac8df627bb3d6c7d8c3ae9bdb0dc9fe1880032c3887817c1fa77f73779

Observation 2110b634-0937-40a8-9cef-88c00df6502c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:11.940289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:11.940289Z digest=sha256:cb1f7814447d4fc48cd4678e4801abc0284e775f190969f7988aac0f9d8cb8de

Observation 3b43b8ad-bc20-443f-a553-8b2bf2522a54 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.800291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:11.995421Z digest=sha256:3ef4660e9734cfd31480e01ef7207f029316074c2ebe4f2b84b96f956a323443

Observation 3b1d654c-9cee-49bc-a702-9e6a175e7a8f · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.078191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.078191Z digest=sha256:7b33179439ba95dbaf5ca81e11365620412840121b26bfc18cbf87fa5285f223

Observation 54962dd3-54f6-4f1b-b9d8-8513fe06f20b · outbound

This paper cites Grit: A generative region-to-text transformer for object understanding.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Grit: A generative region-to-text transformer for object understanding

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.545092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:12.170101Z digest=sha256:d6cd0f93cbe5d0361fc210393eea409bb62037462a1e45a982850273964e0704

Observation ecbdf432-dae1-4225-80a3-8e3fafb23c0e · outbound

This paper cites Mind's eye of llms: Visualization-of-thought elicits spatial reasoning in large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mind's eye of llms: Visualization-of-thought elicits spatial reasoning in large language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.273757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:12.229035Z digest=sha256:5d14c2ac67baac66a5d73bf81740bd2b38b7c9202c90c0c71212f1677085396b

Observation 52378c17-6432-4034-b1cb-652c1758abaf · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.319293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.319293Z digest=sha256:2a32dea8195a1575e24fdc4a903ff4249267ac6a05163bd67c6b064fa5f55600

Observation b15cf5c1-0f16-4079-9e14-187d98653191 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.386906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.386906Z digest=sha256:a96b96d90477cfabb259d61167e00b0b1c0c4608e666cfcd598792d214d5585e

Observation 8006c40d-0ba6-455c-9cb8-4c07c34460d3 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.468494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.468494Z digest=sha256:1bccc1be52bedb828c07095b53929d7458664ccf84c1f61afde47d26a34ff93d

Observation e67b9949-c7f6-4feb-aa0c-0f818da0e6db · outbound

This paper cites Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.565375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.565375Z digest=sha256:3ba682b8a095fab48ccd5c45970a8c5ac420f1885091848eaa31d09c3bb677e6

Observation 8d3a5b4f-6365-49df-ad0d-e2fd06af8d16 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Tree of thoughts: Deliberate problem solving with large language models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.641376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.641376Z digest=sha256:80fa9ef9a6da0ee60405bec12d2ac5572c143a26901c97c6ed032b4fc7fa106d

Observation 8e004ac8-a322-4de4-a87a-5b0d3bacc28b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.722299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.722299Z digest=sha256:6297f939cad37348b0ca57a1c6f6bfa2029680f796e93246075a57d67607647e

Observation 1764ca9d-63de-47ee-a84e-1487f08812a7 · outbound

This paper cites MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.782437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.782437Z digest=sha256:1de9ba1f3e2657ba1ee7684b5b808d986703e7bb47a7e2c4224c7deb1b8f4ad9

Observation 547fe278-652e-451d-879f-2af3e13b2ae0 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.843716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.843716Z digest=sha256:ed61e21eba4c6599d43e048870bf4a93e5cc61b56ce52ae466a7292bcea8966c

Observation 0e7a1d09-8551-47e3-bb06-3a9c18d20f65 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.908645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.908645Z digest=sha256:65c2c3150e8d09dc72e268042fd37c372d92cdf7c4aa75593bc5c253f25bbd56

Observation da366a1c-b9f7-4218-9d91-626a390921c3 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Improve Vision Language Model Chain-of-thought Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:12.982412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:12.982412Z digest=sha256:fb81d773150fb3b998c32ac0f7e3f07f952cac7392b20c15a2901661fd6c68f8

Observation 825be651-6293-4306-aa15-83976889b032 · outbound

This paper cites Multimodal chain-of-thought reasoning in language models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Multimodal chain-of-thought reasoning in language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:40:14.087716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:40:13.071416Z digest=sha256:b1239a60b86c05d95b646dc71f7fef2f3a8cd1f3748c5b1b0062de7944712014

Observation 69ab859d-d281-4114-ba82-84a7bf3038d7 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.125288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.125288Z digest=sha256:ceb84d034d1ad35815f4d6d94955de661258c4d407f429d78bf775cddf984a0e

Observation 1c4c4e5f-7d75-452b-8a3a-e4fc03ffd881 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.207587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.207587Z digest=sha256:e378920ae6dc27f4b73ef3dba238a787c344a61867fe4fb1f3aab3790d886bdf

Observation addd42a4-9d0e-4b7d-883a-740d7ad3fc5d · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.311339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.311339Z digest=sha256:d6ea91aaa4ffc66574d6df9e0fbac90d371293c0f65c42d1f5e50ab4a3688ff6

Observation 833b404f-7f94-40e2-aa27-34de5a8f59f4 · outbound

This paper cites Calibrated Self-Rewarding Vision Language Models.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Calibrated Self-Rewarding Vision Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:13.423620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:13.423620Z digest=sha256:a3fb5f0fce4cf08b9b0a8c41467118a1be89bec60df13ea4a20b8a2a2f13ce8e

Pith citing papers

Observation 9baea096-df66-4209-84a4-cc84c740d227 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.864757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.864757Z digest=sha256:3bdd6e7e3414665e8df9c762c846a862d4e967917e46b664e4765b84f9aa190f

Observation c1007e17-2e9d-42c4-b6b4-27490a6c1045 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:42.673501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:42.673501Z digest=sha256:1a8d5961d6d61d6dcd40d18bc7f0469c9f0f2f9c1521e4cfa3ca996580118c45

Observation 3997fe41-a744-4773-a272-8bbe15188270 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.447802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:4fc3ef012752991cee5220957a885bf18f7169229c657b8196d6c851bdff431e

Observation f3d4dad6-963e-43b0-a461-ffe4a4862636 · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:f6580cd3934ad376c0f2d455cf1fc5695430d246155ddcbb728b15e77784881e

Observation 5d28c6f6-f5b3-443a-99e4-67578dd105ef · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.425487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.425487Z digest=sha256:9085b38501ec4422db9976ed29485caf1d1c47c819e51dc2c8b6c4c5fe7536de