Pith. sign in

Paper Citation Record · LEDGER

VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2503.10291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10291 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:04.949338Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0357deb5-f945-4f28-954b-f79a9557f837 · inbound

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts cites this paper.

T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:02:43.301816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T08:00:12.781392Z digest=sha256:32f858a4ba3e28b225ca4754c4eaf0e5f7600f43985403384ac59d8639da2695

Observation 0c8b6bab-40fd-4ef5-a69b-4f51776f7c37 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.230946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:5f9ce3da8ae8a819883e6f0ac98920686e3a4433e429b52688ce66960ce3884a

Observation 93397a2e-c44b-4da0-b9ff-85e0fa4b7d70 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.295558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:77b47f8542a7170d702979b7b62835a72dbb3092e90af1fdee0c0645af0503be

Observation 4416468e-560f-42f8-b27d-2bfabe7a4572 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.123645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:c29f6b3b91a78d62c2780722c42e00e366c78639a097a172316da871bc15b88b

Observation 8231b32a-9cff-4b97-b07b-433a37fa33fb · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.103457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:c96412b067fdbcccf3c2926c031f41966ca6584fd87bbe8a36c856b7a3fa5bed

Observation ede0b217-4f58-4d77-bde2-4636fe296336 · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.709903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:8c1acb524034d270e40067f8f33b76094ea663c2209e2b0e484e84c820159284

Observation b65f2783-cb81-4d0b-81c6-d4f0b5b9c89e · inbound

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning cites this paper.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.949338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.949338Z digest=sha256:c399c1a235c9deb583afd4e17c32f2a0be6dc2c3d0da147314ae8b1a20a1356e

Observation e58ef32d-61b3-4a67-bba3-05a0c7d124eb · inbound

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design cites this paper.

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:40.281992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:40.281992Z digest=sha256:b5cad2d7b07554baff2e1a198ed1fab156b3cce3ee5785a96bcb5ba1c48210f2

Observation e6bbca02-a527-4cb8-a7e2-05fd2c2b8596 · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.646946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.646946Z digest=sha256:58f8a374f9c5e07c515a15b338543e48bc83933dc3ae91335b5bd529fb2afa0e

Observation 397007ca-f815-49d8-bfee-c49396f092d9 · inbound

EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique cites this paper.

EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:04:54.825440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:04:54.825440Z digest=sha256:a57d07b2cd6dd901be7ca76f100226f30c8cd57009e1fe1d9a1101b86e2390ad

Observation af3cbdf7-50fb-4977-a426-7995841d4cec · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:16.939733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:16.939733Z digest=sha256:0d34d3aa8de89064f0fb39e53fcc29530a0466a4c1a40498359a77f14fdd5bf6

Observation c67da1dd-31f8-4225-9259-3c6597d91784 · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T12:21:30.970706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T12:20:17.430881Z digest=sha256:f3b6fbb9362f82dd7d03d37845ddf5160fe86768420f4ff568d54ce22be18956

Observation c4e37aac-a1b6-47b5-8351-96dc2335b822 · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T04:27:05.367965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:27:05.367965Z digest=sha256:e41ac0256f46c125bb70b5bb01094e5b3b0764f62ebd11c597f2e0bef826f55d

Observation c733a776-8f7c-4f95-a8c5-460a68bb050a · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:53.654101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:53.654101Z digest=sha256:1c3fe860dfcc009c8efbfe861e92d512ad5dbe78c9903700c64552ab5a62f97e

Observation 09a9147e-039b-4d1b-9b85-d77cc603ab32 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.855293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:a3f00b48e2a5a1a4c7cc4c9553c001cfa13646aa6aa87e3b7172ec5546730ea4

Observation 5e30e8a2-f9a7-480f-ac96-142f40427ac7 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.845687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.845687Z digest=sha256:4224e898c8f0d5c227701a9dcb6b9a9a871ae7dc6f2e5b11f4b71dac4209115f

Observation 9563bb18-9800-47b5-8b4d-38580fbf2181 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:05.053667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:05.053667Z digest=sha256:1e181d5eb3df525ae8c7f1839de03eb24bad30cb79e020b6d8f17256dbd0c614

Observation d8ac7af6-dc72-4af8-b27a-13e3231807a1 · inbound

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance cites this paper.

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T22:36:10.373032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:36:10.373032Z digest=sha256:79195b95125446ff899089a81146a181571ac9b21b756b0aeaa563810d9c20dd

Observation fe9e57f4-5e9e-4184-984c-23255ecbc06e · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:25.222181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:25.222181Z digest=sha256:89c47918b9aacca516450de56eb384aab4997f17da7850a606932892960918d3

Observation b47ab510-e6f2-4a28-9631-23b49267ba75 · inbound

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment cites this paper.

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T19:57:33.167864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:57:33.167864Z digest=sha256:1a1623e0e890b631c38837e2212481157ec28ab2e0d762281189058847b560ed

Observation ab5f3789-1d32-4599-bbbe-d830fb9ca578 · inbound

CoVR-R:Reason-Aware Composed Video Retrieval cites this paper.

CoVR-R:Reason-Aware Composed Video Retrieval VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T21:37:55.887477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:37:55.887477Z digest=sha256:16d62e4f6dd7733c4eb1d76a851ca62e1a26e9b0016272801fd6a707ba4b3514

Observation cff69bee-e71b-48b7-82bc-7adcb4899906 · inbound

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards cites this paper.

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:45:21.123931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:40:52.854907Z digest=sha256:f374f1755d81b23682cf78e3dfb71270f05900a99f7ce441aa9ae04cc2c9880a

Observation d2b4ed60-afb0-46a3-a858-67eabcabdc6e · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:04.091377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:f61f61c4c41446d9b8f0e79bd9f43e9f8cffd593953696c085ebb87d34cef486

Observation c97eab3c-2eb9-442d-ad84-ff0870ecc7c8 · inbound

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization cites this paper.

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:01:04.215642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T23:48:32.613988Z digest=sha256:1d66d3164a74955ccd89ce424b3375c00ae8895865bf0cc7d73c6575baa4672a

Observation ba2121d0-d693-4a90-be7d-127fd5e02659 · inbound

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts cites this paper.

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:02.907342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:08:15.393476Z digest=sha256:d83b2ef674063a96f8c3b7a8f004523b18ce4fe15cefc0d7da7eaa721569ad63

Observation e51d29f8-5501-46de-a428-82aec7ed1639 · inbound

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA cites this paper.

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:56:25.498891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:48:09.071812Z digest=sha256:09cf2b6162f0158219cbd8edf7c65a57f6b73e06395d28d3690160cda4072100

Observation 45c67538-77e6-4ca5-b22c-fa3be027f5ca · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:29.175391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:14:48.918959Z digest=sha256:c8519b6710261d37b151d1e0727fe94326275b52cc6b8e0541b7358052419cd2

Observation 4a07c751-3688-4944-89f9-33c75bf95245 · inbound

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model cites this paper.

Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:48:00.564018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:47:50.595481Z digest=sha256:8f25ccd68f3103ca136553bfc7693c06ab4635c67f0a4bfe7950524d63747e30

Observation c8f15d7c-bc4d-4bc7-9d8e-958a90137a98 · inbound

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning cites this paper.

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:22:50.527662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:20:32.435135Z digest=sha256:2273114e5db83a96826382dd3a89efa92cd5b85890ceb7529e08950fc05ae398

Observation 2a218807-cbe8-47ed-901d-fca5d1fefff5 · inbound

LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling cites this paper.

LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:49:43.921154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:49:20.837352Z digest=sha256:ec6485a5b7ba2ac665b82ed3ac9ef5ef297f6c5f38415e4fbc20e16571a85ecc

Observation 47c4b1de-10bd-4aa9-ab26-e8c771386952 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.786474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:ef96eecff2c25943932df7f526a5532daa1fe4c8ea5e8385df417d558d729826

Observation 8e3a314e-a995-4df0-a4dc-2b262804e1c8 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:43:12.481573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:41:25.205368Z digest=sha256:d556c95c054fd996f95876928c7110655297d87613e7542478c654f7f2beff34

Observation 99005f9f-9931-4576-98e8-0ad675cdd3f2 · inbound

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cites this paper.

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T02:23:31.506677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:23:31.506677Z digest=sha256:91808380239c5bb9087c48afd74c7d7851dfdc81fdc34b56bb8735b462971372

Observation edb38eaa-85f3-4dfe-bd80-7f9e3c602e8b · inbound

CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision cites this paper.

CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:18:05.601029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:14:43.288597Z digest=sha256:4ea229dc128b735c9b4fad67db7d174771c4adf5360eef874bcf420053dcf4c7

Observation 79388d26-6c86-48c3-809a-70c48ed6a507 · inbound

PRO-CUA: Process-Reward Optimization for Computer Use Agents cites this paper.

PRO-CUA: Process-Reward Optimization for Computer Use Agents VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T11:53:24.213227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T11:46:38.569528Z digest=sha256:79be9ab31a8f4b87a7ea36a5456dff846b762af1b32ef30e871d7d9bd27b73a4

Observation 4b3b86c1-8001-40c5-b9d7-3665d5975513 · inbound

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning cites this paper.

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.470385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T23:15:56.598968Z digest=sha256:f8329fb79afd8200e72a4ea818c02c4772b1958437376c906d946a219b871d56

Observation e3088156-84dc-4d24-b31a-df4e8c9fe142 · inbound

ATLAS: Agentic Test-time Learning-to-Allocate Scaling cites this paper.

ATLAS: Agentic Test-time Learning-to-Allocate Scaling VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.018851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T15:27:28.290178Z digest=sha256:02d7f693228d931744f78254c72648100cf4babd048b1e46ce0a7f739e8d99b5

Observation c311efaa-8915-439f-9605-8e9a7ce9317b · inbound

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification cites this paper.

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:46.511202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:38:37.969788Z digest=sha256:e88469053a93a77cb31f4b8c6315bd9445112d621be9b49f5d27d542bb8c701c

Observation ed2bd02a-a54d-46d2-bded-3debe6d62791 · inbound

Improving Multimodal Reasoning via Worst Dimension Optimization cites this paper.

Improving Multimodal Reasoning via Worst Dimension Optimization VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.806679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:55:57.691185Z digest=sha256:c01b75fa5b5b2d2288f742a61e9b76e84aafbb256321f606ae3c41d9831419c9

Observation 8172c0e5-dec6-45cc-be10-b308f5d40ae3 · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:25.318226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:44a27ad4ccb96facc7d96cbb1bdeec51593c772a4f218a017755954cc38c1fee

Observation 88dd4483-2fe0-4c59-aaed-a601380ca739 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.733804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:8d855e113e182f0c74acf53b709aaf55fa95109b77b8c1214042f4047661ab0e

Observation 58d63079-3feb-46bd-9be9-c55760f13cbe · inbound

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ cites this paper.

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T03:00:51.318412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:00:51.318412Z digest=sha256:7bb2e60804f5c03fc76e77802e049923f28b1166c1871ea44918dbb28094d608

Observation 27a4cdd8-8675-4686-940c-1dea5a4259e8 · inbound

Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners cites this paper.

Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T10:52:57.434420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:52:57.434420Z digest=sha256:23467d15ddcda8efb5462c275262e7c3d52de77f1006147258e2fbfa8a38c38f

Observation f9026cc3-10b4-4545-9445-6bb030723ba1 · inbound

Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners cites this paper.

Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T01:24:27.043856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:24:27.043856Z digest=sha256:b0675f0c4c1635c29aa0a1d95a8890e24b746c883089146ee76e76fdd6d8762e