Pith. sign in

Paper Citation Record · LEDGER

Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2503.18013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.18013 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:01:36.044543Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.981702Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd411d1f-0a5f-44d6-b5df-ac076a45f325 · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.750683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:f5e6958cf269a16a785fa5367819665a1684b1bf3bec1e2390f0629eec25db63

Observation c26de6c9-a946-4bec-b57f-2df7882f7bce · inbound

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO cites this paper.

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:36.044543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:36.044543Z digest=sha256:fec789dec2660cc32fed552a63a6a3ad498c8d4582c17b143ea30e25781c1a16

Observation 2f88aef6-8697-47bb-a8ff-b2997eefa401 · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.657935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.657935Z digest=sha256:b1af8430e7ae24362e2d0c3d56d6a5902ad9d45f59c2f592a9bb9bf72f0d4e93

Observation 346c0276-220e-4c6d-b86f-e7d4bf69744d · inbound

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning cites this paper.

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.090547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.090547Z digest=sha256:db0c9698b075f22cf3e995b70fba805dbfba178d13d08b6863bbffae352140d8

Observation 850ec3cd-0c61-4c3f-8211-822fc8fc9f39 · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:15.259254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:15.259254Z digest=sha256:b021c25520b80147a2294673d14fddb4c2879dae64959d5bab2cd58b97d16ca6

Observation 9b77ae93-91de-49b5-8e52-1196ea0c1596 · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.037158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.037158Z digest=sha256:067aa104530a6bf6a7a74dc26663e511cd4981f6b49374561d2884d006e5ada2

Observation 0613e258-0573-4692-ae18-f61740a8bef4 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:12.120909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:12.120909Z digest=sha256:6f54e2f07e18501b355c989a8154399c50c9e2ee7a6e087b6048a21a43c2ebe0

Observation b6525d68-4675-45d3-ad5f-a6f602a39e3a · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.736748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.736748Z digest=sha256:ce0b37ca4aa1c4d4c60f01132aa8510c6025aa0e1bd2ac5c953496ed55b93d1e

Observation b19bc3e2-5c75-401a-aca3-6dbb135c814f · inbound

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints cites this paper.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.569467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.569467Z digest=sha256:f99633a7def766b437d935b23ce3a091ea6bb9ffedfd1c0ee6b68619b625d77a

Observation 11926f27-4162-4895-8f2b-c855a4d22f50 · inbound

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning cites this paper.

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:12.695884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:12.695884Z digest=sha256:fcd5ba51617466ce73af68baf7c15b8ab39dae3470196fd426b8e7fff0b9f259

Observation 005b77f1-e017-4840-a79d-34eb69205190 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.012516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:bb4b6819af926613b5ad4789f942546475342c78fac685d7ddadecc2f9802e71

Observation 20453af6-1290-458e-a709-3063cac7be52 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:05.080004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:05.080004Z digest=sha256:f8a57b0b64675b861ad8469de51fe803338788597930214da2d94b6225717c1a

Observation 5eb4008c-f784-4649-a0cc-4dd41817ff9d · inbound

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization cites this paper.

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:11:30.555836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:11:30.555836Z digest=sha256:191654bf64ab61c927945063b1146724ca13e11a6290cbe3c0464c8d7805326c

Observation b41ac04a-da97-4869-a536-9d821f622145 · inbound

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation cites this paper.

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.936784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:42.936784Z digest=sha256:399d7fa1fd504b8df5d252a731ffc5bb8ccbbf4302c630aec1802948b6af08fc

Observation 964003eb-ae4d-4fe1-946d-78a03b8837d5 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.475041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:9a3014cb66713824d79b77153cc9115d219db6ce062833bf2adbbc933b6718f0

Observation 03dde0a9-2f71-4042-b8cc-9b0a3fa53698 · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.494204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:bc9818da2a2fc64025f5642210bbbb9116f14925140779c242353ea45c89827c

Observation 143cbac7-a3aa-4e84-b091-6391208307bc · inbound

Generalization Under Scrutiny: Cross-Domain Detection Progresses, Pitfalls, and Persistent Challenges cites this paper.

Generalization Under Scrutiny: Cross-Domain Detection Progresses, Pitfalls, and Persistent Challenges Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:06:14.056710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:17:55.617945Z digest=sha256:21583da6336e08d3a3f8b896d93aed97ac470dc5e9963a91056c52d51446ca1e

Observation bfeae30e-7bbb-4965-9a04-bb709ce3a7c6 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 192

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.665008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:6f1f3ffcdb1dc0fe2ab051b328d8e07f4f24409963065af86608e3f250e7bc08

Observation 93726af1-dd57-4995-81b6-d035274dbf44 · inbound

Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification cites this paper.

Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:04.959897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:20:44.688117Z digest=sha256:7faa074fd3a4e9e4cc8cffe5cbc06d38ba3950ed0cff921404e0cb460aa9d690

Observation 0863ece9-6c7f-4e32-94d6-292022c50a0d · inbound

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs cites this paper.

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:06.437566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:30:56.297653Z digest=sha256:edc9c4e730f54027f1b1cb093ed886633b422b1c1b0ed63db42a65b96f248b94

Observation 2cbe4ccb-6387-4101-b6b8-869a564221a4 · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:15.641687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:452e9514e0cdf5f964d2bb04af34eb48e4a70d329eaf48e0d916ec186bfda8e8

Observation 84577f29-4403-40ff-9fef-de410433aa4d · inbound

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning cites this paper.

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:09.550329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T14:02:29.480442Z digest=sha256:c806e2a9864ab7e007191ca6d35977b628667a9dcecb448208edee8a6228a302

Observation d19c5210-e34f-47db-85a0-ed0e69f31057 · inbound

Towards One-to-Many Temporal Grounding cites this paper.

Towards One-to-Many Temporal Grounding Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.677234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:11:48.455492Z digest=sha256:4bc85db2749d2f4ab853d04e08c2cd6231ae35c0f508f96ca4b6f7dce73233d2

Observation 578c924b-c8e4-4c09-beec-4cddfe498dc2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 280

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.341513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:4b48debcf0d85f910d39b14c8c12157754c470d4d259add24ade837e7291439d

Observation 90b25e58-57d5-4e74-bbf2-7e484b9c4699 · inbound

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views cites this paper.

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:45.396738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:38:46.044079Z digest=sha256:4013c3fcc177ea006991c2f6bf045e9f8dd313b32f8a81fe12f57c60cbd8913c

Observation a7c47baa-2951-4fd0-b56a-17dfa008c6c9 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.983303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:62a1e1cf1f23604371cdbb1fceb17496b075d30759cba2e996154c474c4b58aa

Observation 082de28b-1b62-4c6b-9689-8b7081ba8bbe · inbound

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation cites this paper.

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T14:00:00.388339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:00:00.388339Z digest=sha256:bfd8863400b172e5c811c9cb8b165d2d0a2807f0ccef35eb7baa20731c81c1a2