Pith. sign in

Paper Citation Record · LEDGER

Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2503.18013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.18013 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:54.929294Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.981702Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd411d1f-0a5f-44d6-b5df-ac076a45f325 · inbound

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning cites this paper.

VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:56:07.750683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:56:07.247122Z digest=sha256:25a978f40b804841c045bcc117ebb5e63067b0eaff6d7eea0e1864c410e450dc

Observation 05a4234d-b062-4944-b633-d5c847f3be3c · inbound

An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents cites this paper.

An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:54.929294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:54.929294Z digest=sha256:8a287bb455aa517757957c6957b96a7a105c7bf00bb57971be8c555846c2055f

Observation c26de6c9-a946-4bec-b57f-2df7882f7bce · inbound

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO cites this paper.

R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:36.044543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:36.044543Z digest=sha256:57ff2c4e4d3aef93609d5eb425548cb3b8d310f3ac4d7c5e6b92339d3d0240e2

Observation 2f88aef6-8697-47bb-a8ff-b2997eefa401 · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:04.657935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:04.657935Z digest=sha256:7c490605024bf145f05255e90d525138b3c61a7ae0b0daea1ad34dd542bc5f7e

Observation 346c0276-220e-4c6d-b86f-e7d4bf69744d · inbound

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning cites this paper.

Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:52.090547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:52.090547Z digest=sha256:205e91934cfe396c847a968c99d915faef6976ed0d1cf0cd3e25cdd110478f56

Observation 850ec3cd-0c61-4c3f-8211-822fc8fc9f39 · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:15.259254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:15.259254Z digest=sha256:dc1bba7ac61fba4d7904c04d10cfc8eaf0e3197fb0434bde8e09977bc0356d0c

Observation 9b77ae93-91de-49b5-8e52-1196ea0c1596 · inbound

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models cites this paper.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.037158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.037158Z digest=sha256:ec33743b2eaf08145a0b215787c7c69e3e4b940eecb0e7d8751090c9ecfb6145

Observation 0613e258-0573-4692-ae18-f61740a8bef4 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:12.120909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:12.120909Z digest=sha256:2c52c072c93e293be7dd2f1a971a1b89ff62c245531707bbb1b753f6175a1825

Observation b6525d68-4675-45d3-ad5f-a6f602a39e3a · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:47.736748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:47.736748Z digest=sha256:5868f8b01ed90c7068aa622502c5c54c74f4d6504344d60687d5b74e37730f73

Observation b19bc3e2-5c75-401a-aca3-6dbb135c814f · inbound

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints cites this paper.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.569467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.569467Z digest=sha256:79716b3f0ba5ba0faecccc69eab7b73b253ab28fa1f3af748c0c05395968bfcd

Observation 11926f27-4162-4895-8f2b-c855a4d22f50 · inbound

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning cites this paper.

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:12.695884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:12.695884Z digest=sha256:78dd9f5be0119110fc52acb242cf8a543710e2b33c096db78f498f8f5c7a8633

Observation 005b77f1-e017-4840-a79d-34eb69205190 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.012516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:e9455699dc5c970a546018b7c6c58fdd7a62636f37d4ad46a478b733b7f494cf

Observation 20453af6-1290-458e-a709-3063cac7be52 · inbound

Reinforced Visual Perception with Tools cites this paper.

Reinforced Visual Perception with Tools Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:05.080004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:27:05.080004Z digest=sha256:8f53bf7edbf7a3a7cabfd27884e70146df659cf87e5987870606fdd455796459

Observation 5eb4008c-f784-4649-a0cc-4dd41817ff9d · inbound

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization cites this paper.

Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:11:30.555836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:11:30.555836Z digest=sha256:6413b0f2e869ace69506f18d0b21ed3854a01bbc7f001dbc0ff2797a3e03b4b6

Observation b41ac04a-da97-4869-a536-9d821f622145 · inbound

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation cites this paper.

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.936784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:42.936784Z digest=sha256:c27232b4b292653a960e9d39b5c7a9a96e25e420e76c7a2f2b33b0b2273aad4f

Observation 964003eb-ae4d-4fe1-946d-78a03b8837d5 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.475041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:a5a42f0a72bf8285da9a4c8929839eddccbea4d3be765bc3a1bcdc518fbe295a

Observation 03dde0a9-2f71-4042-b8cc-9b0a3fa53698 · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.494204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:67d7b150b3f4b4218655b80e45929f597c880b6027ff754de16501196480d5b4

Observation 143cbac7-a3aa-4e84-b091-6391208307bc · inbound

Generalization Under Scrutiny: Cross-Domain Detection Progresses, Pitfalls, and Persistent Challenges cites this paper.

Generalization Under Scrutiny: Cross-Domain Detection Progresses, Pitfalls, and Persistent Challenges Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 117

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:06:14.056710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:17:55.617945Z digest=sha256:423cf4908c58613166461c232b0670983372a76598199f3552f73eadd7e5371b

Observation bfeae30e-7bbb-4965-9a04-bb709ce3a7c6 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 192

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.665008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:96f81cf75246e6bc1b3fb08dd3a6dd99e489c191f3c547aec8879514655a1a31

Observation 93726af1-dd57-4995-81b6-d035274dbf44 · inbound

Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification cites this paper.

Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:04.959897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:20:44.688117Z digest=sha256:3b7009f2e0f8f79797c0fa8c7d50e4ab6c70fad8025144e053e786d109d10f36

Observation 0863ece9-6c7f-4e32-94d6-292022c50a0d · inbound

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs cites this paper.

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:06.437566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T14:30:56.297653Z digest=sha256:76502aeec34bd42589e5c5c9467bed2a3a2f07d8b53e18b0d749b16b6112b085

Observation 2cbe4ccb-6387-4101-b6b8-869a564221a4 · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:11:15.641687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:e0f0b66193dcbed11a9c3818c87e7aae1890a6e93106ac986062890b9772f537

Observation 84577f29-4403-40ff-9fef-de410433aa4d · inbound

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning cites this paper.

Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:09.550329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T14:02:29.480442Z digest=sha256:abb3edec94ef0bf5bd901e1bd2116fe76e16eeb8ba66e91555b5a9196e6960f4

Observation d19c5210-e34f-47db-85a0-ed0e69f31057 · inbound

Towards One-to-Many Temporal Grounding cites this paper.

Towards One-to-Many Temporal Grounding Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.677234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:11:48.455492Z digest=sha256:7c5d160a606f3ad6e20dbe2fbf6c094cfc26889070e4ffa781f9e338ad98b6e5

Observation 578c924b-c8e4-4c09-beec-4cddfe498dc2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 280

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.341513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:30f889681ddac1875c8b90da5b642397c3e5d66cb2d603ea3dddb8e0f83fc0a4

Observation 90b25e58-57d5-4e74-bbf2-7e484b9c4699 · inbound

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views cites this paper.

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:45.396738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:38:46.044079Z digest=sha256:58c85b17b43f6cdd53e4b204ef302260698c34757b1b8b135bfb4cc08f2120db

Observation a7c47baa-2951-4fd0-b56a-17dfa008c6c9 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.983303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:d21ae115f10a41d1ac7c5e1bca83a01d276a7a80447aedf3bc96dba4120a2356

Observation 082de28b-1b62-4c6b-9689-8b7081ba8bbe · inbound

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation cites this paper.

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T14:00:00.388339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:00:00.388339Z digest=sha256:7014248526320f2d29c37777c9481fb3f3006fa76379fb172bc0699619b6a57d