Pith. sign in

REVIEW 6 cited by

Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.18013 v1 pith:SQKJEUHT submitted 2025-03-23 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelsrewardlvlmsmodelpreferencereinforcementvision-r1data
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Vision-Language Models (LVLMs) typically follow a two-stage training paradigm-pretraining and supervised fine-tuning. Recently, preference optimization, derived from the language domain, has emerged as an effective post-training reinforcement strategy to enhance capabilities of LVLMs. However, constructing high-quality human-annotated preference data and developing robust reward models to mimic these preferences are both costly and challenging. Motivated by this observation, we propose Vision-R1, a novel vision-guided R1-like reinforcement learning algorithm for LVLMs that rewards models with definitive vision feedback. It only leverages curated instruction data, eliminating the need for specialized reward models and handcrafted preference datasets. We incorporate a criterion-driven reward function that further integrates multi-dimensional feedback to evaluate model completions comprehensively based on the vision task logic. Furthermore, we introduce a progressive rule refinement strategy that dynamically adjusts the reward criteria during training, enabling continuous model improvement and mitigating reward hacking. Extensive experiments on both in-distribution and out-of-distribution benchmarks demonstrate that fine-tuning the 7B LVLMs with Vision-R1 achieves consistent performance gains, with even up to 50% improvement and surpassing the state-of-the-art 10x size model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    REVA-PO stabilizes GRPO-style RL for CXR report generation via response-level adaptive KL weights and validation-anchored policy resets, reporting new SOTA BLEU and clinical F1 scores.

  2. Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Surgery-R1 uses supervised fine-tuning and reinforcement fine-tuning to give a surgical visual question answering model chain-of-thought reasoning, improving accuracy and localization on two EndoVis benchmarks.

  3. VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation

    cs.CL 2025-10 conditional novelty 5.0 of 10

    EVisRAG, an evidence-guided observe-record-reason-answer pipeline trained with reward-scoped GRPO, improves multi-image visual QA accuracy by about 19% over its backbone VLM across five benchmarks.

  4. Reinforced Visual Perception with Tools

    cs.CV 2025-09 conditional novelty 5.0 of 10

    ReVPT uses GRPO reinforcement learning with a cold-start SFT phase to make Qwen2.5-VL models call visual tools, improving perception benchmarks over SFT and text-only RL baselines.

  5. Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A 3B VLM trained with GRPO to call a zoom tool improves V*Bench accuracy by 5.7% over its base model but degrades TextVQA and HR-Bench performance.

  6. Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization

    cs.LG 2025-09 unverdicted

    A survey categorizing deep reinforcement learning and direct preference optimization methods for aligning large vision-language models, with tables of studies and datasets and no new experimental result.

Pith tools