Pith. sign in

REVIEW 14 cited by

Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.18013 v1 pith:SQKJEUHT submitted 2025-03-23 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelsrewardlvlmsmodelpreferencereinforcementvision-r1data
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Vision-Language Models (LVLMs) typically follow a two-stage training paradigm-pretraining and supervised fine-tuning. Recently, preference optimization, derived from the language domain, has emerged as an effective post-training reinforcement strategy to enhance capabilities of LVLMs. However, constructing high-quality human-annotated preference data and developing robust reward models to mimic these preferences are both costly and challenging. Motivated by this observation, we propose Vision-R1, a novel vision-guided R1-like reinforcement learning algorithm for LVLMs that rewards models with definitive vision feedback. It only leverages curated instruction data, eliminating the need for specialized reward models and handcrafted preference datasets. We incorporate a criterion-driven reward function that further integrates multi-dimensional feedback to evaluate model completions comprehensively based on the vision task logic. Furthermore, we introduce a progressive rule refinement strategy that dynamically adjusts the reward criteria during training, enabling continuous model improvement and mitigating reward hacking. Extensive experiments on both in-distribution and out-of-distribution benchmarks demonstrate that fine-tuning the 7B LVLMs with Vision-R1 achieves consistent performance gains, with even up to 50% improvement and surpassing the state-of-the-art 10x size model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    REVA-PO stabilizes GRPO-style RL for CXR report generation via response-level adaptive KL weights and validation-anchored policy resets, reporting new SOTA BLEU and clinical F1 scores.

  2. Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Surgery-R1 uses supervised fine-tuning and reinforcement fine-tuning to give a surgical visual question answering model chain-of-thought reasoning, improving accuracy and localization on two EndoVis benchmarks.

  3. Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A GRPO-based post-training recipe for video LLMs using discrete QA rewards plus continuous temporal IoU rewards with variance-based data selection outperforms SFT and Video-R1.

  4. Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Mixed-R1 uses four reward types under GRPO, including a new bidirectional max-average token similarity (BMAS) reward, and lifts MLLM reasoning benchmarks by 2-5%.

  5. Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage trained MLLM with a perception-to-cognition chain-of-thought and a self-verification RL reward outperforms prior models on video anomaly detection and reasoning.

  6. Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A curriculum reinforcement-learning framework with in-context refocus examples improves camouflaged object classification and detection for a vision-language model, and the authors report surpassing human performance ...

  7. Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

    cs.AI 2025-05 conditional novelty 6.0 of 10

    RL-trained VLMs generalize compositionally far better than SFT-trained ones on synthetic geometry and spatial tasks, but cross-modal combination remains weak, and a caption-before-thinking plus progress-reward recipe ...

  8. R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Share-GRPO creates paraphrased and visually augmented versions of reasoning questions and shares answers and reward signals across versions, improving multimodal reasoning without cold-start SFT.

  9. VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation

    cs.CL 2025-10 conditional novelty 5.0 of 10

    EVisRAG, an evidence-guided observe-record-reason-answer pipeline trained with reward-scoped GRPO, improves multi-image visual QA accuracy by about 19% over its backbone VLM across five benchmarks.

  10. Reinforced Visual Perception with Tools

    cs.CV 2025-09 conditional novelty 5.0 of 10

    ReVPT uses GRPO reinforcement learning with a cold-start SFT phase to make Qwen2.5-VL models call visual tools, improving perception benchmarks over SFT and text-only RL baselines.

  11. Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A 3B VLM trained with GRPO to call a zoom tool improves V*Bench accuracy by 5.7% over its base model but degrades TextVQA and HR-Bench performance.

  12. Reinforcing Video Reasoning with Focused Thinking

    cs.CV 2025-05 reject novelty 5.0 of 10

    A GRPO variant with token-level KL weighting and partial-credit rewards improves video-QA on modified multi-answer benchmarks, but transfer to original single-answer benchmarks is not established.

  13. An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Format rewards help base LLMs more than intermediate retrieval rewards, general-purpose backbones beat reasoning-distilled ones, and stronger search engines stabilize RL training for LLM search agents.

  14. Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization

    cs.LG 2025-09 unverdicted

    A survey categorizing deep reinforcement learning and direct preference optimization methods for aligning large vision-language models, with tables of studies and datasets and no new experimental result.

Pith tools