Pith. sign in

REVIEW 39 references

Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read On a PMC-VQA subset, GRPO-based RL fine-tuning of Qwen2-VL-2B-Instruct outperforms SFT in accuracy, but the study has no error bars and several prose claims conflict with its own table.

arxiv 2505.13973 v1 pith:IGX4FC6H submitted 2025-05-20 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords medicalmodelmodelstuningfine-tuninglearningmllmsreasoning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Medical visual question answering asks an AI to look at a scan or X-ray and answer a clinical question. Most systems are trained by supervised fine-tuning (SFT), which shows the model many question-answer pairs until it imitates them. This paper instead uses reinforcement learning, where the model is rewarded for answers that are marked correct and for reasoning that a pair of biomedical language models deems clinically coherent. The authors compare these two training styles on a small vision-language model, Qwen2-VL-2B, using 10,000 training questions from the PMC-VQA benchmark and 7,000 test questions.

Their main table shows the RL-tuned model reaching 58.04% accuracy, versus 52.00% for full SFT, 45.98% for LoRA SFT, and 46.97% for DPO. Removing the standard deviation normalization in the advantage calculation, a variant called Dr.GRPO, pushed accuracy to 61.09%. The paper also reports that rewarding longer reasoning chains hurts accuracy, and that adding a semantic alignment reward helps accuracy but makes the model's language less fluent according to their perplexity measure.

The findings are plausible but fragile. Accuracy numbers have no error bars, the SFT baselines are not described with enough detail, and some prose conclusions contradict the table, such as claiming better token efficiency when the model produces more thinking tokens. The comparison also lacks any direct matchup against existing medical RL-VQA systems like Med-R1.

Extended reading notes

Core claim

The abstract states: "GRPO-based RL tuning consistently outperforms standard supervised fine-tuning (SFT) in both accuracy and reasoning quality." If true, medical VQA models should be post-trained with GRPO-style RL rather than SFT alone, and reward design choices (semantic alignment, normalization removal) become first-order factors.

Load-bearing premise

The SFT baselines are trained to a fair, competitive standard. The paper cites Lee (2024) for SFT implementation but provides no SFT hyperparameters (epochs, LR, LoRA rank). The LoRA and DPO results (45.98%, 46.97%) are below the zero-shot base model (47.29%), which is unusual and suggests the SFT configurations may be suboptimal. If SFT is undertrained, the headline GRPO-vs-SFT gap is not apples-to-apples.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The empirical claims rest on the GRPO objective (standard math), the assumption that PMC-VQA ground truth and LLM-judged rewards are valid clinical signals, and the evaluation metrics being meaningful proxies. The paper introduces no new theoretical entities. The main free choices are hyperparameters borrowed from prior work and hand-designed reward prompts.

free parameters (5)
  • GRPO hyperparameters (G=8, KL=0.04, LR=1e-6, temperature 1.0, 1500 steps) = G=8, KL=0.04, LR=1e-6, temp=1.0, steps=1500
    Adopted from Zhou et al. 2025b; not swept, so sensitivity of the conclusions to these choices is unknown.
  • Semantic alignment reward prompt and binary threshold = Yes/No from BioGPT/BioMistral, reward 1 or 0
    Hand-crafted prompt (Fig. 1) and binary scoring drive the reported 1.82% accuracy gain; no validation against expert clinical judgment.
  • ECR and CWR length reward scaling = Not specified numerically
    The exact reward function for length is not given; the observed verbosity effects depend on this unspecified choice.
  • PMC-VQA test subset = 7K samples
    Evaluation set is a fixed subset with no reported test-set variance or confidence intervals.
  • Base model choice = Qwen2-VL-2B / Qwen2-VL-2B-Instruct
    All conclusions are specific to this 2B model family; no evidence for larger models.
assumptions (4)
  • standard math GRPO objective and advantage estimator (Eqs. 1-2) are a valid training objective.
    Used as the optimization target; inherited from Shao et al. 2024 without modification.
  • domain assumption PMC-VQA ground-truth answers are correct labels for the subset used.
    Accuracy is measured against these labels without clinician verification on the specific 10K/7K split.
  • domain assumption LLM judges (BioGPT, BioMistral) provide a valid clinical semantic signal.
    The reward is a yes/no on <think> reasoning; there is no evidence that these models' judgment matches expert opinion.
  • domain assumption Evaluation metrics (similarity, perplexity, thinking reward) capture clinical quality.
    These automated metrics are treated as proxies for clinical alignment; no human evaluation is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models." pith.science (2026). https://pith.science/paper/IGX4FC6H

@misc{pith2026250513973,
  author       = {Pith},
  title        = {Pith review of: Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IGX4FC6H}},
  note         = {Machine review of arXiv:2505.13973}
}
read the original abstract

Recently, reinforcement learning (RL)-based tuning has shifted the trajectory of Multimodal Large Language Models (MLLMs), particularly following the introduction of Group Relative Policy Optimization (GRPO). However, directly applying it to medical tasks remains challenging for achieving clinically grounded model behavior. Motivated by the need to align model response with clinical expectations, we investigate four critical dimensions that affect the effectiveness of RL-based tuning in medical visual question answering (VQA): base model initialization strategy, the role of medical semantic alignment, the impact of length-based rewards on long-chain reasoning, and the influence of bias. We conduct extensive experiments to analyze these factors for medical MLLMs, providing new insights into how models are domain-specifically fine-tuned. Additionally, our results also demonstrate that GRPO-based RL tuning consistently outperforms standard supervised fine-tuning (SFT) in both accuracy and reasoning quality.

Figures

Figures reproduced from arXiv: 2505.13973 by the authors.

Figure 1
Figure 1. Illustration of the prompt template used to [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

39 extracted references · 7 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Alisson Azzolini, Hannah Brandon, Prithvijit Chattopadhyay, Huayu Chen, Jinju Chu, Yin Cui, Jenna Diamond, Yifan Ding, Francesco Ferroni, Rama Govindaraju, and 1 others. 2025. Cosmos-reason1: From physical common sense to embodied reasoning. arXiv preprint arXiv:2503.15558

  4. [4]

    Liang Chen, Lei Li, Haozhe Zhao, Yifan Song, and Vinci. 2025 a . R1-v: Reinforcing super generalization ability in vision-language models with less than \ 3. https://github.com/Deep-Agent/R1-V. Accessed: 2025-02-02

  5. [5]

    Xiwen Chen, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Hao Wang, Haiyu Wu, Huayu Li, Aristeidis Sotiras, Yalin Wang, and Abolfazl Razi. 2025 b . Dra-grpo: Exploring diversity-aware reward adjustment for r1-zero-like training of large language models. arXiv preprint arXiv:2505.09655

  6. [6]

    Kanzhi Cheng, Yantao Li, Fangzhi Xu, Jianbing Zhang, Hao Zhou, and Yang Liu. 2024. Vision-language models can self-improve reasoning via reflection. arXiv preprint arXiv:2411.00855

  7. [7]

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, and 1 others. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1--53

  8. [8]

    Matthew Chung. 2025. Training a qwen 2.5 model for medical reasoning with grpo: A tutorial and “aha!” moment. Accessed: 2025-05-16

Show all 39 references
  1. [9]

    Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu. 2024. Insight-v: Exploring long-chain visual reasoning with multimodal large language models. arXiv preprint arXiv:2411.14432

  2. [10]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  3. [11]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, and 1 others. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948

  4. [12]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, and 1 others. 2022. Lora: Low-rank adaptation of large language models. ICLR, 1(2):3

  5. [13]

    Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao, Zheyu Ye, Fei Zhao, Zhe Xu, Yao Hu, and Shaohui Lin. 2025. Vision-r1: Incentivizing reasoning capability in multimodal large language models. arXiv preprint arXiv:2503.06749

  6. [14]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L \'e lio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Tho...

  7. [15]

    Komal Kumar, Tajamul Ashraf, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, Phillip HS Torr, Salman Khan, and Fahad Shahbaz Khan. 2025. Llm post-training: A deep dive into reasoning large language models. arXiv preprint arXiv:2502.21321

  8. [16]

    Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024. Biomistral: A collection of open-source pretrained large language models for medical domains. arXiv preprint arXiv:2402.10373

  9. [17]

    Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, and Xiaofeng Yang. 2025. Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models. arXiv preprint arXiv:2503.13939

  10. [18]

    Yuwon Lee. 2024. https://github.com/2U1/Qwen2-VL-Finetune Qwen2-vl-finetune

  11. [19]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Advances in neural information processing systems, 36:34892--34916

  12. [20]

    Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. 2025. Understanding r1-zero-like training: A critical perspective. arXiv preprint arXiv:2503.20783

  13. [21]

    INGIN LLMS. 2025. Demystifying long chain-of-thought reason. arXiv

  14. [22]

    Zhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin, Liang Liu, Hao Wang, Han Xiao, Shuai Ren, Guanjing Xiong, and Hongsheng Li. 2025. Ui-r1: Enhancing action prediction of gui agents by reinforcement learning. arXiv preprint arXiv:2503.21620

  15. [23]

    Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in bioinformatics, 23(6):bbac409

  16. [24]

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36:53728--53741

  17. [25]

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. https://arxiv.org/abs/1707.06347 Proximal policy optimization algorithms . Preprint, arXiv:1707.06347

  18. [26]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300

  19. [27]

    Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang, Yibo Ma, Jiajia Liao, Qiaoli Shen, Zilun Zhang, Kangjia Zhao, Qianqian Zhang, and 1 others. 2025 a . Vlm-r1: A stable and generalizable r1-style large vision-language model. arXiv preprint arXiv:2504.07615

  20. [28]

    Haozhan Shen, Zilun Zhang, Qianqian Zhang, Ruochen Xu, and Tiancheng Zhao. 2025 b . Vlm-r1: A stable and generalizable r1-style large vision-language model. https://github.com/om-ai-lab/VLM-R1. Accessed: 2025-02-15

  21. [29]

    Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, and 1 others. 2025. Kimi k1. 5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599

  22. [30]

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. https://arxiv.org/abs/2409.12191 Qwen2-v...

  23. [31]

    Xiaodong Wang and Peixi Peng. 2025. Open-r1-video. https://github.com/Wang-Xiaodong1899/Open-R1-Video

  24. [32]

    Xiaobo Xia and Run Luo. 2025. Gui-r1: A generalist r1-style vision-language action model for gui agents. arXiv preprint arXiv:2504.10458

  25. [33]

    Ruohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, and Yiming Yang. 2024. Improve vision language model chain-of-thought reasoning. arXiv preprint arXiv:2410.16198

  26. [34]

    Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and 1 others. 2023 a . Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792

  27. [35]

    Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao, Weixiong Lin, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023 b . Pmc-vqa: Visual instruction tuning for medical visual question answering. arXiv preprint arXiv:2305.10415

  28. [36]

    Jiaxing Zhao, Xihan Wei, and Liefeng Bo. 2025. R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning. arXiv preprint arXiv:2503.05379

  29. [37]

    Yaowei Zheng, Junting Lu, Shenzhi Wang, Zhangchi Feng andDongdong Kuang, and Yuwen Xiong. 2025. Easyr1: An efficient, scalable, multi-modality rl training framework. https://github.com/hiyouga/EasyR1

  30. [39]

    aha moment

    Hengguang Zhou, Xirui Li, Ruochen Wang, Minhao Cheng, Tianyi Zhou, and Cho-Jui Hsieh. 2025 b . R1-zero's" aha moment" in visual reasoning on a 2b non-sft model. arXiv preprint arXiv:2503.05132

  31. [40]

    Wenhui Zhu, Xin Li, Xiwen Chen, Peijie Qiu, Vamsi Krishna Vasa, Xuanzhao Dong, Yanxi Chen, Natasha Lepore, Oana Dumitrascu, Yi Su, and 1 others. 2025. Retinalgpt: A retinal clinical preference conversational assistant powered by large vision-language models. arXiv preprint arX...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.