A confidence-aware reward inside GRPO, called CAR, jointly improves diagnostic accuracy and confidence calibration of a 7B medical VQA model on VQA-RAD, SLAKE, and PathVQA.
In: International Conference on Machine Learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
A confidence-aware reward inside GRPO, called CAR, jointly improves diagnostic accuracy and confidence calibration of a 7B medical VQA model on VQA-RAD, SLAKE, and PathVQA.