Multi-task LoRA fine-tuning with Grad-CAM grounding and terminology-free descriptions raises small-VLM GI VQA accuracy and implicit answer-to-region alignment on in- and out-of-distribution data.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
Multi-task LoRA fine-tuning with Grad-CAM grounding and terminology-free descriptions raises small-VLM GI VQA accuracy and implicit answer-to-region alignment on in- and out-of-distribution data.