Prompting multimodal LLMs with ground-truth bounding boxes and gaze durations improves chest X-ray report metrics, but the effect is inconsistent and relies on privileged annotations.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation
Prompting multimodal LLMs with ground-truth bounding boxes and gaze durations improves chest X-ray report metrics, but the effect is inconsistent and relies on privileged annotations.