Fine-tuning medical vision-language models on the Med-MIM multi-image instruction dataset improves their scores on the authors' multi-image benchmarks, but the held-in benchmark is drawn from the same data used for training.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Medical Large Vision Language Models with Multi-Image Visual Ability
Fine-tuning medical vision-language models on the Med-MIM multi-image instruction dataset improves their scores on the authors' multi-image benchmarks, but the held-in benchmark is drawn from the same data used for training.