Fine-tuning VLMs on video-tracking conversations with made-up object names teaches them to localize a specific object in a new image from only a few in-context examples.
MiniGPT-v2: Large Language Model as a Unified Interface for Vision-Language Multi-task Learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Teaching VLMs to Localize Specific Objects from In-context Examples
Fine-tuning VLMs on video-tracking conversations with made-up object names teaches them to localize a specific object in a new image from only a few in-context examples.