Training VLMs to explicitly convert images to text before reasoning transfers simple-to-hard generalization from text to image, and this conversion skill can be internalized to keep inference cheap.
• To reason about the query: the model needs to correctly enumerate the attribute values for each image in the query similarly
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
Training VLMs to explicitly convert images to text before reasoning transfers simple-to-hard generalization from text to image, and this conversion skill can be internalized to keep inference cheap.