Fine-tuning vision-language models with sentence-formatted outputs instead of tuple outputs improves shape attribute and coordinate prediction for larger models, and scaling the loss on numeric tokens sharpens numerical accuracy.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models
Fine-tuning vision-language models with sentence-formatted outputs instead of tuple outputs improves shape attribute and coordinate prediction for larger models, and scaling the loss on numeric tokens sharpens numerical accuracy.