A LoRA fine-tuned Molmo-7B model, trained on synthetic automotive UI data with reasoning and pass/fail evaluation labels, improves visual grounding on a new automotive benchmark and on the external ScreenSpot test.
Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
A LoRA fine-tuned Molmo-7B model, trained on synthetic automotive UI data with reasoning and pass/fail evaluation labels, improves visual grounding on a new automotive benchmark and on the external ScreenSpot test.