A LoRA fine-tuned Molmo-7B model, trained on synthetic automotive UI data with reasoning and pass/fail evaluation labels, improves visual grounding on a new automotive benchmark and on the external ScreenSpot test.
Website screenshots dataset
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
A LoRA fine-tuned Molmo-7B model, trained on synthetic automotive UI data with reasoning and pass/fail evaluation labels, improves visual grounding on a new automotive benchmark and on the external ScreenSpot test.