Spatial-LLaVA, fine-tuned on the SUN-Spot v2.0 dataset with set-of-marks prompts, reports 56.14% accuracy on Visual Spatial Reasoning, about 3.8 points above LLaVA v1.5 13B.
Meyer, Yuning Chai, and Yong Jae Lee
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding
Spatial-LLaVA, fine-tuned on the SUN-Spot v2.0 dataset with set-of-marks prompts, reports 56.14% accuracy on Visual Spatial Reasoning, about 3.8 points above LLaVA v1.5 13B.