LLaVA-SpaceSGG is a visual instruction-tuned model that uses a new 2D+3D spatial scene graph dataset to improve open-vocabulary scene graph generation and spatial relation accuracy.
Llama 3 model card
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations
LLaVA-SpaceSGG is a visual instruction-tuned model that uses a new 2D+3D spatial scene graph dataset to improve open-vocabulary scene graph generation and spatial relation accuracy.