G2 generates a location-free scene graph from an image and feeds it, with confidence-based token weighting, into an LLM to produce visual commonsense answers and explanations.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
G2 generates a location-free scene graph from an image and feeds it, with confidence-based token weighting, into an LLM to produce visual commonsense answers and explanations.