Training a scene graph model on LLM-generated relationship labels, with iterative self-refinement, improves mean recall on a custom Visual Genome benchmark, including predicates absent from human annotations.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
Training a scene graph model on LLM-generated relationship labels, with iterative self-refinement, improves mean recall on a custom Visual Genome benchmark, including predicates absent from human annotations.