Diff-VRD generates visual relation phrases with a diffusion model conditioned on CLIP features, aiming to detect interactions beyond dataset labels and scoring them with text-to-image retrieval and SPICE.
Mining the benefits of two-stage and one-stage hoi detection,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generalized Visual Relation Detection with Diffusion Models
Diff-VRD generates visual relation phrases with a diffusion model conditioned on CLIP features, aiming to detect interactions beyond dataset labels and scoring them with text-to-image retrieval and SPICE.