Seq-CVAE learns a latent variable for every word position in an image caption, guided by a backward language model, and produces more diverse yet accurate captions than previous approaches.
Mind’s eye: A recur- rent visual representation for image caption generation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning
Seq-CVAE learns a latent variable for every word position in an image caption, guided by a backward language model, and produces more diverse yet accurate captions than previous approaches.