A scene-conditioned factored attention module, which multiplies attention weights by the image's predicted scene vector, improves MS COCO captioning scores over an Up-Down re-implementation.
Bottom-up and top-down attention for image captioning and visual question answering
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Scene-based Factored Attention for Image Captioning
A scene-conditioned factored attention module, which multiplies attention weights by the image's predicted scene vector, improves MS COCO captioning scores over an Up-Down re-implementation.