Stack-VS stacks LSTM decoder cells that jointly attend to visual features and semantic attributes to refine image captions stage by stage, with reported gains over 2018 baselines.
Boosting image captioning with attributes,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2019 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
Stack-VS stacks LSTM decoder cells that jointly attend to visual features and semantic attributes to refine image captions stage by stage, with reported gains over 2018 baselines.