The paper shows that mAP-selected semantic features, multinomial scheduled sampling, and a length-modulated training loss improve video captioning numbers on YouTube2Text and roughly match prior state of the art on MSR-VTT.
Long-term recurrent convolutional networks for visual recognition and description,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Semantics-Assisted Video Captioning Model Trained with Scheduled Sampling
The paper shows that mAP-selected semantic features, multinomial scheduled sampling, and a length-modulated training loss improve video captioning numbers on YouTube2Text and roughly match prior state of the art on MSR-VTT.