The paper presents GEST, an event-graph representation of videos that is converted automatically into natural language and is also used as a teacher to pre-train end-to-end video captioning models.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
The paper presents GEST, an event-graph representation of videos that is converted automatically into natural language and is also used as a teacher to pre-train end-to-end video captioning models.