VSD2M, a 2.09 million sample bilingual sticker dataset with animated GIFs, plus a Spatial Temporal Interaction layer, improves animated sticker generation over standard video diffusion baselines.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.HC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VSD2M: A Large-scale Vision-language Sticker Dataset for Multi-frame Animated Sticker Generation
VSD2M, a 2.09 million sample bilingual sticker dataset with animated GIFs, plus a Spatial Temporal Interaction layer, improves animated sticker generation over standard video diffusion baselines.