CoordTok encodes a 128-frame video into three 2D triplane latents (1280 tokens total) and reconstructs randomly sampled patch coordinates, enabling efficient long-video tokenization and 128-frame generation.
An image is worth 16x16 words: Trans- formers for image recognition at scale
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
CoordTok encodes a 128-frame video into three 2D triplane latents (1280 tokens total) and reconstructs randomly sampled patch coordinates, enabling efficient long-video tokenization and 128-frame generation.