CAT uses LLM-scored captions to assign each image an 8x, 16x, or 32x compression and trains a nested VAE with variable-length latents, improving ImageNet generation FID and throughput over fixed-ratio baselines.
Taming transformers for high-resolution image synthesis, 2020
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
CAT: Content-Adaptive Image Tokenization
CAT uses LLM-scored captions to assign each image an 8x, 16x, or 32x compression and trains a nested VAE with variable-length latents, improving ImageNet generation FID and throughput over fixed-ratio baselines.