ResGen predicts cumulative vector embeddings of masked RVQ tokens, decoupling generative sampling cost from token depth and improving FID and TTS metrics over autoregressive baselines.
To increase the depth of RVQ, we warm-start from the 4-depth RQ-V AE checkpoint (Lee et al., 2022), excluding the attention layers, and reduce the latent dimension from 256 to
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
ResGen predicts cumulative vector embeddings of masked RVQ tokens, decoupling generative sampling cost from token depth and improving FID and TTS metrics over autoregressive baselines.