GEAR jointly trains VQ tokenizer and AR generator end-to-end via dual hard/soft read-out and representation alignment, achieving up to 10x faster ImageNet gFID convergence than LlamaGen-REPA while generalizing across quantizers and to text-to-image.
iFSQ: Improving FSQ for image generation with 1 line of code
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
ClariCodec applies GRPO reinforcement learning to a 300 bps neural speech codec, using ASR word-error rate as reward to cut LibriSpeech test-clean WER from 4.64% to 3.55%.
citing papers explorer
-
GEAR: Guided End-to-End AutoRegression for Image Synthesis
GEAR jointly trains VQ tokenizer and AR generator end-to-end via dual hard/soft read-out and representation alignment, achieving up to 10x faster ImageNet gFID convergence than LlamaGen-REPA while generalizing across quantizers and to text-to-image.
-
Optimising Neural Speech Codecs for 300bps Communication using Reinforcement Learning
ClariCodec applies GRPO reinforcement learning to a 300 bps neural speech codec, using ASR word-error rate as reward to cut LibriSpeech test-clean WER from 4.64% to 3.55%.