Pith. sign in

REVIEW 17 cited by

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04431 v2 pith:DDTFBREX submitted 2024-12-05 cs.CV

classification cs.CV
keywords infinityautoregressivebitwisescalingmodelingmodelssd3-mediumtokenizer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Infinity, a Bitwise Visual AutoRegressive Modeling capable of generating high-resolution, photorealistic images following language instruction. Infinity redefines visual autoregressive model under a bitwise token prediction framework with an infinite-vocabulary tokenizer & classifier and bitwise self-correction mechanism, remarkably improving the generation capacity and details. By theoretically scaling the tokenizer vocabulary size to infinity and concurrently scaling the transformer size, our method significantly unleashes powerful scaling capabilities compared to vanilla VAR. Infinity sets a new record for autoregressive text-to-image models, outperforming top-tier diffusion models like SD3-Medium and SDXL. Notably, Infinity surpasses SD3-Medium by improving the GenEval benchmark score from 0.62 to 0.73 and the ImageReward benchmark score from 0.87 to 0.96, achieving a win rate of 66%. Without extra optimization, Infinity generates a high-quality 1024x1024 image in 0.8 seconds, making it 2.6x faster than SD3-Medium and establishing it as the fastest text-to-image model. Models and codes will be released to promote further exploration of Infinity for visual generation and unified tokenizer modeling.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    HACK++ is a head-aware KV cache compression framework for VAR models that decouples current-scale attention from historical cache under adaptive per-head budgets to achieve near-lossless generation at 30% attention an...

  2. ImageAttributionBench: How Far Are We from Generalizable Attribution?

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    ImageAttributionBench is a benchmark dataset demonstrating that state-of-the-art image attribution methods lack robustness to image degradation and fail to generalize to semantically disjoint domains.

  3. VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    VARestorer converts a text-to-image VAR model into a fast one-step real-world image super-resolution model via distribution matching distillation and pyramid image conditioning.

  4. Generative Refinement Networks for Visual Synthesis

    cs.CV 2026-04 accept novelty 7.0 of 10

    Hierarchical Binary Quantization plus global refinement AR yields 0.56 rFID reconstruction and 1.81 gFID class-conditional generation on ImageNet, with competitive T2I/T2V at 2B scale.

  5. WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

    cs.CV 2025-03 unverdicted novelty 7.0 of 10

    Text-to-image models show significant limitations in integrating world knowledge, as measured by the new WISE benchmark and WiScore metric across 20 models.

  6. Generative Refinement Networks for Visual Synthesis

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    GRN uses hierarchical binary quantization and entropy-guided refinement to set new ImageNet records of 0.56 rFID for reconstruction and 1.81 gFID for class-conditional generation while releasing code and models.

  7. Progressive Checkerboards for Autoregressive Multiscale Image Generation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A balanced multiscale checkerboard sampling order for autoregressive image generation allows large scale-up factors without quality loss, because only the total number of serial steps matters.

  8. Scalable GANs with Transformers

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A transformer-only GAN trained in VAE latent space with multi-level noise supervision and width-scaled learning rates achieves FID 2.96 on ImageNet-256 in 40 epochs.

  9. A Unified Low-level Foundation Model for Enhancing Pathology Image Quality

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A prompt-guided diffusion model pretrained on 190 million pathology patches outperforms task-specific models across most restoration and virtual staining benchmarks.

  10. Localizing and Mitigating Memorization in Image Autoregressive Models

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Memorization in image autoregressive models sits in early blocks at coarse scales for VAR models and in middle/late blocks for RAR models; halving the flagged neurons' weights cuts extractable images by 65 to 84 percent.

  11. HPSv3: Towards Wide-Spectrum Human Preference Score

    cs.CV 2025-08 conditional novelty 6.0 of 10

    HPSv3, trained on the new 1.08M-pair HPDv3 dataset, reaches 76.9% pairwise preference accuracy on its own test set and Spearman 0.94 against human model rankings, and is used to iteratively refine generated images (CoHP).

  12. Implementing Adaptations for Vision AutoRegressive Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Fine-tuned Vision AutoRegressive models mostly beat a strong diffusion baseline on downstream image generation, but DP fine-tuning yields poor FID scores.

  13. ImgEdit: A Unified Image Editing Dataset and Benchmark

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ImgEdit supplies 1.2 million curated edit pairs and a three-part benchmark that let a VLM-based model outperform prior open-source editors on adherence, quality, and detail preservation.

  14. Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

    cs.CV 2025-05 unverdicted novelty 6.0 of 10

    Mogao presents a causal unified model with deep fusion, dual encoders, and interleaved position embeddings that achieves strong performance on multi-modal understanding, text-to-image generation, and coherent interlea...

  15. Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance

    cs.CV 2025-09 conditional novelty 5.0 of 10

    IGG, an attention-based reweighting of classifier-free guidance, concentrates guidance on important tokens and modestly improves FID/IS in scale-wise autoregressive image generation.

  16. SMPL-GPTexture: Dual-View 3D Human Texture Estimation using Text-to-Image Generation Models

    cs.GR 2025-04 unverdicted novelty 5.0 of 10

    SMPL-GPTexture uses text-to-image generation to produce dual-view human images, aligns them to SMPL meshes via 2D-to-3D recovery, projects colors to UV space, and applies diffusion inpainting to create full high-resol...

  17. MolPIF: A Parameter Interpolation Flow Model for Molecule Generation

    cs.LG 2025-07 conditional novelty 4.0 of 10

    MolPIF generates 3D ligands by interpolating the parameters of Gaussian coordinate and Dirichlet atom-type distributions, reporting stronger docking scores and geometric fidelity than prior flow and diffusion models o...

Pith tools