Pith. sign in

REVIEW 8 cited by

Taming Transformers for High-Resolution Image Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.09841 v3 pith:CWCQ2R6H submitted 2020-12-17 cs.CV

classification cs.CV
keywords transformershigh-resolutionimagescnnsimagesynthesisbiasinductive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Designed to learn long-range interactions on sequential data, transformers continue to show state-of-the-art results on a wide variety of tasks. In contrast to CNNs, they contain no inductive bias that prioritizes local interactions. This makes them expressive, but also computationally infeasible for long sequences, such as high-resolution images. We demonstrate how combining the effectiveness of the inductive bias of CNNs with the expressivity of transformers enables them to model and thereby synthesize high-resolution images. We show how to (i) use CNNs to learn a context-rich vocabulary of image constituents, and in turn (ii) utilize transformers to efficiently model their composition within high-resolution images. Our approach is readily applied to conditional synthesis tasks, where both non-spatial information, such as object classes, and spatial information, such as segmentations, can control the generated image. In particular, we present the first results on semantically-guided synthesis of megapixel images with transformers and obtain the state of the art among autoregressive models on class-conditional ImageNet. Code and pretrained models can be found at https://github.com/CompVis/taming-transformers .

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents

    cs.LG 2026-07 conditional novelty 6.0 of 10

    FairDiffuseVQVAE reaches state-of-the-art fairness on the standard tabular benchmark (DPR 0.702, EOR 0.686) by uniform protected-attribute sampling at inference, paying ~15 AUC points of utility.

  2. InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames

    cs.LG 2025-10 conditional novelty 6.0 of 10

    An autoregressive transformer with inertial-frame tokenization and geometric rotary positional encoding reports state-of-the-art validity and stability on QM9, GEOM-Drugs, and B3LYP, plus strong functional-group-condi...

  3. Transition Matching: Scalable and Flexible Generative Modeling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Transition Matching unifies flow matching and continuous autoregressive generation as discrete-time Markov processes, with three variants that improve text-to-image quality and speed.

  4. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.

  5. Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A lesion-measurement-conditioned VAR model with a lesion-focused VQVAE achieves best average FID (0.74) for controllable skin lesion synthesis.

  6. StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    A training-free inference-time pipeline uses masked cross-image attention sharing and region harmonization to keep subjects consistent across generated story images.

  7. Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission

    cs.SD 2025-09 conditional novelty 4.0 of 10

    Neural audio codecs match or beat Opus for speaker verification on VoxCeleb1 below 12 kbps and stay within about 1.5 percentage points EER above it.

  8. Latent Sensor Fusion: Multimedia Learning of Physiological Signals for Resource-Constrained Devices

    eess.SP 2025-07 conditional novelty 4.0 of 10

    A single image-pretrained VQ-VAE encoder, applied to spectrograms of six physiological signals, matches a modality-specific fusion baseline on WESAD stress classification while using less compute and memory.

Pith tools