Pith. sign in

REVIEW 5 cited by

All are Worth Words: A ViT Backbone for Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.12152 v4 pith:OYQQRJLP submitted 2022-09-25 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelsu-vitdiffusiongenerationimageu-netclass-conditionalcnn-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision transformers (ViT) have shown promise in various vision tasks while the U-Net based on a convolutional neural network (CNN) remains dominant in diffusion models. We design a simple and general ViT-based architecture (named U-ViT) for image generation with diffusion models. U-ViT is characterized by treating all inputs including the time, condition and noisy image patches as tokens and employing long skip connections between shallow and deep layers. We evaluate U-ViT in unconditional and class-conditional image generation, as well as text-to-image generation tasks, where U-ViT is comparable if not superior to a CNN-based U-Net of a similar size. In particular, latent diffusion models with U-ViT achieve record-breaking FID scores of 2.29 in class-conditional image generation on ImageNet 256x256, and 5.48 in text-to-image generation on MS-COCO, among methods without accessing large external datasets during the training of generative models. Our results suggest that, for diffusion-based image modeling, the long skip connection is crucial while the down-sampling and up-sampling operators in CNN-based U-Net are not always necessary. We believe that U-ViT can provide insights for future research on backbones in diffusion models and benefit generative modeling on large scale cross-modality datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference

    cs.DC 2026-07 conditional novelty 7.0 of 10

    HeteroMosaic uses micro-batching and trace-guided co-optimization to split edge LLM prefill across iGPU and NPU, achieving up to 1.73-2.05x speedups and 45.3% energy reduction on AMD Ryzen AI.

  2. Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings

    q-bio.QM 2026-04 unverdicted novelty 7.0 of 10

    Dual Triangle Attention achieves effective bidirectional attention with built-in positional inductive bias via dual triangular masks, outperforming standard bidirectional attention on position-sensitive tasks and show...

  3. Synthesizing Images on Perceptual Boundaries of ANNs for Uncovering Human Perceptual Variability on Facial Expressions

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Faces generated on ANN decision boundaries raise inter-individual variability in emotion labeling, and fine-tuning on those labels improves both group-level and individual-level prediction.

  4. HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference

    cs.DC 2026-07 conditional novelty 5.0 of 10

    HeteroMosaic co-schedules edge LLM inference across iGPU and NPU via roofline analysis and micro-batches, claiming up to ~2× speedup and ~45% energy reduction on AMD Ryzen AI SoCs.

  5. Improving Joint Embedding Predictive Architecture with Diffusion Noise

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Injecting EDM-style noise into masked-token position embeddings and adding two auxiliary losses improves I-JEPA's linear-probing accuracy by about 1.5 points on ImageNet-1K.

Pith tools