Pith. sign in

REVIEW 9 cited by

Return of Unconditional Generation: A Self-supervised Representation Generation Method

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03701 v4 pith:G3UGS56D submitted 2023-12-06 cs.CV

classification cs.CV
keywords generationunconditionallabelsproblemdatafundamentalmethodquality
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Unconditional generation -- the problem of modeling data distribution without relying on human-annotated labels -- is a long-standing and fundamental challenge in generative models, creating a potential of learning from large-scale unlabeled data. In the literature, the generation quality of an unconditional method has been much worse than that of its conditional counterpart. This gap can be attributed to the lack of semantic information provided by labels. In this work, we show that one can close this gap by generating semantic representations in the representation space produced by a self-supervised encoder. These representations can be used to condition the image generator. This framework, called Representation-Conditioned Generation (RCG), provides an effective solution to the unconditional generation problem without using labels. Through comprehensive experiments, we observe that RCG significantly improves unconditional generation quality: e.g., it achieves a new state-of-the-art FID of 2.15 on ImageNet 256x256, largely reducing the previous best of 5.91 by a relative 64%. Our unconditional results are situated in the same tier as the leading class-conditional ones. We hope these encouraging observations will attract the community's attention to the fundamental problem of unconditional generation. Code is available at https://github.com/LTH14/rcg.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Representation Consistency Model via Contrastive Denoising

    cs.CV 2025-01 accept novelty 8.0 of 10

    rRCM, a contrastive denoising pre-training and fine-tuning scheme, gives a single-pass robust classifier that beats diffusion-based defenses on ImageNet and CIFAR-10 while reducing inference cost by up to 85x.

  2. XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative Modeling

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Conditioning each patch's denoising on the full trajectories of earlier patches lets XYZFlow generate ImageNet images with FID 1.22 to 1.63 in only 2 to 5 steps per patch, at 7.2 to 8.5x teacher speedups.

  3. Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A self-supervised diffusion framework with a low-bitrate vector-quantization bottleneck learns disentangled motion and content latents supporting motion transfer and auto-regressive generation.

  4. Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Token-based language-style vision models (VAR, LlamaGen) tolerate quantization better than diffusion models, and a custom TopKLD distillation loss pushes their low-bit scaling roughly one precision level higher.

  5. Learning Visual Generative Priors without Text

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An image-to-image diffusion model pretrained on 190 million unlabeled images serves as a transferable visual generative prior for text-to-image, novel-view synthesis, and image-to-video tasks.

  6. Nested Diffusion Models Using Hierarchical Latent Priors

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Nested diffusion models that generate images by progressively synthesizing hierarchical semantic latents from a frozen pretrained encoder improve image quality over single-level baselines at modest extra cost.

  7. M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    M-VAR decouples VAR's scale-wise autoregressive image generation into intra-scale attention and inter-scale Mamba, achieving 1.78 FID on ImageNet 256x256 with faster inference than VAR.

  8. Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    DUSA adapts classifiers and segmenters at test time by matching their predictions to conditional noise estimates from a pre-trained diffusion model, using a single timestep and active class selection.

  9. GMem: A Modular Approach for Ultra-Efficient Generative Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GMem conditions diffusion models on a fixed bank of DINOv2 features and reports much lower FID at far fewer epochs than SiT and REPA baselines on ImageNet.

Pith tools