Pith. sign in

REVIEW 20 cited by

Synthetic Data from Diffusion Models Improves ImageNet Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08466 v1 pith:UYFC2VQR submitted 2023-04-17 cs.CV cs.AIcs.CLcs.LG

Synthetic Data from Diffusion Models Improves ImageNet Classification

classification cs.CV cs.AIcs.CLcs.LG
keywords modelssamplesclassificationgenerativeimagenetx256accuracydata
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Deep generative models are becoming increasingly powerful, now generating diverse high fidelity photo-realistic samples given text prompts. Have they reached the point where models of natural images can be used for generative data augmentation, helping to improve challenging discriminative tasks? We show that large-scale text-to image diffusion models can be fine-tuned to produce class conditional models with SOTA FID (1.76 at 256x256 resolution) and Inception Score (239 at 256x256). The model also yields a new SOTA in Classification Accuracy Scores (64.96 for 256x256 generative samples, improving to 69.24 for 1024x1024 samples). Augmenting the ImageNet training set with samples from the resulting models yields significant improvements in ImageNet classification accuracy over strong ResNet and Vision Transformer baselines.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

    cs.CL 2023-09 unverdicted novelty 8.0

    Promptbreeder evolves both task prompts and the mutation prompts that improve them using LLMs, outperforming Chain-of-Thought and Plan-and-Solve on arithmetic and commonsense reasoning benchmarks.

  2. Active Flow Expansion for Out-of-Distribution Discovery: from Theory to Molecules

    cs.LG 2026-06 unverdicted novelty 7.0

    ActFlow expands the generable set of pre-trained flow models for out-of-distribution molecular and sequence design via active synthetic data generation and verifier feedback, with new statistical guarantees.

  3. S3OD: Towards Generalizable Salient Object Detection with Synthetic Data

    cs.CV 2025-10 conditional novelty 7.0

    A 139k-image synthetic dataset with diffusion- and DINO-derived masks, trained with a multi-mask decoder, improves cross-dataset salient-object detection and reaches state-of-the-art after fine-tuning.

  4. Learning Interactive Real-World Simulators

    cs.AI 2023-10 conditional novelty 7.0

    UniSim learns a universal real-world simulator from orchestrated diverse datasets, enabling zero-shot deployment of policies trained purely in simulation.

  5. MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentation

    cs.CV 2026-06 unverdicted novelty 6.0

    MedDiffuseMix uses classifier saliency maps to restrict diffusion-based mixing to non-diagnostic areas of medical images, yielding accuracy, F1, and AUC gains over standard, Mixup, and diffusion baselines on four publ...

  6. Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems

    cs.MA 2026-06 unverdicted novelty 6.0

    OCL is a governance layer for LLM agents that cuts unsafe executions from 88% to near-zero and raises valid success from 12% to 96% in adversarial buyer-seller negotiations across frontier LLMs.

  7. ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search

    cs.CV 2026-06 unverdicted novelty 6.0

    ROGLE automates region-level supervision via Region-to-Sentence Matching and introduces the P-VLG benchmark to improve fine-grained alignment in text-based person search over CLIP-based models.

  8. What Makes Synthetic Data Effective in Image Segmentation

    cs.CV 2026-05 unverdicted novelty 6.0

    Dense scene composition and instance fidelity in synthetic diffusion images drive better segmentation performance; SENSE framework exploits this to improve models on Cityscapes, COCO, and ADE20K.

  9. LiBaGS: Lightweight Boundary Gap Synthesis for Targeted Synthetic Data Selection

    cs.LG 2026-05 unverdicted novelty 6.0

    LiBaGS scores and selects synthetic data near decision boundaries using proximity, uncertainty, density, and validity, with boundary-gap allocation and marginal stopping to improve training accuracy.

  10. Stylistic Attribute Control in Latent Diffusion Models

    cs.CV 2026-05 unverdicted novelty 6.0

    A technique for parametric stylistic control in latent diffusion models learns disentangled directions from synthetic datasets and applies them via guidance composition while preserving semantics.

  11. All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding

    cs.CV 2026-04 unverdicted novelty 6.0

    A unified synthetic data generation pipeline produces unlimited annotated multimodal video data across multiple tasks, enabling models trained mostly on synthetic data to generalize effectively to real-world video und...

  12. Towards Continual Expansion of Data Coverage: Automatic Text-guided Edge-case Synthesis

    cs.CV 2025-09 unverdicted novelty 6.0

    Automated LLM-based prompt engineering for text-to-image edge-case synthesis improves object detection robustness on the FishEye8K benchmark over naive augmentation and manual prompts.

  13. Masked Language Prompting for Generative Data Augmentation in Few-shot Fashion Style Recognition

    cs.CV 2025-04 unverdicted novelty 6.0

    Masked Language Prompting masks selected words in reference captions and leverages LLMs to produce diverse, semantically coherent completions for style-consistent generative image augmentation without fine-tuning.

  14. From SRA to Self-Flow: Data Augmentation or Self-Supervision?

    cs.CV 2026-07 unverdicted novelty 5.0

    Attention Separation ablations show that gains from SRA to Self-Flow in diffusion transformers arise mainly from noise-dimension data augmentation rather than token-level self-supervision.

  15. AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    AC3S adds a self-supervised visual prompt modulator to ControlNet diffusion and a multi-agent VLM prompt composer to generate photorealistic images with accurate 2D/3D annotations while avoiding over-conditioning.

  16. MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentation

    cs.CV 2026-06 conditional novelty 5.0

    Saliency-guided diffusion mixing that preserves Grad-CAM-highlighted diagnostic regions improves medical image classification accuracy and AUC across four public datasets.

  17. ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search

    cs.CV 2026-06 unverdicted novelty 5.0

    ROGLE introduces automated pseudo region-sentence pairs via RSM and multi-granular learning to boost fine-grained alignment in text-based person search, plus the P-VLG benchmark with over 100k annotated regions.

  18. Personalized Generative Models for Contextual Debiasing

    cs.CV 2026-05 unverdicted novelty 5.0

    DecoupleGen personalizes diffusion models to create images with uncommon contexts for debiasing object recognition, yielding consistent gains on scene classification tasks.

  19. LiBaGS: Lightweight Boundary Gap Synthesis for Targeted Synthetic Data Selection

    cs.LG 2026-05 unverdicted novelty 5.0

    LiBaGS is a lightweight method that picks synthetic data near decision boundaries while checking density and validity to improve training accuracy over standard oversampling or uncertainty sampling.

  20. Class-specific diffusion models improve military object detection in a low-data domain

    cs.CV 2026-04 unverdicted novelty 5.0

    Class-specific diffusion models fine-tuned on 8-24 real images per class generate synthetic data that improves military vehicle detection by up to 8% mAP50 in low-data regimes, with further gains from ControlNet edge ...