Pith. sign in

REVIEW 3 cited by

Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.02253 v2 pith:VV2JY2O6 submitted 2023-12-04 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords syntheticimagestrainingdatadomaingenerativeimagenetmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advances in generative deep learning have enabled the creation of high-quality synthetic images in text-to-image generation. Prior work shows that fine-tuning a pretrained diffusion model on ImageNet and generating synthetic training images from the finetuned model can enhance an ImageNet classifier's performance. However, performance degrades as synthetic images outnumber real ones. In this paper, we explore whether generative fine-tuning is essential for this improvement and whether it is possible to further scale up training using more synthetic data. We present a new framework leveraging off-the-shelf generative models to generate synthetic training images, addressing multiple challenges: class name ambiguity, lack of diversity in naive prompts, and domain shifts. Specifically, we leverage large language models (LLMs) and CLIP to resolve class name ambiguity. To diversify images, we propose contextualized diversification (CD) and stylized diversification (SD) methods, also prompted by LLMs. Finally, to mitigate domain shifts, we leverage domain adaptation techniques with auxiliary batch normalization for synthetic images. Our framework consistently enhances recognition model performance with more synthetic data, up to 6x of original ImageNet size showcasing the potential of synthetic data for improved recognition models and strong out-of-domain generalization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Per-image LoRA adapters fused at inference time produce synthetic training data that improves few-shot image classification accuracy over existing synthetic-data methods.

  2. Does Feasibility Matter? Understanding the Impact of Feasibility on Synthetic Training Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Feasibility of synthetic images has little effect on fine-tuned CLIP accuracy; the edited attribute (background, color, or texture) matters more than whether the attribute is realistic.

  3. AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Synthetic lecture slides generated by an LLM pipeline improve few-shot slide element detection and text-based retrieval when used as pre-training data.

Pith tools