Pith. sign in

REVIEW 2 cited by

Diversity is Definitely Needed: Improving Model-Agnostic Zero-shot Classification via Stable Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.03298 v4 pith:RQPO7T3A submitted 2023-02-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords classificationimagesmodelsapproachdiffusionma-zscarchitecturesdiversity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this work, we investigate the problem of Model-Agnostic Zero-Shot Classification (MA-ZSC), which refers to training non-specific classification architectures (downstream models) to classify real images without using any real images during training. Recent research has demonstrated that generating synthetic training images using diffusion models provides a potential solution to address MA-ZSC. However, the performance of this approach currently falls short of that achieved by large-scale vision-language models. One possible explanation is a potential significant domain gap between synthetic and real images. Our work offers a fresh perspective on the problem by providing initial insights that MA-ZSC performance can be improved by improving the diversity of images in the generated dataset. We propose a set of modifications to the text-to-image generation process using a pre-trained diffusion model to enhance diversity, which we refer to as our $\textbf{bag of tricks}$. Our approach shows notable improvements in various classification architectures, with results comparable to state-of-the-art models such as CLIP. To validate our approach, we conduct experiments on CIFAR10, CIFAR100, and EuroSAT, which is particularly difficult for zero-shot classification due to its satellite image domain. We evaluate our approach with five classification architectures, including ResNet and ViT. Our findings provide initial insights into the problem of MA-ZSC using diffusion models. All code will be available on GitHub.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restoration

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A controlled diffusion model, GenDeg, generates a large dataset of diverse paired degradations that improves out-of-distribution performance of all-in-one image restoration models when used as training data.

  2. Dataset Augmentation by Mixing Visual Concepts

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MVC fine-tunes Stable Diffusion with mixed CLIP caption embeddings to produce in-domain synthetic images, improving classifier accuracy on several benchmarks.

Pith tools