Pith. sign in

REVIEW 3 cited by

ZeroDiff: Solidified Visual-Semantic Correlation in Zero-Shot Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02929 v2 pith:3EMNA4U2 submitted 2024-06-05 cs.CV cs.LG

classification cs.CVcs.LG
keywords zerodiffclassesgenerativevisual-semanticcorrelationsdatadiffusionlearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient number of samples from seen classes. Our study reveals that a scarcity of seen class samples results in a marked decrease in performance across many generative ZSL techniques. We argue, quantify, and empirically demonstrate that this decline is largely attributable to spurious visual-semantic correlations. To address this issue, we introduce ZeroDiff, an innovative generative framework for ZSL that incorporates diffusion mechanisms and contrastive representations to enhance visual-semantic correlations. ZeroDiff comprises three key components: (1) Diffusion augmentation, which naturally transforms limited data into an expanded set of noised data to mitigate generative model overfitting; (2) Supervised-contrastive (SC)-based representations that dynamically characterize each limited sample to support visual feature generation; and (3) Multiple feature discriminators employing a Wasserstein-distance-based mutual learning approach, evaluating generated features from various perspectives, including pre-defined semantics, SC-based representations, and the diffusion process. Extensive experiments on three popular ZSL benchmarks demonstrate that ZeroDiff not only achieves significant improvements over existing ZSL methods but also maintains robust performance even with scarce training data. Our codes are available at https://github.com/FouriYe/ZeroDiff_ICLR25.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interpretable Zero-shot Learning with Infinite Class Concepts

    cs.CV 2025-05 conditional novelty 6.0 of 10

    InfZSL generates unlimited LLM-based visual concepts for each class, selects transferable and discriminative ones via a concept-entropy ranking, and builds interpretable class embeddings that improve zero-shot recogni...

  2. Improved Feature Generating Framework for Transductive Zero-shot Learning

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A new transductive zero-shot learning framework uses predicted semantic labels as pseudo-conditions to reduce the accuracy drop caused by biased unseen-class priors, achieving small gains over prior methods.

  3. Discriminative Image Generation with Diffusion Models for Zero-Shot Learning

    cs.CV 2024-12 conditional novelty 4.0 of 10

    DIG-ZSL learns a per-class token that steers Stable Diffusion to generate discriminative images for unseen classes, then trains a ZSL classifier on those images.

Pith tools