REVIEW 2 cited by
Zero-Shot Learning -- A Comprehensive Evaluation of the Good, the Bad and the Ugly
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Due to the importance of zero-shot learning, i.e. classifying images where there is a lack of labeled training data, the number of proposed approaches has recently increased steadily. We argue that it is time to take a step back and to analyze the status quo of the area. The purpose of this paper is three-fold. First, given the fact that there is no agreed upon zero-shot learning benchmark, we first define a new benchmark by unifying both the evaluation protocols and data splits of publicly available datasets used for this task. This is an important contribution as published results are often not comparable and sometimes even flawed due to, e.g. pre-training on zero-shot test classes. Moreover, we propose a new zero-shot learning dataset, the Animals with Attributes 2 (AWA2) dataset which we make publicly available both in terms of image features and the images themselves. Second, we compare and analyze a significant number of the state-of-the-art methods in depth, both in the classic zero-shot setting but also in the more realistic generalized zero-shot setting. Finally, we discuss in detail the limitations of the current status of the area which can be taken as a basis for advancing it.
Forward citations
Cited by 2 Pith papers
-
ZoRI: Towards Discriminative Zero-Shot Remote Sensing Instance Segmentation
ZoRI combines CLIP text-channel selection, partial fine-tuning, and a pseudo-label cache bank to segment unseen aerial classes, but the cache bank is seeded with the model's own test-set predictions.
-
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
Knowledge-CLIP distills Llama 2 embeddings into CLIP and uses k-means soft concept labels to slightly improve CLIP text and image encoder scores on three benchmarks.
Discussion (0). Continue with ORCID to comment.