Pith. sign in

REVIEW 2 cited by

Text Descriptions are Compressive and Invariant Representations for Visual Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04317 v2 pith:5CHVAC6I submitted 2023-07-10 cs.CV cs.LG

classification cs.CVcs.LG
keywords visualdescriptionsfeaturesimageinvariantclasslearningapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern image classification is based upon directly predicting classes via large discriminative networks, which do not directly contain information about the intuitive visual features that may constitute a classification decision. Recently, work in vision-language models (VLM) such as CLIP has provided ways to specify natural language descriptions of image classes, but typically focuses on providing single descriptions for each class. In this work, we demonstrate that an alternative approach, in line with humans' understanding of multiple visual features per class, can also provide compelling performance in the robust few-shot learning setting. In particular, we introduce a novel method, \textit{SLR-AVD (Sparse Logistic Regression using Augmented Visual Descriptors)}. This method first automatically generates multiple visual descriptions of each class via a large language model (LLM), then uses a VLM to translate these descriptions to a set of visual feature embeddings of each image, and finally uses sparse logistic regression to select a relevant subset of these features to classify each image. Core to our approach is the fact that, information-theoretically, these descriptive features are more invariant to domain shift than traditional image embeddings, even though the VLM training process is not explicitly designed for invariant representation learning. These invariant descriptive features also compose a better input compression scheme. When combined with finetuning, we show that SLR-AVD is able to outperform existing state-of-the-art finetuning approaches on both in-distribution and out-of-distribution performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Does VLM Classification Benefit from LLM Description Semantics?

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LLM-generated descriptions improve VLM classification only when selected to discriminate among ambiguous classes, not when simply ensembled.

  2. MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A concept-guided distillation method lets a fundus-image model learn from unpaired OCT scans during training, improving retinal disease classification when only fundus photos are available at test time.

Pith tools