REVIEW 3 cited by
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Vision-language models (VLMs) offer a promising paradigm for image classification by comparing the similarity between images and class embeddings. A critical challenge lies in crafting precise textual representations for class names. While previous studies have leveraged recent advancements in large language models (LLMs) to enhance these descriptors, their outputs often suffer from ambiguity and inaccuracy. We attribute this to two primary factors: 1) the reliance on single-turn textual interactions with LLMs, leading to a mismatch between generated text and visual concepts for VLMs; 2) the oversight of the inter-class relationships, resulting in descriptors that fail to differentiate similar classes effectively. In this paper, we propose a novel framework that integrates LLMs and VLMs to find the optimal class descriptors. Our training-free approach develops an LLM-based agent with an evolutionary optimization strategy to iteratively refine class descriptors. We demonstrate our optimized descriptors are of high quality which effectively improves classification accuracy on a wide range of benchmarks. Additionally, these descriptors offer explainable and robust features, boosting performance across various backbone models and complementing fine-tuning-based methods.
Forward citations
Cited by 3 Pith papers
-
DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery
DiSciPLE uses LLM-guided evolution to discover interpretable Python programs that predict geospatial quantities, outperforming black-box deep nets on population density and on out-of-distribution generalization.
-
MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images
A concept-guided distillation method lets a fundus-image model learn from unpaired OCT scans during training, improving retinal disease classification when only fundus photos are available at test time.
-
From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors via LLM-guided Symbolic Reasoning
SymbolicDet adds an LLM-guided evolutionary search over object detector outputs to produce interpretable rules for event recognition, reporting large AUROC gains across fishing, safety, and crowd benchmarks.
Discussion (0). Continue with ORCID to comment.