REVIEW 5 cited by
African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent Large Vision-Language Models (LVLMs) demonstrate impressive abilities on numerous image understanding and reasoning tasks. The task of fine-grained object classification (e.g., distinction between \textit{animal species}), however, has been probed insufficiently, despite its downstream importance. We fill this evaluation gap by creating \texttt{FOCI} (\textbf{F}ine-grained \textbf{O}bject \textbf{C}lass\textbf{I}fication), a difficult multiple-choice benchmark for fine-grained object classification, from existing object classification datasets: (1) multiple-choice avoids ambiguous answers associated with casting classification as open-ended QA task; (2) we retain classification difficulty by mining negative labels with a CLIP model. \texttt{FOCI}\xspace complements five popular classification datasets with four domain-specific subsets from ImageNet-21k. We benchmark 12 public LVLMs on \texttt{FOCI} and show that it tests for a \textit{complementary skill} to established image understanding and reasoning benchmarks. Crucially, CLIP models exhibit dramatically better performance than LVLMs. Since the image encoders of LVLMs come from these CLIP models, this points to inadequate alignment for fine-grained object distinction between the encoder and the LLM and warrants (pre)training data with more fine-grained annotation. We release our code at \url{https://github.com/gregor-ge/FOCI-Benchmark}.
Forward citations
Cited by 5 Pith papers
-
Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models
HiR² extracts coarse-to-fine visual features from LMM layers and regularizes them with Lorentz entailment cones and unit-sphere dispersive loss, improving hierarchical consistency across models and fine-tuning methods.
-
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
Finedefics improves fine-grained image classification in multimodal LLMs by contrastively aligning image, attribute, and category embeddings, though its headline gains are measured against zero-shot baselines rather t...
-
Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
KeCO updates the visual feature keys of a small coreset with all leftover support images, and its diversity-based update outperforms retrieval from the five-times-larger full support set for LVLM in-context image clas...
-
Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models
A prompt-ranking metric that subtracts semantic similarity to confounding classes selects higher-precision prompts for zero-shot vision-language object detection.
-
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.
Discussion (0). Continue with ORCID to comment.