Pith. sign in

REVIEW 5 cited by

African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14496 v1 pith:7J4GVTWM submitted 2024-06-20 cs.CV cs.CL

classification cs.CVcs.CL
keywords classificationfine-grainedobjectlvlmsmodelstextbfclipfoci
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent Large Vision-Language Models (LVLMs) demonstrate impressive abilities on numerous image understanding and reasoning tasks. The task of fine-grained object classification (e.g., distinction between \textit{animal species}), however, has been probed insufficiently, despite its downstream importance. We fill this evaluation gap by creating \texttt{FOCI} (\textbf{F}ine-grained \textbf{O}bject \textbf{C}lass\textbf{I}fication), a difficult multiple-choice benchmark for fine-grained object classification, from existing object classification datasets: (1) multiple-choice avoids ambiguous answers associated with casting classification as open-ended QA task; (2) we retain classification difficulty by mining negative labels with a CLIP model. \texttt{FOCI}\xspace complements five popular classification datasets with four domain-specific subsets from ImageNet-21k. We benchmark 12 public LVLMs on \texttt{FOCI} and show that it tests for a \textit{complementary skill} to established image understanding and reasoning benchmarks. Crucially, CLIP models exhibit dramatically better performance than LVLMs. Since the image encoders of LVLMs come from these CLIP models, this points to inadequate alignment for fine-grained object distinction between the encoder and the LLM and warrants (pre)training data with more fine-grained annotation. We release our code at \url{https://github.com/gregor-ge/FOCI-Benchmark}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    HiR² extracts coarse-to-fine visual features from LMM layers and regularizes them with Lorentz entailment cones and unit-sphere dispersive loss, improving hierarchical consistency across models and fine-tuning methods.

  2. Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models

    cs.CV 2025-01 reject novelty 6.0 of 10

    Finedefics improves fine-grained image classification in multimodal LLMs by contrastively aligning image, attribute, and category embeddings, though its headline gains are measured against zero-shot baselines rather t...

  3. Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization

    cs.CV 2025-04 conditional novelty 5.0 of 10

    KeCO updates the visual feature keys of a small coreset with all leftover support images, and its diversity-based update outperforms retrieval from the five-times-larger full support set for LVLM in-context image clas...

  4. Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A prompt-ranking metric that subtracts semantic similarity to confounding classes selects higher-precision prompts for zero-shot vision-language object detection.

  5. MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A broad survey that organizes MLLM evaluation benchmarks into capability categories, explains benchmark construction and scoring methods, and identifies gaps in current evaluation practice.

Pith tools