Pith. sign in

REVIEW 2 cited by

Taxonomy-Aware Evaluation of Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.05457 v1 pith:PDJU25XP submitted 2025-04-07 cs.CV

classification cs.CV
keywords evaluationtextconifergeneratedmeasurespredictionssimilaritytaxonomy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When a vision-language model (VLM) is prompted to identify an entity depicted in an image, it may answer 'I see a conifer,' rather than the specific label 'norway spruce'. This raises two issues for evaluation: First, the unconstrained generated text needs to be mapped to the evaluation label space (i.e., 'conifer'). Second, a useful classification measure should give partial credit to less-specific, but not incorrect, answers ('norway spruce' being a type of 'conifer'). To meet these requirements, we propose a framework for evaluating unconstrained text predictions, such as those generated from a vision-language model, against a taxonomy. Specifically, we propose the use of hierarchical precision and recall measures to assess the level of correctness and specificity of predictions with regard to a taxonomy. Experimentally, we first show that existing text similarity measures do not capture taxonomic similarity well. We then develop and compare different methods to map textual VLM predictions onto a taxonomy. This allows us to compute hierarchical similarity measures between the generated text and the ground truth labels. Finally, we analyze modern VLMs on fine-grained visual classification tasks based on our proposed taxonomic evaluation scheme.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    HiR² extracts coarse-to-fine visual features from LMM layers and regularizes them with Lorentz entailment cones and unit-sphere dispersive loss, improving hierarchical consistency across models and fine-tuning methods.

  2. GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A zero-shot geospatial classifier that turns satellite images into text descriptions and uses a language model to assign labels, plus an optional LLM-based hierarchical clustering for many-class datasets.

Pith tools