Pith. sign in

Repope: Impact of annotation errors on the pope benchmark

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it
abstract

Since data annotation is costly, benchmark datasets often incorporate labels from established image datasets. In this work, we assess the impact of label errors in MSCOCO on the frequently used object hallucination benchmark POPE. We re-annotate the benchmark images and identify an imbalance in annotation errors across different subsets. Evaluating multiple models on the revised labels, which we denote as RePOPE, we observe notable shifts in model rankings, highlighting the impact of label quality. Code and data are available at https://github.com/YanNeu/RePOPE .

fields

cs.CV 2 cs.AI 1

years

2026 3

representative citing papers

Diagnosing Visual Ignorance in Vision-Language Models

cs.CV · 2026-06-05 · unverdicted · novelty 6.0

VLMs show language-prior reliance via multi-stage bottlenecks in visual retrieval and suppression, with many benchmark examples remaining answerable under severe visual obfuscation.

citing papers explorer

Showing 3 of 3 citing papers.