Pith. sign in

REVIEW 2 cited by

Transparent Human Evaluation for Image Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.08940 v2 pith:ZMBA4EAP submitted 2021-11-17 cs.CL cs.CV

classification cs.CLcs.CV
keywords evaluationimagecaptioninghumanmetricsrecallautomaticcaptions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We establish THumB, a rubric-based human evaluation protocol for image captioning models. Our scoring rubrics and their definitions are carefully developed based on machine- and human-generated captions on the MSCOCO dataset. Each caption is evaluated along two main dimensions in a tradeoff (precision and recall) as well as other aspects that measure the text quality (fluency, conciseness, and inclusive language). Our evaluations demonstrate several critical problems of the current evaluation practice. Human-generated captions show substantially higher quality than machine-generated ones, especially in coverage of salient information (i.e., recall), while most automatic metrics say the opposite. Our rubric-based results reveal that CLIPScore, a recent metric that uses image features, better correlates with human judgments than conventional text-only metrics because it is more sensitive to recall. We hope that this work will promote a more transparent evaluation protocol for image captioning and its automatic metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence

    cs.HC 2024-12 conditional novelty 6.0 of 10

    A wearable AI system that turns the live view into one sentence and back into an image lets users experientially confront how linguistic mediation filters and biases perception.

  2. Perception of Visual Content: Differences Between Humans and Foundation Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Machine-generated captions are more similar to human labels than object-detection labels, and ML captions give the best region-classification performance while combined ML objects and captions give the best income prediction.

Pith tools