Pith. sign in

REVIEW 1 cited by

Room for improvement in automatic image description: an error analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1704.04198 v1 pith:6QJBKSEL submitted 2017-04-13 cs.CL

classification cs.CL
keywords descriptionsanalysiserrorerrorsimageautomaticbeendescription
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years we have seen rapid and significant progress in automatic image description but what are the open problems in this area? Most work has been evaluated using text-based similarity metrics, which only indicate that there have been improvements, without explaining what has improved. In this paper, we present a detailed error analysis of the descriptions generated by a state-of-the-art attention-based model. Our analysis operates on two levels: first we check the descriptions for accuracy, and then we categorize the types of errors we observe in the inaccurate descriptions. We find only 20% of the descriptions are free from errors, and surprisingly that 26% are unrelated to the image. Finally, we manually correct the most frequently occurring error types (e.g. gender identification) to estimate the performance reward for addressing these errors, observing gains of 0.2--1 BLEU point per type.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Large Language Models Judge Error Severity Like Humans?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Most tested LLMs excuse colour errors that people consider highly severe, while both groups penalize gender errors; only one model, Doubao, reproduced the human severity ordering.

Pith tools