Pith. sign in

REVIEW 1 cited by

Analysis of diversity-accuracy tradeoff in image captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.11848 v1 pith:FBSNP2RA submitted 2020-02-27 cs.CL cs.CV

classification cs.CLcs.CV
keywords decodingdiversitycaptionsimagetrainingaccuracyaccurateaddition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate the effect of different model architectures, training objectives, hyperparameter settings and decoding procedures on the diversity of automatically generated image captions. Our results show that 1) simple decoding by naive sampling, coupled with low temperature is a competitive and fast method to produce diverse and accurate caption sets; 2) training with CIDEr-based reward using Reinforcement learning harms the diversity properties of the resulting generator, which cannot be mitigated by manipulating decoding parameters. In addition, we propose a new metric AllSPICE for evaluating both accuracy and diversity of a set of captions by a single value.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single LLaVA-based captioning model continuously controls caption length, descriptiveness, and word uniqueness by interpolating between learned endpoint conditioning vectors.

Pith tools