A single LLaVA-based captioning model continuously controls caption length, descriptiveness, and word uniqueness by interpolating between learned endpoint conditioning vectors.
Analysis of diversity-accuracy tradeoff in image captioning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We investigate the effect of different model architectures, training objectives, hyperparameter settings and decoding procedures on the diversity of automatically generated image captions. Our results show that 1) simple decoding by naive sampling, coupled with low temperature is a competitive and fast method to produce diverse and accurate caption sets; 2) training with CIDEr-based reward using Reinforcement learning harms the diversity properties of the resulting generator, which cannot be mitigated by manipulating decoding parameters. In addition, we propose a new metric AllSPICE for evaluating both accuracy and diversity of a set of captions by a single value.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
A single LLaVA-based captioning model continuously controls caption length, descriptiveness, and word uniqueness by interpolating between learned endpoint conditioning vectors.