Pith. sign in

REVIEW 2 cited by

Understanding Guided Image Captioning Performance across Domains

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.02339 v3 pith:WVUUAVRJ submitted 2020-12-04 cs.CV cs.CL

classification cs.CVcs.CL
keywords imageguidedcaptioningmodelsguidingcaptionconceptsgenerally
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image captioning models generally lack the capability to take into account user interest, and usually default to global descriptions that try to balance readability, informativeness, and information overload. On the other hand, VQA models generally lack the ability to provide long descriptive answers, while expecting the textual question to be quite precise. We present a method to control the concepts that an image caption should focus on, using an additional input called the guiding text that refers to either groundable or ungroundable concepts in the image. Our model consists of a Transformer-based multimodal encoder that uses the guiding text together with global and object-level image features to derive early-fusion representations used to generate the guided caption. While models trained on Visual Genome data have an in-domain advantage of fitting well when guided with automatic object labels, we find that guided captioning models trained on Conceptual Captions generalize better on out-of-domain images and guiding texts. Our human-evaluation results indicate that attempting in-the-wild guided image captioning requires access to large, unrestricted-domain training datasets, and that increased style diversity (even without increasing the number of unique tokens) is a key factor for improved performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SEMANTIC SEE-THROUGH GOGGLES: Wearing Linguistic Virtual Reality in (Artificial) Intelligence

    cs.HC 2024-12 conditional novelty 6.0 of 10

    A wearable AI system that turns the live view into one sentence and back into an image lets users experientially confront how linguistic mediation filters and biases perception.

  2. All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ALM-bench is a 100-language, 19-domain cultural visual QA benchmark on which GPT-4o reaches 78.8% and the best open model, GLM-4V-9B, reaches 51.9%.

Pith tools