Pith. sign in

REVIEW 2 cited by

Aligning where to see and what to tell: image caption with region-based attention and scene factorization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1506.06272 v1 pith:VAVHLJCE submitted 2015-06-20 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords imagevisualattentioncontextsscenesystemcaptiongeneration
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent progress on automatic generation of image captions has shown that it is possible to describe the most salient information conveyed by images with accurate and meaningful sentences. In this paper, we propose an image caption system that exploits the parallel structures between images and sentences. In our model, the process of generating the next word, given the previously generated ones, is aligned with the visual perception experience where the attention shifting among the visual regions imposes a thread of visual ordering. This alignment characterizes the flow of "abstract meaning", encoding what is semantically shared by both the visual scene and the text description. Our system also makes another novel modeling contribution by introducing scene-specific contexts that capture higher-level semantic information encoded in an image. The contexts adapt language models for word generation to specific scene types. We benchmark our system and contrast to published results on several popular datasets. We show that using either region-based attention or scene-specific contexts improves systems without those components. Furthermore, combining these two modeling ingredients attains the state-of-the-art performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aesthetic Image Captioning From Weakly-Labelled Photographs

    cs.CV 2019-08 conditional novelty 6.0 of 10

    By filtering noisy web comments, the authors built AVA-Captions, a 230,000-image aesthetic captioning dataset, and showed a weakly supervised CNN can match ImageNet-pretrained features for this task.

  2. Image Captioning using Facial Expression and Attention

    cs.CV 2019-08 conditional novelty 6.0 of 10

    Facial expression features, especially with attention, yield small captioning improvements on face-containing Flickr images, driven mostly by more diverse verbs.

Pith tools