Pith. sign in

REVIEW 2 cited by

Guiding Long-Short Term Memory for Image Caption Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1509.04942 v1 pith:LBHPIW4X submitted 2015-09-16 cs.CV

classification cs.CV
keywords imageshortcaptiongenerationguidinglstmmemorymodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this work we focus on the problem of image caption generation. We propose an extension of the long short term memory (LSTM) model, which we coin gLSTM for short. In particular, we add semantic information extracted from the image as extra input to each unit of the LSTM block, with the aim of guiding the model towards solutions that are more tightly coupled to the image content. Additionally, we explore different length normalization strategies for beam search in order to prevent from favoring short sentences. On various benchmark datasets such as Flickr8K, Flickr30K and MS COCO, we obtain results that are on par with or even outperform the current state-of-the-art.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Selecting a sparse uniform set of 25% of LLM layers for LoRA tuning preserves about 99% of visual task performance across four LVLMs and speeds up training by 12 to 23%.

  2. UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer

    cs.CV 2024-12 reject novelty 4.0 of 10

    The authors combine factual and stylized image captioning with a transformer summarizer to output a single caption containing factual, romantic, and humorous elements.

Pith tools