Selecting a sparse uniform set of 25% of LLM layers for LoRA tuning preserves about 99% of visual task performance across four LVLMs and speeds up training by 12 to 23%.
Guiding Long-Short Term Memory for Image Caption Generation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this work we focus on the problem of image caption generation. We propose an extension of the long short term memory (LSTM) model, which we coin gLSTM for short. In particular, we add semantic information extracted from the image as extra input to each unit of the LSTM block, with the aim of guiding the model towards solutions that are more tightly coupled to the image content. Additionally, we explore different length normalization strategies for beam search in order to prevent from favoring short sentences. On various benchmark datasets such as Flickr8K, Flickr30K and MS COCO, we obtain results that are on par with or even outperform the current state-of-the-art.
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
Selecting a sparse uniform set of 25% of LLM layers for LoRA tuning preserves about 99% of visual task performance across four LVLMs and speeds up training by 12 to 23%.