Pith. sign in

Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Visual storytelling systems generate multi-sentence stories from image sequences. In this task, capturing contextual information and bridging visual variation bring additional challenges. We propose a simple yet effective framework that leverages the generalization capabilities of pretrained foundation models, only training a lightweight vision-language mapping network to connect modalities, while incorporating context to enhance coherence. We introduce a multimodal contrastive objective that also improves visual relevance and story informativeness. Extensive experimental results, across both automatic metrics and human evaluations, demonstrate that the stories generated by our framework are diverse, coherent, informative, and interesting.

fields

cs.CL 1

years

2025 1

verdicts

UNVERDICTED 1

representative citing papers

From Image Captioning to Visual Storytelling

cs.CL · 2025-07-31 · unverdicted · novelty 4.0

Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.

citing papers explorer

Showing 1 of 1 citing paper.

  • From Image Captioning to Visual Storytelling cs.CL · 2025-07-31 · unverdicted · none · ref 67 · internal anchor

    Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.