Pith. sign in

SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Visual storytelling aims to automatically generate a coherent story based on a given image sequence. Unlike tasks like image captioning, visual stories should contain factual descriptions, worldviews, and human social commonsense to put disjointed elements together to form a coherent and engaging human-writeable story. However, most models mainly focus on applying factual information and using taxonomic/lexical external knowledge when attempting to create stories. This paper introduces SCO-VIST, a framework representing the image sequence as a graph with objects and relations that includes human action motivation and its social interaction commonsense knowledge. SCO-VIST then takes this graph representing plot points and creates bridges between plot points with semantic and occurrence-based edge weights. This weighted story graph produces the storyline in a sequence of events using Floyd-Warshall's algorithm. Our proposed framework produces stories superior across multiple metrics in terms of visual grounding, coherence, diversity, and humanness, per both automatic and human evaluations.

fields

cs.CL 1

years

2025 1

verdicts

UNVERDICTED 1

representative citing papers

From Image Captioning to Visual Storytelling

cs.CL · 2025-07-31 · unverdicted · novelty 4.0

Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.

citing papers explorer

Showing 1 of 1 citing paper.

  • From Image Captioning to Visual Storytelling cs.CL · 2025-07-31 · unverdicted · none · ref 76 · internal anchor

    Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.