Pith. sign in

REVIEW 7 cited by

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.19267 v3 pith:J7YAFU5V submitted 2025-04-27 cs.CL cs.AIcs.CVcs.LG

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

classification cs.CL cs.AIcs.CVcs.LG
keywords visualstorytellingmetricsevaluationmodelsmultimodalnarrativesnovel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in multimodal models, specifically adapting transformer-based architectures and large multimodal models, for the visual storytelling task. Leveraging the large-scale Visual Storytelling (VIST) dataset, our VIST-GPT model produces visually grounded, contextually appropriate narratives. We address the limitations of traditional evaluation metrics, such as BLEU, METEOR, ROUGE, and CIDEr, which are not suitable for this task. Instead, we utilize RoViST and GROOVIST, novel reference-free metrics designed to assess visual storytelling, focusing on visual grounding, coherence, and non-redundancy. These metrics provide a more nuanced evaluation of narrative quality, aligning closely with human judgment.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs

    cs.LG 2026-05 unverdicted novelty 7.0

    Fine-tuned 7B LLMs generating unified diffs for neural architecture refinement achieve 66-75% valid rates and 64-66% mean first-epoch accuracy, outperforming full-generation baselines by large margins while cutting ou...

  2. Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models

    cs.CV 2026-01 unverdicted novelty 6.0

    Closed-loop LLM search with AST-generated examples discovers non-standard channel widths that improve vision model performance over initial architectures on CIFAR-100.

  3. Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design

    cs.CV 2025-12 conditional novelty 6.0

    Three-example few-shot prompting optimizes LLM-generated vision architectures while a whitespace-normalized hash provides 100x faster duplicate detection than AST parsing across seven benchmarks.

  4. LEMUR 2: Unlocking Neural Network Diversity for AI

    cs.LG 2026-07 conditional novelty 5.5

    LEMUR 2 releases a multi-generator, multi-task neural-architecture corpus with real-device latency metadata intended as fuel for LLM-driven AutoML.

  5. Controllable Narrative Rendering for Enhanced Assisted Writing

    cs.CL 2026-05 unverdicted novelty 4.0

    Loom is a framework using intent-centered semiotic chain-of-thought in a three-layer pipeline to separate perceptual material generation from syntactic insertion, achieving higher factual integrity and descriptive int...

  6. Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis

    cs.LG 2025-11 unverdicted novelty 4.0

    FractalNet automatically generates and tests over 1,200 CNN architectures based on recursive fractal templates, achieving up to 80.18% accuracy on CIFAR-10 after five training epochs.

  7. Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis

    cs.LG 2025-11 unverdicted novelty 3.0

    Fractal templates enable systematic creation of more than 1,200 neural network variants that show strong performance and computational efficiency when trained on CIFAR-10 for five epochs.