Pith. sign in

REVIEW 1 cited by

DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07066 v1 pith:ZUNYL24N submitted 2023-12-12 cs.CL cs.CV

classification cs.CLcs.CV
keywords diffuvstinferencemodelsscenesvisualautoregressivedenoisingfictional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in image and video creation, especially AI-based image synthesis, have led to the production of numerous visual scenes that exhibit a high level of abstractness and diversity. Consequently, Visual Storytelling (VST), a task that involves generating meaningful and coherent narratives from a collection of images, has become even more challenging and is increasingly desired beyond real-world imagery. While existing VST techniques, which typically use autoregressive decoders, have made significant progress, they suffer from low inference speed and are not well-suited for synthetic scenes. To this end, we propose a novel diffusion-based system DiffuVST, which models the generation of a series of visual descriptions as a single conditional denoising process. The stochastic and non-autoregressive nature of DiffuVST at inference time allows it to generate highly diverse narratives more efficiently. In addition, DiffuVST features a unique design with bi-directional text history guidance and multimodal adapter modules, which effectively improve inter-sentence coherence and image-to-text fidelity. Extensive experiments on the story generation task covering four fictional visual-story datasets demonstrate the superiority of DiffuVST over traditional autoregressive models in terms of both text quality and inference speed.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structured Relevance Assessment for Robust Retrieval-Augmented Language Models

    cs.AI 2025-07 reject novelty 4.0 of 10

    A retrieval-augmented framework that scores document relevance, balances internal and external knowledge, and abstains when uncertain, claims to cut hallucinations, but the reported evidence is thin and inconsistent.

Pith tools