Pith. sign in

REVIEW 3 cited by

Zero-shot Generation of Coherent Storybook from Plain Text Story using Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.03900 v1 pith:JGRT2YXX submitted 2023-02-08 cs.CV cs.AIcs.LGstat.ML

classification cs.CVcs.AIcs.LGstat.ML
keywords coherentimagesgenerationmodelmodelsstorystorybookcoherency
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advancements in large scale text-to-image models have opened new possibilities for guiding the creation of images through human-devised natural language. However, while prior literature has primarily focused on the generation of individual images, it is essential to consider the capability of these models to ensure coherency within a sequence of images to fulfill the demands of real-world applications such as storytelling. To address this, here we present a novel neural pipeline for generating a coherent storybook from the plain text of a story. Specifically, we leverage a combination of a pre-trained Large Language Model and a text-guided Latent Diffusion Model to generate coherent images. While previous story synthesis frameworks typically require a large-scale text-to-image model trained on expensive image-caption pairs to maintain the coherency, we employ simple textual inversion techniques along with detector-based semantic image editing which allows zero-shot generation of the coherent storybook. Experimental results show that our proposed method outperforms state-of-the-art image editing baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    Query features in video diffusion models encode both motion and identity, enabling efficient zero-shot motion transfer and training-free multi-shot character consistency.

  2. Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A dual-stream diffusion model that jointly denoises the output image and an insertion mask, trained on a new 3.16 million pair dataset, outperforms prior baselines on affordance-aware object insertion.

  3. StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    A training-free inference-time pipeline uses masked cross-image attention sharing and region harmonization to keep subjects consistent across generated story images.

Pith tools