Pith. sign in

REVIEW 11 cited by

SEED-Story: Multimodal Long Story Generation with Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08683 v2 pith:LVUKULFV submitted 2024-07-11 cs.CV

classification cs.CV
keywords generationmultimodalmodelimagesstorytasktextscomprehension
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the remarkable advancements in image generation and open-form text generation, the creation of interleaved image-text content has become an increasingly intriguing field. Multimodal story generation, characterized by producing narrative texts and vivid images in an interleaved manner, has emerged as a valuable and practical task with broad applications. However, this task poses significant challenges, as it necessitates the comprehension of the complex interplay between texts and images, and the ability to generate long sequences of coherent, contextually relevant texts and visuals. In this work, we propose SEED-Story, a novel method that leverages a Multimodal Large Language Model (MLLM) to generate extended multimodal stories. Our model, built upon the powerful comprehension capability of MLLM, predicts text tokens as well as visual tokens, which are subsequently processed with an adapted visual de-tokenizer to produce images with consistent characters and styles. We further propose multimodal attention sink mechanism to enable the generation of stories with up to 25 sequences (only 10 for training) in a highly efficient autoregressive manner. Additionally, we present a large-scale and high-resolution dataset named StoryStream for training our model and quantitatively evaluating the task of multimodal story generation in various aspects.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Story2Board: A Training-Free Approach for Expressive Storyboard Generation

    cs.CV 2025-08 conditional novelty 7.0 of 10

    Story2Board uses reciprocal attention value mixing and latent panel anchoring to generate consistent yet visually diverse storyboards from text without any training.

  2. FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

    cs.CV 2026-06 conditional novelty 6.0 of 10

    FreeStory keeps character identity consistent across free-form story images by grounding pronouns/type mentions to descriptions and reusing attention features with dynamic masks, correspondence matching, KV injection,...

  3. FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    FreeStory reformulates character consistency as entity-grounded feature reuse for free-form prompts, introduces FreeStoryBench, and reports stronger consistency than baselines among training-free methods.

  4. LongLive: Real-time Interactive Long Video Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    LongLive is a causal autoregressive video generator that produces up to 240-second interactive videos at 20.7 FPS on one H100 GPU after 32 GPU-days of fine-tuning from a 1.3B short-clip model.

  5. Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Sa2VA unifies SAM-2 segmentation with MLLM reasoning into a single model for referring segmentation and conversation on images and videos, supported by a new 72k-expression Ref-SAV dataset.

  6. When Attention Sink Emerges in Language Models: An Empirical View

    cs.CL 2024-10 accept novelty 6.0 of 10

    Attention sinks emerge in language models from softmax-induced token dependence on attention scores and do not appear when using sigmoid attention without normalization in models up to 1B parameters.

  7. Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization

    cs.CL 2025-11 unverdicted novelty 5.0 of 10

    MAGIC-HMO is a multi-agent framework that treats Chinese short-form creative NLG as heterogeneous multi-objective optimization over personalized constraints plus explanation reliability and outperforms baselines on a ...

  8. StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    A training-free inference-time pipeline uses masked cross-image attention sharing and region harmonization to keep subjects consistent across generated story images.

  9. Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An integrated storytelling framework that generates text, scene graphs, images, and sound together reports higher expert-rated coherence than a sequential baseline, but the evaluation is small and qualitative.

  10. Captain Cinema: Towards Short Movie Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A two-stage text-to-movie system that plans keyframes for the story and then synthesizes video between them, using a compressed memory bank to keep long narratives consistent.

  11. Character-Centered Dialogue Generation from Scene-Level Prompts

    cs.CV 2025-05 unverdicted novelty 4.0 of 10

    A training-free framework generates expressive, character-grounded dialogue and speech from scene prompts using vision-language encoders, LLMs, and a recursive narrative memory bank for cross-scene consistency.

Pith tools