Pith. sign in

REVIEW 18 cited by

SEED-Story: Multimodal Long Story Generation with Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08683 v2 pith:LVUKULFV submitted 2024-07-11 cs.CV

classification cs.CV
keywords generationmultimodalmodelimagesstorytasktextscomprehension
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the remarkable advancements in image generation and open-form text generation, the creation of interleaved image-text content has become an increasingly intriguing field. Multimodal story generation, characterized by producing narrative texts and vivid images in an interleaved manner, has emerged as a valuable and practical task with broad applications. However, this task poses significant challenges, as it necessitates the comprehension of the complex interplay between texts and images, and the ability to generate long sequences of coherent, contextually relevant texts and visuals. In this work, we propose SEED-Story, a novel method that leverages a Multimodal Large Language Model (MLLM) to generate extended multimodal stories. Our model, built upon the powerful comprehension capability of MLLM, predicts text tokens as well as visual tokens, which are subsequently processed with an adapted visual de-tokenizer to produce images with consistent characters and styles. We further propose multimodal attention sink mechanism to enable the generation of stories with up to 25 sequences (only 10 for training) in a highly efficient autoregressive manner. Additionally, we present a large-scale and high-resolution dataset named StoryStream for training our model and quantitatively evaluating the task of multimodal story generation in various aspects.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Story2Board: A Training-Free Approach for Expressive Storyboard Generation

    cs.CV 2025-08 conditional novelty 7.0 of 10

    Story2Board uses reciprocal attention value mixing and latent panel anchoring to generate consistent yet visually diverse storyboards from text without any training.

  2. IDEA-Bench: How Far are Generative Models from Professional Designing?

    cs.CV 2024-12 conditional novelty 7.0 of 10

    IDEA-Bench measures generative models on 100 professional design tasks and finds the best tested system scores only 22.48 out of 100.

  3. FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    FreeStory reformulates character consistency as entity-grounded feature reuse for free-form prompts, introduces FreeStoryBench, and reports stronger consistency than baselines among training-free methods.

  4. FairyGen: Storied Cartoon Video from a Single Child-Drawn Character

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A pipeline that generates story-driven cartoon videos from one child-drawn character by separating foreground style, background synthesis, and motion learning.

  5. AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AnimeShooter provides hierarchical story and shot annotations plus reference images for 148K one-minute animation stories, and AnimeShooterGen trained on it shows improved cross-shot consistency.

  6. Rethinking Causal Mask Attention for Vision-Language Inference

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Relaxing causal masking so image tokens can preview future image and text context during prefill improves several vision-language benchmarks, and pooling future attention into a single prefix token preserves most of the gain.

  7. UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

    cs.CV 2025-02 conditional novelty 6.0 of 10

    UniMoD prunes tokens with task-specific routers in unified multimodal transformers, cutting training FLOPs by 15-40% while roughly maintaining benchmark performance.

  8. One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Concatenating all frame prompts into a single prompt, then reweighting singular values and re-anchoring cross-attention, yields training-free identity-consistent text-to-image generation.

  9. VideoAuteur: Towards Long Narrative Video Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    VideoAuteur builds a cooking narrative dataset and an autoregressive pipeline that generates coherent long-form cooking videos by predicting actions, captions, and CLIP-based visual embeddings step by step.

  10. Olympus: A Universal Task Router for Computer Vision Tasks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Olympus is a trained MLLM router that delegates 20 vision tasks to specialist models and supports chain-of-action execution of up to five tasks per instruction.

  11. DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DiffSensei combines an SDXL diffusion generator with a multimodal LLM adapter and masked attention to generate manga pages with multiple characters whose poses and expressions follow panel captions.

  12. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Diffusion-conditioned denoising trains a video tokenizer whose features support both video question answering and text-to-video generation in a single LLM.

  13. OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A large new benchmark and an offline judge model for open-ended interleaved image-text generation, with IntJudge matching human agreement better than GPT-4o.

  14. StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    A training-free inference-time pipeline uses masked cross-image attention sharing and region harmonization to keep subjects consistent across generated story images.

  15. Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An integrated storytelling framework that generates text, scene graphs, images, and sound together reports higher expert-rated coherence than a sequential baseline, but the evaluation is small and qualitative.

  16. Captain Cinema: Towards Short Movie Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A two-stage text-to-movie system that plans keyframes for the story and then synthesizes video between them, using a compressed memory bank to keep long narratives consistent.

  17. SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation

    cs.CR 2024-12 conditional novelty 5.0 of 10

    SpearBot uses jailbreak prompts and iterative LLM-critic feedback to generate spear-phishing emails that evade machine detectors and appear human-like.

  18. Curiosity-Driven Reinforcement Learning from Human Feedback

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Adding a prediction-error curiosity reward, masked to non-top-k tokens, improves output diversity of RLHF-trained LLMs while keeping reward-model judged quality roughly unchanged.

Pith tools