Pith. sign in

REVIEW 10 cited by

MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.05242 v1 pith:BALCNMVH submitted 2025-03-07 cs.CL

MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio

classification cs.CL
keywords mm-storyagentstoryimmersivestorytellingacrossattractivenessaudioevaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The rapid advancement of large language models (LLMs) and artificial intelligence-generated content (AIGC) has accelerated AI-native applications, such as AI-based storybooks that automate engaging story production for children. However, challenges remain in improving story attractiveness, enriching storytelling expressiveness, and developing open-source evaluation benchmarks and frameworks. Therefore, we propose and opensource MM-StoryAgent, which creates immersive narrated video storybooks with refined plots, role-consistent images, and multi-channel audio. MM-StoryAgent designs a multi-agent framework that employs LLMs and diverse expert tools (generative models and APIs) across several modalities to produce expressive storytelling videos. The framework enhances story attractiveness through a multi-stage writing pipeline. In addition, it improves the immersive storytelling experience by integrating sound effects with visual, music and narrative assets. MM-StoryAgent offers a flexible, open-source platform for further development, where generative modules can be substituted. Both objective and subjective evaluation regarding textual story quality and alignment between modalities validate the effectiveness of our proposed MM-StoryAgent system. The demo and source code are available.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

    cs.CV 2026-07 conditional novelty 7.0

    FilmWorld generates multi-scene films from novels by materializing an explicit evolving world-state trajectory and rendering shots in parallel, beating five agents on its own FilmEval benchmark.

  2. KathaTrace: Diagnosing Semantic Trajectory Collapse in Generated Visual Narratives

    cs.CV 2026-07 unverdicted novelty 7.0

    Introduces KathaTrace protocol and KathaBench-25K benchmark to quantify Semantic Trajectory Gap (STG) as the loss of transition meaning in visualized narratives, reporting STG of 23.5 +/- 1.3 across generators.

  3. GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

    cs.CV 2026-06 unverdicted novelty 7.0

    GroundShot introduces entity-grounded shot scheduling with online visual memory to improve consistency in multi-shot video generation and presents GroundBench for entity-level evaluation.

  4. DramaDirector: Geometry-Guided Short Drama Generation

    cs.CV 2026-06 conditional novelty 6.0

    Geometry-indexed depth–pose retrieval plus schema SFT and GRPO planning improves faithfulness, consistency, and controllability of plot-to-short-drama video generation over multi-agent and text-only baselines.

  5. GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

    cs.CV 2026-06 conditional novelty 6.0

    A training-free framework that reorders shot generation and maintains per-entity visual memory improves cross-shot character, object, and scene consistency over narrative-order memory baselines.

  6. Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

    cs.CV 2026-05 conditional novelty 6.0

    Crayotter, a multimodal multi-agent video editor that exposes every planning and tool step, scored 3.40/5 in human evaluation on 23 themes, beating two baselines (2.44 and 1.70).

  7. Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

    cs.CV 2026-05 unverdicted novelty 6.0

    Crayotter introduces a traceable three-phase multi-agent workflow for long-form video editing that scores 3.40/5 in human evaluations, outperforming two baselines on 23 themes.

  8. DramaDirector: Geometry-Guided Short Drama Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    DramaDirector retrieves depth-pose references from real drama shots to guide first-frame and image-to-video synthesis for plot-driven short dramas, paired with the DramaBoard benchmark.

  9. Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation

    cs.CV 2026-04 unverdicted novelty 5.0

    A self-reasoning agentic framework constructs a Product Narrative Framework, generates constraint-aware unified grid collages, and refines outputs via failure attribution to improve narrative coherence and aesthetics ...

  10. Character-Centered Dialogue Generation from Scene-Level Prompts

    cs.CV 2025-05 unverdicted novelty 4.0

    A training-free framework generates expressive, character-grounded dialogue and speech from scene prompts using vision-language encoders, LLMs, and a recursive narrative memory bank for cross-scene consistency.