Pith. sign in

REVIEW 19 cited by

DreamFactory: Pioneering Multi-Scene Long Video Generation with a Multi-Agent Framework

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11788 v1 pith:JPLSPCTJ submitted 2024-08-21 cs.AI cs.CLcs.CVcs.SE

DreamFactory: Pioneering Multi-Scene Long Video Generation with a Multi-Agent Framework

classification cs.AI cs.CLcs.CVcs.SE
keywords videosdreamfactorylongmulti-scenetextttchallengeconsistencycross-scene
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Current video generation models excel at creating short, realistic clips, but struggle with longer, multi-scene videos. We introduce \texttt{DreamFactory}, an LLM-based framework that tackles this challenge. \texttt{DreamFactory} leverages multi-agent collaboration principles and a Key Frames Iteration Design Method to ensure consistency and style across long videos. It utilizes Chain of Thought (COT) to address uncertainties inherent in large language models. \texttt{DreamFactory} generates long, stylistically coherent, and complex videos. Evaluating these long-form videos presents a challenge. We propose novel metrics such as Cross-Scene Face Distance Score and Cross-Scene Style Consistency Score. To further research in this area, we contribute the Multi-Scene Videos Dataset containing over 150 human-rated videos.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

    cs.CV 2026-06 unverdicted novelty 7.0

    GroundShot introduces entity-grounded shot scheduling with online visual memory to improve consistency in multi-shot video generation and presents GroundBench for entity-level evaluation.

  2. EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

    cs.CV 2026-05 conditional novelty 7.0

    EntityBench is a new benchmark with detailed per-shot entity schedules from real media, and the EntityMem baseline using persistent per-entity memory achieves the highest character fidelity with Cohen's d of +2.33.

  3. CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives

    cs.CV 2026-05 unverdicted novelty 7.0

    CausalCine enables real-time causal autoregressive multi-shot video generation via multi-shot training, content-aware memory routing for coherence, and distillation to few-step inference.

  4. TIE: Time Interval Encoding for Video Generation over Events

    cs.CV 2026-05 unverdicted novelty 7.0

    TIE derives a sinc-based interval encoding from Temporal Integrability and Duration Invariance principles, raising human-verified temporal constraint satisfaction from 77.34% to 96.03% while preserving visual quality ...

  5. TIE: Time Interval Encoding for Video Generation over Events

    cs.CV 2026-05 unverdicted novelty 7.0

    TIE derives a sinc-based interval encoding from temporal integrability and duration invariance principles, raising temporal constraint satisfaction from 77% to 96% on the OmniEvents dataset while preserving visual quality.

  6. AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

    cs.CV 2026-07 conditional novelty 6.0

    AgentHOI generates human-object interaction videos from text plus one human image and one object image, using multi-agent action planning and implicit text-to-motion feature alignment inside a video diffusion model.

  7. SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation

    cs.CV 2026-07 conditional novelty 6.0

    SlotMem keeps a compact, updateable memory slot for each recurring character and injects it only into that character's tokens, reporting improved long-range identity consistency in narrative video generation.

  8. GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

    cs.CV 2026-06 conditional novelty 6.0

    A training-free framework that reorders shot generation and maintains per-entity visual memory improves cross-shot character, object, and scene consistency over narrative-order memory baselines.

  9. Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

    cs.CV 2026-05 conditional novelty 6.0

    Crayotter, a multimodal multi-agent video editor that exposes every planning and tool step, scored 3.40/5 in human evaluation on 23 themes, beating two baselines (2.44 and 1.70).

  10. Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

    cs.CV 2026-05 unverdicted novelty 6.0

    Crayotter introduces a traceable three-phase multi-agent workflow for long-form video editing that scores 3.40/5 in human evaluations, outperforming two baselines on 23 themes.

  11. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    MAVEN introduces a multi-agent system for refining prompts in multicultural text-to-video generation and releases a benchmark of 243 prompts and 972 videos showing improved cultural relevance via parallel agent specia...

  12. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    MAVEN is a multi-agent prompt refinement framework that improves cultural fidelity in text-to-video generation, demonstrated on a new benchmark of 243 prompts and 972 videos across Chinese, American, and Romanian cultures.

  13. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 conditional novelty 6.0

    Parallel, role-specialized prompt agents improve cultural relevance in text-to-video generation, with a new cross-cultural benchmark showing the largest gains for location cues.

  14. StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics

    cs.CV 2026-04 unverdicted novelty 6.0

    StoryBlender generates inter-shot consistent editable 3D storyboards using a three-stage pipeline of semantic-spatial grounding, canonical asset materialization, and spatial-temporal dynamics with agent-based verification.

  15. Rolling Forcing: Autoregressive Long Video Diffusion in Real Time

    cs.CV 2025-09 unverdicted novelty 6.0

    Rolling Forcing generates multi-minute videos in real time by jointly denoising frames at increasing noise levels, anchoring attention to early frames, and using windowed distillation to limit error accumulation.

  16. AmbiGraph-Eval: Can LLMs Effectively Handle Ambiguous Graph Queries?

    cs.DB 2025-08 unverdicted novelty 6.0

    A benchmark organized by a six-type taxonomy of ambiguous graph queries reportedly shows that nine LLMs, including top models, frequently produce wrong query translations.

  17. Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops

    cs.CV 2026-06 unverdicted novelty 5.0

    CHIEF is a human-in-the-loop video generation system that combines creator direction with LLM-based audience-perspective critiques to improve narrative coherence in AI-generated videos, tested on student-made films up...

  18. CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    Introduces CineDance-1M dataset for multi-shot long-form text-to-audio-video generation along with CineBench and a model adaptation.

  19. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

    cs.CV 2025-03 unverdicted novelty 2.0

    The paper provides the first comprehensive survey of multimodal chain-of-thought reasoning, including foundational concepts, a taxonomy of methodologies, application analyses, challenges, and future directions.