Pith. sign in

Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt. This study focuses on zero-shot text-to-video generation considering the data- and cost-efficient. To generate a semantic-coherent video, exhibiting a rich portrayal of temporal semantics such as the whole process of flower blooming rather than a set of "moving images", we propose a novel Free-Bloom pipeline that harnesses large language models (LLMs) as the director to generate a semantic-coherence prompt sequence, while pre-trained latent diffusion models (LDMs) as the animator to generate the high fidelity frames. Furthermore, to ensure temporal and identical coherence while maintaining semantic coherence, we propose a series of annotative modifications to adapting LDMs in the reverse process, including joint noise sampling, step-aware attention shift, and dual-path interpolation. Without any video data and training requirements, Free-Bloom generates vivid and high-quality videos, awe-inspiring in generating complex scenes with semantic meaningful frame sequences. In addition, Free-Bloom is naturally compatible with LDMs-based extensions.

fields

cs.CV 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Infinite Worlds with Versatile Interactions

cs.CV · 2026-07-08 · conditional · novelty 5.0

An open-source causal video world model sustains hour-long, 720p/60fps interactive generation without visual drift, paired with a VLM-based director-pilot agentic harness for rich, open-ended interaction.

citing papers explorer

Showing 1 of 1 citing paper.

  • Infinite Worlds with Versatile Interactions cs.CV · 2026-07-08 · conditional · none · ref 25 · internal anchor

    An open-source causal video world model sustains hour-long, 720p/60fps interactive generation without visual drift, paired with a VLM-based director-pilot agentic harness for rich, open-ended interaction.