Pith. sign in

Fifo-diffusion: Generating infinite videos from text without training

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

fields

cs.CV 3

representative citing papers

TeleMorpher: Toward Robust Simultaneous Motion-Location Editing

cs.CV · 2026-06-18 · unverdicted · novelty 6.0

TeleMorpher introduces a training-free pose-warping pipeline plus two LPIPS-based metrics for simultaneous motion and location editing in videos, claiming superior results on in-the-wild and TaiChi data.

VRAG: Learning World Models for Interactive Video Generation

cs.CV · 2025-05-28 · conditional · novelty 5.0

VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft and RealEstate10K benchmarks.

Movie Gen: A Cast of Media Foundation Models

cs.CV · 2024-10-17 · unverdicted · novelty 5.0

A 30B-parameter transformer and related models generate high-quality videos and audio, claiming state-of-the-art results on text-to-video, video editing, personalization, and audio generation tasks.

citing papers explorer

Showing 3 of 3 citing papers.

  • TeleMorpher: Toward Robust Simultaneous Motion-Location Editing cs.CV · 2026-06-18 · unverdicted · none · ref 4

    TeleMorpher introduces a training-free pose-warping pipeline plus two LPIPS-based metrics for simultaneous motion and location editing in videos, claiming superior results on in-the-wild and TaiChi data.

  • VRAG: Learning World Models for Interactive Video Generation cs.CV · 2025-05-28 · conditional · none · ref 36

    VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft and RealEstate10K benchmarks.

  • Movie Gen: A Cast of Media Foundation Models cs.CV · 2024-10-17 · unverdicted · none · ref 33

    A 30B-parameter transformer and related models generate high-quality videos and audio, claiming state-of-the-art results on text-to-video, video editing, personalization, and audio generation tasks.