Pith. sign in

Slowfast-vgen: Slow-fast learning for action-driven long video generation

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

citation-role summary

background 2

citation-polarity summary

fields

cs.CV 4 cs.RO 1

years

2026 3 2025 2

roles

background 2

polarities

background 2

representative citing papers

Lyra 2.0: Explorable Generative 3D Worlds

cs.CV · 2026-04-14 · unverdicted · novelty 6.0

Lyra 2.0 produces persistent 3D-consistent video sequences for large explorable worlds by using per-frame geometry for information routing and self-augmented training to correct temporal drift.

VRAG: Learning World Models for Interactive Video Generation

cs.CV · 2025-05-28 · conditional · novelty 5.0

VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft and RealEstate10K benchmarks.

Towards Error-Free Long Video Generation

cs.CV · 2026-06-21 · unverdicted · novelty 4.0

An autoregressive diffusion framework with causal inter-clip attention, KV caching, and truncation-rectified flow produces coherent minute-level videos while reducing error accumulation.

citing papers explorer

Showing 5 of 5 citing papers.

  • CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives cs.CV · 2026-05-12 · unverdicted · none · ref 15

    CausalCine enables real-time causal autoregressive multi-shot video generation via multi-shot training, content-aware memory routing for coherence, and distillation to few-step inference.

  • Large Video Planner Enables Generalizable Robot Control cs.RO · 2025-12-17 · conditional · none · ref 39

    A video foundation model trained on human demonstrations generates zero-shot plans that convert to executable robot actions on novel scenes and tasks.

  • Lyra 2.0: Explorable Generative 3D Worlds cs.CV · 2026-04-14 · unverdicted · none · ref 32

    Lyra 2.0 produces persistent 3D-consistent video sequences for large explorable worlds by using per-frame geometry for information routing and self-augmented training to correct temporal drift.

  • VRAG: Learning World Models for Interactive Video Generation cs.CV · 2025-05-28 · conditional · none · ref 34

    VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft and RealEstate10K benchmarks.

  • Towards Error-Free Long Video Generation cs.CV · 2026-06-21 · unverdicted · none · ref 18

    An autoregressive diffusion framework with causal inter-clip attention, KV caching, and truncation-rectified flow produces coherent minute-level videos while reducing error accumulation.