Pith. sign in

REVIEW 3 cited by

Streaming Video Diffusion: Online Video Editing with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19726 v1 pith:FTG2N7D6 submitted 2024-05-30 cs.CV

classification cs.CV
keywords videoeditingstreamingdiffusiononlinetemporalvideosedit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a novel task called online video editing, which is designed to edit \textbf{streaming} frames while maintaining temporal consistency. Unlike existing offline video editing assuming all frames are pre-established and accessible, online video editing is tailored to real-life applications such as live streaming and online chat, requiring (1) fast continual step inference, (2) long-term temporal modeling, and (3) zero-shot video editing capability. To solve these issues, we propose Streaming Video Diffusion (SVDiff), which incorporates the compact spatial-aware temporal recurrence into off-the-shelf Stable Diffusion and is trained with the segment-level scheme on large-scale long videos. This simple yet effective setup allows us to obtain a single model that is capable of executing a broad range of videos and editing each streaming frame with temporal coherence. Our experiments indicate that our model can edit long, high-quality videos with remarkable results, achieving a real-time inference speed of 15.2 FPS at a resolution of 512x512.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Inverting the Streaming-Diffusion Bottleneck: Video-Rate MLLM-Conditioned Edit Diffusion on a Consumer GPU

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Reports a streaming pipeline with asymmetric CUDA pipelining and batched MLLM amortization that sustains 27.4 fps at 512x512 on RTX 3090 Ti for oil-painting stylization.

  2. In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion

    cs.CV 2026-08 conditional novelty 6.0 of 10

    In-Context Forcing conditions each denoising step of the current video frame on previous frames with decreasing noise levels, improving VBench dynamic scores and enabling parallel inference.

  3. Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!

    cs.CV 2025-10 conditional novelty 6.0 of 10

    DragStream enables real-time drag, deform, and rotate edits on autoregressively generated videos without retraining, by correcting latent drift and selectively filtering context features.

Pith tools