Pith. sign in

REVIEW 2 cited by

VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.18837 v1 pith:CBO22KLK submitted 2023-11-30 cs.CV cs.AIcs.LGcs.MM

classification cs.CVcs.AIcs.LGcs.MM
keywords videotaskseditingvideosdiffusioninstructionsvidiffgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have achieved significant success in image and video generation. This motivates a growing interest in video editing tasks, where videos are edited according to provided text descriptions. However, most existing approaches only focus on video editing for short clips and rely on time-consuming tuning or inference. We are the first to propose Video Instruction Diffusion (VIDiff), a unified foundation model designed for a wide range of video tasks. These tasks encompass both understanding tasks (such as language-guided video object segmentation) and generative tasks (video editing and enhancement). Our model can edit and translate the desired results within seconds based on user instructions. Moreover, we design an iterative auto-regressive method to ensure consistency in editing and enhancing long videos. We provide convincing generative results for diverse input videos and written instructions, both qualitatively and quantitatively. More examples can be found at our website https://ChenHsing.github.io/VIDiff.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A siamese-branch layout adapter lets multimodal diffusion transformers follow detailed region captions and bounding boxes, beating prior layout-to-image methods on a new 2.7M-pair dataset and benchmark.

  2. StableAnimator: High-Quality Identity-Preserving Human Image Animation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    StableAnimator uses a video diffusion model with face-embedding adapters and per-step latent optimization to generate pose-driven videos that preserve the reference person's identity end-to-end.

Pith tools