Pith. sign in

REVIEW 1 cited by

SinFusion: Training Diffusion Models on a Single Image or Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11743 v3 pith:JVA3EKHA submitted 2022-11-21 cs.CV cs.LG

classification cs.CVcs.LG
keywords videoimagediffusionsingleinputmodelmodelsvideo-specific
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models exhibited tremendous progress in image and video generation, exceeding GANs in quality and diversity. However, they are usually trained on very large datasets and are not naturally adapted to manipulate a given input image or video. In this paper we show how this can be resolved by training a diffusion model on a single input image or video. Our image/video-specific diffusion model (SinFusion) learns the appearance and dynamics of the single image or video, while utilizing the conditioning capabilities of diffusion models. It can solve a wide array of image/video-specific manipulation tasks. In particular, our model can learn from few frames the motion and dynamics of a single input video. It can then generate diverse new video samples of the same dynamic scene, extrapolate short videos into long ones (both forward and backward in time) and perform video upsampling. Most of these tasks are not realizable by current video-specific generation methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Programmatic Video Prediction Using Large Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProgGen uses language-model-written programs for perception, dynamics, and rendering to predict future video frames from about ten training examples, beating large diffusion baselines on two synthetic benchmarks.

Pith tools