Pith. sign in

REVIEW 2 cited by

Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11814 v3 pith:4UNA4NGN submitted 2024-07-16 cs.CV

classification cs.CV
keywords videoconsistencymulti-scenescenecontrastivedescriptionsgeneratednext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generated video scenes for action-centric sequence descriptions, such as recipe instructions and do-it-yourself projects, often include non-linear patterns, where the next video may need to be visually consistent not with the immediately preceding video but with earlier ones. Current multi-scene video synthesis approaches fail to meet these consistency requirements. To address this, we propose a contrastive sequential video diffusion method that selects the most suitable previously generated scene to guide and condition the denoising process of the next scene. The result is a multi-scene video that is grounded in the scene descriptions and coherent w.r.t. the scenes that require visual consistency. Experiments with action-centered data from the real world demonstrate the practicality and improved consistency of our model compared to previous work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Presto lets each temporal segment of a video attend to a matching sub-caption, enabling long, content-rich videos that beat prior open-source and commercial models on VBench semantic and dynamic scores.

  2. MatchDiffusion: Training-free Generation of Match-cuts

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Training-free match-cut synthesis: run two text-to-video generations from one shared noisy latent for the first K denoising steps, then let each follow its own prompt, so the outputs align in structure and motion.

Pith tools