Pith. sign in

REVIEW 10 cited by

Video-to-Video Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.06601 v2 pith:N62IT3ID submitted 2018-08-20 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords synthesisvideovideo-to-videoinputproblemadversarialapproachimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of the source video. While its image counterpart, the image-to-image synthesis problem, is a popular topic, the video-to-video synthesis problem is less explored in the literature. Without understanding temporal dynamics, directly applying existing image synthesis approaches to an input video often results in temporally incoherent videos of low visual quality. In this paper, we propose a novel video-to-video synthesis approach under the generative adversarial learning framework. Through carefully-designed generator and discriminator architectures, coupled with a spatio-temporal adversarial objective, we achieve high-resolution, photorealistic, temporally coherent video results on a diverse set of input formats including segmentation masks, sketches, and poses. Experiments on multiple benchmarks show the advantage of our method compared to strong baselines. In particular, our model is capable of synthesizing 2K resolution videos of street scenes up to 30 seconds long, which significantly advances the state-of-the-art of video synthesis. Finally, we apply our approach to future video prediction, outperforming several state-of-the-art competing systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer

    cs.CV 2025-09 conditional novelty 7.0 of 10

    Durian introduces a dual-reference diffusion model trained via self-reconstruction on video frames to enable cross-identity attribute transfer in portrait animations, supporting multi-attribute composition and interpolation.

  2. From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A one-step, draft-free video object removal model trained by distilling a ground-truth-conditioned teacher reaches comparable or better quality than multi-step diffusion methods while running in about 1 second.

  3. EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    EverAnimate restores drifted latent flow trajectories in chunked video generation via persistent latent propagation and restorative flow matching, achieving measurable gains in PSNR, SSIM, LPIPS, and FID over prior lo...

  4. IN2OUT: Fine-Tuning Video Inpainting Model for Video Outpainting Using Hierarchical Discriminator

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical discriminator with separate local and global objectives lets a video inpainting model be fine-tuned for video outpainting with reduced blur.

  5. PreciseCam: Precise Camera Control for Text-to-Image Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PreciseCam enables precise camera control (roll, pitch, vFoV, distortion) in text-to-image generation by conditioning SDXL with Perspective Field maps and a new dataset of 57,380 images.

  6. AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Timestep-segment preference optimization with separate motion and fidelity LoRAs improves audio-driven human animation quality and allows a 3.3x inference speedup.

  7. Multivariate-Information Adversarial Ensemble for Scalable Joint Distribution Matching

    cs.LG 2019-07 unverdicted novelty 5.0 of 10

    MMI-ALI extends pairwise ALI models into an m-domain ensemble by maximizing MMI on joint variables to achieve scalable joint distribution matching with linear scaling in m.

  8. UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation

    cs.CV 2025-02 reject novelty 4.0 of 10

    UniCP combines error-aware caching and PCA-based pruning to speed up diffusion-transformer video generation by about 1.6x.

  9. Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

    cs.CV 2025-01 reject novelty 4.0 of 10

    GE-Adapter combines a temporal smoothness loss, bilateral-filtered DDIM inversion, and shared plus frame-specific prompt tokens to improve text-to-video editing, though the reported evidence is inconsistent.

  10. Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques

    cs.LG 2025-07 conditional novelty 3.0 of 10

    A survey that categorizes tabular data synthesis by generation objectives and adds a benchmark comparison of six models on Adult and CreditRisk.

Pith tools