Pith. sign in

Improved Conditional VRNNs for Video Prediction

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Predicting future frames for a video sequence is a challenging generative modeling task. Promising approaches include probabilistic latent variable models such as the Variational Auto-Encoder. While VAEs can handle uncertainty and model multiple possible future outcomes, they have a tendency to produce blurry predictions. In this work we argue that this is a sign of underfitting. To address this issue, we propose to increase the expressiveness of the latent distributions and to use higher capacity likelihood models. Our approach relies on a hierarchy of latent variables, which defines a family of flexible prior and posterior distributions in order to better model the probability of future sequences. We validate our proposal through a series of ablation experiments and compare our approach to current state-of-the-art latent variable models. Our method performs favorably under several metrics in three different datasets.

fields

cs.CV 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Efficient Continuous Video Flow Model for Video Prediction

cs.CV · 2024-12-07 · conditional · novelty 4.0

The paper adapts the authors' prior continuous-video-process framework to latent space, reporting state-of-the-art FVD on KTH, BAIR, Human3.6M, and UCF101 with fewer parameters and sampling steps.

citing papers explorer

Showing 1 of 1 citing paper.

  • Efficient Continuous Video Flow Model for Video Prediction cs.CV · 2024-12-07 · conditional · none · ref 4 · internal anchor

    The paper adapts the authors' prior continuous-video-process framework to latent space, reporting state-of-the-art FVD on KTH, BAIR, Human3.6M, and UCF101 with fewer parameters and sampling steps.