Pith. sign in

REVIEW 7 cited by

FitVid: Overfitting in Pixel-Level Video Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.13195 v1 pith:ZREX7JIJ submitted 2021-06-24 cs.CV cs.LG

classification cs.CVcs.LG
keywords videomodelsbenchmarkscurrentfitvidoverfittingpredictionquality
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

An agent that is capable of predicting what happens next can perform a variety of tasks through planning with no additional training. Furthermore, such an agent can internally represent the complex dynamics of the real-world and therefore can acquire a representation useful for a variety of visual perception tasks. This makes predicting the future frames of a video, conditioned on the observed past and potentially future actions, an interesting task which remains exceptionally challenging despite many recent advances. Existing video prediction models have shown promising results on simple narrow benchmarks but they generate low quality predictions on real-life datasets with more complicated dynamics or broader domain. There is a growing body of evidence that underfitting on the training data is one of the primary causes for the low quality predictions. In this paper, we argue that the inefficient use of parameters in the current video models is the main reason for underfitting. Therefore, we introduce a new architecture, named FitVid, which is capable of severe overfitting on the common benchmarks while having similar parameter count as the current state-of-the-art models. We analyze the consequences of overfitting, illustrating how it can produce unexpected outcomes such as generating high quality output by repeating the training data, and how it can be mitigated using existing image augmentation techniques. As a result, FitVid outperforms the current state-of-the-art models across four different video prediction benchmarks on four different metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Overcoming Statistical Bias in Action-Controllable World Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Counterfactual consistency training makes action-conditioned world models' predictions respond to actions, reducing zero-action drift and improving average visual planning success from 70.1% to 73.1%.

  2. Cell as Point: One-Stage Framework for Efficient Cell Tracking

    eess.IV 2024-11 conditional novelty 6.0 of 10

    A point-based transformer framework tracks cells directly from raw microscopy frames, without segmentation, using event-guided sampling and rolling windows.

  3. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  4. Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A diffusion model generates dense 4D point trajectories from a single image, and a separate view-synthesis module renders them into novel-view videos.

  5. FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Explicitly predicting 3D scene flow before diffusion-based image generation improves future-frame prediction and visual planning in RGB-D robot manipulation world models.

  6. Efficient Continuous Video Flow Model for Video Prediction

    cs.CV 2024-12 conditional novelty 4.0 of 10

    The paper adapts the authors' prior continuous-video-process framework to latent space, reporting state-of-the-art FVD on KTH, BAIR, Human3.6M, and UCF101 with fewer parameters and sampling steps.

  7. Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction

    cs.CV 2024-12 conditional novelty 4.0 of 10

    CVP trains a network to reverse a continuous interpolation between past and future frames, reporting competitive FVD scores and 25-step sampling on KTH, BAIR, Human3.6M, and UCF101.

Pith tools