Pith. sign in

Epipolar geometry improves video generation models

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it
abstract

Video generation models have advanced significantly through the latent diffusion transformers trained with rectified flow techniques. Yet these models still struggle with geometric inconsistencies, unstable motion, and visual artifacts that break the illusion of realistic 3D scenes. 3D-consistent video generation could significantly impact numerous downstream applications in generation and reconstruction tasks. We explore how epipolar geometry constraints improve modern video diffusion models. Despite using massive training data, these models fail to capture fundamental geometric principles. We align diffusion models using pairwise epipolar geometry constraints via preference-based optimization, directly addressing unstable trajectories and geometric artifacts through mathematically principled geometric enforcement. Our approach efficiently enforces geometric principles without requiring end-to-end differentiability. Evaluation demonstrates that classical geometric constraints provide more stable optimization signals than modern learned metrics. Training on static scenes with dynamic cameras ensures metric quality while the model generalizes to various dynamic scenes. By bridging data-driven learning with classical computer vision, we reduce epipolar error by 31% and improve human-rated consistency from 54% to 72% without compromising visual quality.

fields

cs.CV 8

years

2026 8

representative citing papers

Probing into Camera Control of Video Models

cs.CV · 2026-05-14 · unverdicted · novelty 7.0

A training-free method reformulates camera control as geometric displacement fields applied via differentiable latent resampling, enabling control and bias probing in video diffusion models.

Geometry-Aware Implicit Memory for Video World Models

cs.CV · 2026-06-01 · unverdicted · novelty 6.0

GIM-World adds a camera-queryable geometry distillation head and pruning rule to implicit memory in video world models, claiming better long-horizon geometric consistency on the MIND benchmark than explicit and implicit baselines.

Feed-Forward Gaussian Splatting from Sparse Aerial Views

cs.CV · 2026-05-19 · unverdicted · novelty 5.0

AnyCity reconstructs coherent 3D Gaussian urban scenes from sparse aerial views in one feed-forward pass by anchoring observation-supported geometry and applying gated residual updates conditioned on an aerial-adapted video diffusion prior.

citing papers explorer

Showing 8 of 8 citing papers.

  • FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation cs.CV · 2026-06-23 · unverdicted · none · ref 36 · internal anchor

    FLAT maps compressed video diffusion latents to explicit triangle splats via ray-centered rotation parameterization and a product window function, reporting better geometric accuracy than 3D Gaussian baselines under identical training.

  • Geo-Align: Video Generation Alignment via Metric Geometry Reward cs.CV · 2026-05-22 · unverdicted · none · ref 59 · internal anchor

    Geo-Align applies RL with a perceptual reward derived from 3D camera trajectory estimation to improve controllability and fidelity in video generation without paired training data.

  • Probing into Camera Control of Video Models cs.CV · 2026-05-14 · unverdicted · none · ref 23 · internal anchor

    A training-free method reformulates camera control as geometric displacement fields applied via differentiable latent resampling, enabling control and bias probing in video diffusion models.

  • Geometry-Aware Implicit Memory for Video World Models cs.CV · 2026-06-01 · unverdicted · none · ref 31 · internal anchor

    GIM-World adds a camera-queryable geometry distillation head and pruning rule to implicit memory in video world models, claiming better long-horizon geometric consistency on the MIND benchmark than explicit and implicit baselines.

  • GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation cs.CV · 2026-05-18 · unverdicted · none · ref 35 · internal anchor

    GeoFlow adds a geometry-consistency reward based on rigid camera flow and object appearance preservation, integrated via reinforcement fine-tuning to improve geometric coherence in video generation.

  • Drift-Resistant Navigation World Model with Anchored Epipolar Guidance cs.CV · 2026-05-23 · unverdicted · none · ref 17 · internal anchor

    A generative navigation world model that uses sparse anchored rollout with epipolar constraints to reduce perceptual and geometric drift.

  • Feed-Forward Gaussian Splatting from Sparse Aerial Views cs.CV · 2026-05-19 · unverdicted · none · ref 10 · internal anchor

    AnyCity reconstructs coherent 3D Gaussian urban scenes from sparse aerial views in one feed-forward pass by anchoring observation-supported geometry and applying gated residual updates conditioned on an aerial-adapted video diffusion prior.

  • VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation cs.CV · 2026-01-30 · conditional · none · ref 7 · internal anchor

    Using VGGT's reconstruction error as a self-supervised reward, DPO post-training with ~2,500 preference pairs improves 3D consistency of CogVideoX-based video generation while preserving perceptual quality.