Pith. sign in

REVIEW 26 cited by

Video Depth Anything: Consistent Depth Estimation for Super-Long Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.12375 v3 pith:C47QNCKD submitted 2025-01-21 cs.CV cs.AI

classification cs.CVcs.AI
keywords depthvideoanythingvideosestimationmodeltemporalability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Depth Anything has achieved remarkable success in monocular depth estimation with strong generalization ability. However, it suffers from temporal inconsistency in videos, hindering its practical applications. Various methods have been proposed to alleviate this issue by leveraging video generation models or introducing priors from optical flow and camera poses. Nonetheless, these methods are only applicable to short videos (< 10 seconds) and require a trade-off between quality and computational efficiency. We propose Video Depth Anything for high-quality, consistent depth estimation in super-long videos (over several minutes) without sacrificing efficiency. We base our model on Depth Anything V2 and replace its head with an efficient spatial-temporal head. We design a straightforward yet effective temporal consistency loss by constraining the temporal depth gradient, eliminating the need for additional geometric priors. The model is trained on a joint dataset of video depth and unlabeled images, similar to Depth Anything V2. Moreover, a novel key-frame-based strategy is developed for long video inference. Experiments show that our model can be applied to arbitrarily long videos without compromising quality, consistency, or generalization ability. Comprehensive evaluations on multiple video benchmarks demonstrate that our approach sets a new state-of-the-art in zero-shot video depth estimation. We offer models of different scales to support a range of scenarios, with our smallest model capable of real-time performance at 30 FPS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LuMon: A Comprehensive Benchmark and Development Suite with Novel Datasets for Lunar Monocular Depth Estimation

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    A new benchmark with real lunar stereo ground truth and analog data shows that sim-to-real fine-tuned monocular depth models achieve large in-domain gains but minimal generalization to actual lunar images.

  2. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

    cs.CV 2025-12 unverdicted novelty 7.0 of 10

    CARI4D is the first category-agnostic pipeline that produces metric-scale, spatially and temporally consistent 4D reconstructions of human-object interactions from monocular RGB videos via foundation-model hypothesis ...

  3. UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.

  4. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A 0.04B-parameter feed-forward model estimates metric depth from variable calibrated fisheye and pinhole views using calibration tokens and Jacobian distortion bias, with a new multi-view synthetic dataset.

  5. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 conditional novelty 6.0 of 10

    X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.

  6. Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    A self-supervised framework learns implicit 3D physics by lifting V-JEPA features into voxels and performing volumetric feature advection conditioned on actions.

  7. Stabilizing Streaming Video Geometry via Dynamic Feature Normalization

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    DyFN is a lightweight recurrent module that dynamically normalizes latent feature statistics to remove scale-shift drift and achieve state-of-the-art temporal consistency in streaming monocular geometry estimation whi...

  8. UfM*: Uncertainty from Motion* for DNN Depth Estimation Using Gaussians

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    UfM* uses Gaussian mixtures to compute multiview disagreement for uncertainty in depth estimation with single inference per image, reducing energy and memory use.

  9. UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    UniVidX unifies diverse video generation tasks into one conditional diffusion model using stochastic condition masking, decoupled gated LoRAs, and cross-modal self-attention.

  10. WildLIFT: Lifting monocular drone video to 3D for species-agnostic wildlife monitoring

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    WildLIFT lifts monocular drone video to 3D for species-agnostic wildlife detection, tracking, and viewpoint analysis by integrating scene geometry with open-vocabulary segmentation.

  11. GenMatter: Perceiving Physical Objects with Generative Matter Models

    cs.CV 2026-04 conditional novelty 6.0 of 10

    A hierarchical probabilistic model with parallelized Gibbs sampling segments moving matter across random-dot, camouflaged-texture, and naturalistic-video domains, matching supervised baselines and human perceptual judgments.

  12. HOIGS: Human-Object Interaction Gaussian Splatting

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    HOIGS adds a cross-attention HOI module to Gaussian Splatting that combines HexPlane human features with Cubic Hermite Spline object features to model interaction-induced deformations.

  13. From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data

    cs.RO 2026-04 accept novelty 6.0 of 10

    Video-to-robot control methods cluster into three interface families, and the field’s main bottleneck is grounding video-derived predictions into dependable closed-loop robot behavior.

  14. Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    SGC quantifies 3D geometric consistency of generated videos by measuring divergence among local camera poses estimated only on static background sub-regions.

  15. Feedback Matters: Augmenting Autonomous Dissection with Visual and Topological Feedback

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A stretch-based tissue connectivity estimator plus an exposure-maximizing controller and recovery planner raised autonomous dissection success on a da Vinci robot to 80%.

  16. Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    GVF-TAPE predicts future RGB-D frames from an image and text, then extracts end-effector poses to control a robot, achieving strong success rates without action-labeled data.

  17. SpatialTrackerV2: 3D Point Tracking Made Easy

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.

  18. RCG: Safety-Critical Scenario Generation for Robust Autonomous Driving via Real-World Crash Grounding

    cs.RO 2025-07 conditional novelty 6.0 of 10

    RCG replaces handcrafted adversarial scenario scoring with a crash-grounded embedding and k-NN selection, yielding a 9.2% average relative improvement in ego success.

  19. DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation

    cs.CV 2024-11 unverdicted novelty 6.0 of 10

    DissolveStereo injects coarse dissolved depth maps into video diffusion latents via noisy restart and iterative refinement to produce temporally coherent stereo videos zero-shot.

  20. PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    PhysisForcing applies trajectory and relational alignment losses to DiT features in video models, improving physical plausibility on R-Bench, PAI-Bench, and EZS-Bench while raising closed-loop robotic success rates fr...

  21. GenMatter: Perceiving Physical Objects with Generative Matter Models

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    GenMatter is a generative hierarchical model that groups low-level motion and high-level features into particles and clusters representing independently moveable physical entities, validated across dot kinematograms, ...

  22. Controllable Video Object Insertion via Multiview Priors

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    A multi-view prior-based framework for video object insertion that uses dual-path conditioning and an integration-aware consistency module to improve appearance stability and occlusion handling.

  23. From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data

    cs.RO 2026-04 accept novelty 5.0 of 10

    A survey introduces an interface-centric taxonomy for video-to-control methods in robotic manipulation and identifies the robotics integration layer as the central open challenge.

  24. ViPE: Video Pose Engine for 3D Geometric Perception

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.

  25. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

  26. MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Combining joint, limb, RGB, Taylor-video, optical-flow, and depth streams with two video backbones and a validation-tuned weighted ensemble reaches 73.213% top-1 accuracy on iMiGUE, the best MiGA challenge result to date.

Pith tools