REVIEW 26 cited by
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Depth Anything has achieved remarkable success in monocular depth estimation with strong generalization ability. However, it suffers from temporal inconsistency in videos, hindering its practical applications. Various methods have been proposed to alleviate this issue by leveraging video generation models or introducing priors from optical flow and camera poses. Nonetheless, these methods are only applicable to short videos (< 10 seconds) and require a trade-off between quality and computational efficiency. We propose Video Depth Anything for high-quality, consistent depth estimation in super-long videos (over several minutes) without sacrificing efficiency. We base our model on Depth Anything V2 and replace its head with an efficient spatial-temporal head. We design a straightforward yet effective temporal consistency loss by constraining the temporal depth gradient, eliminating the need for additional geometric priors. The model is trained on a joint dataset of video depth and unlabeled images, similar to Depth Anything V2. Moreover, a novel key-frame-based strategy is developed for long video inference. Experiments show that our model can be applied to arbitrarily long videos without compromising quality, consistency, or generalization ability. Comprehensive evaluations on multiple video benchmarks demonstrate that our approach sets a new state-of-the-art in zero-shot video depth estimation. We offer models of different scales to support a range of scenarios, with our smallest model capable of real-time performance at 30 FPS.
Forward citations
Cited by 26 Pith papers
-
LuMon: A Comprehensive Benchmark and Development Suite with Novel Datasets for Lunar Monocular Depth Estimation
A new benchmark with real lunar stereo ground truth and analog data shows that sim-to-real fine-tuned monocular depth models achieve large in-domain gains but minimal generalization to actual lunar images.
-
CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
CARI4D is the first category-agnostic pipeline that produces metric-scale, spatially and temporally consistent 4D reconstructions of human-object interactions from monocular RGB videos via foundation-model hypothesis ...
-
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.
-
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras
A 0.04B-parameter feed-forward model estimates metric depth from variable calibrated fisheye and pinhole views using calibration tokens and Jacobian distortion bias, with a new multi-view synthetic dataset.
-
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras
X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.
-
Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection
A self-supervised framework learns implicit 3D physics by lifting V-JEPA features into voxels and performing volumetric feature advection conditioned on actions.
-
Stabilizing Streaming Video Geometry via Dynamic Feature Normalization
DyFN is a lightweight recurrent module that dynamically normalizes latent feature statistics to remove scale-shift drift and achieve state-of-the-art temporal consistency in streaming monocular geometry estimation whi...
-
UfM*: Uncertainty from Motion* for DNN Depth Estimation Using Gaussians
UfM* uses Gaussian mixtures to compute multiview disagreement for uncertainty in depth estimation with single inference per image, reducing energy and memory use.
-
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
UniVidX unifies diverse video generation tasks into one conditional diffusion model using stochastic condition masking, decoupled gated LoRAs, and cross-modal self-attention.
-
WildLIFT: Lifting monocular drone video to 3D for species-agnostic wildlife monitoring
WildLIFT lifts monocular drone video to 3D for species-agnostic wildlife detection, tracking, and viewpoint analysis by integrating scene geometry with open-vocabulary segmentation.
-
GenMatter: Perceiving Physical Objects with Generative Matter Models
A hierarchical probabilistic model with parallelized Gibbs sampling segments moving matter across random-dot, camouflaged-texture, and naturalistic-video domains, matching supervised baselines and human perceptual judgments.
-
HOIGS: Human-Object Interaction Gaussian Splatting
HOIGS adds a cross-attention HOI module to Gaussian Splatting that combines HexPlane human features with Cubic Hermite Spline object features to model interaction-induced deformations.
-
From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data
Video-to-robot control methods cluster into three interface families, and the field’s main bottleneck is grounding video-derived predictions into dependable closed-loop robot behavior.
-
Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation
SGC quantifies 3D geometric consistency of generated videos by measuring divergence among local camera poses estimated only on static background sub-regions.
-
Feedback Matters: Augmenting Autonomous Dissection with Visual and Topological Feedback
A stretch-based tissue connectivity estimator plus an exposure-maximizing controller and recovery planner raised autonomous dissection success on a da Vinci robot to 80%.
-
Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
GVF-TAPE predicts future RGB-D frames from an image and text, then extracts end-effector poses to control a robot, achieving strong success rates without action-labeled data.
-
SpatialTrackerV2: 3D Point Tracking Made Easy
A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.
-
RCG: Safety-Critical Scenario Generation for Robust Autonomous Driving via Real-World Crash Grounding
RCG replaces handcrafted adversarial scenario scoring with a crash-grounded embedding and k-NN selection, yielding a 9.2% average relative improvement in ego success.
-
DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation
DissolveStereo injects coarse dissolved depth maps into video diffusion latents via noisy restart and iterative refinement to produce temporally coherent stereo videos zero-shot.
-
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
PhysisForcing applies trajectory and relational alignment losses to DiT features in video models, improving physical plausibility on R-Bench, PAI-Bench, and EZS-Bench while raising closed-loop robotic success rates fr...
-
GenMatter: Perceiving Physical Objects with Generative Matter Models
GenMatter is a generative hierarchical model that groups low-level motion and high-level features into particles and clusters representing independently moveable physical entities, validated across dot kinematograms, ...
-
Controllable Video Object Insertion via Multiview Priors
A multi-view prior-based framework for video object insertion that uses dual-path conditioning and an integration-aware consistency module to improve appearance stability and occlusion handling.
-
From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data
A survey introduces an interface-centric taxonomy for video-to-control methods in robotic manipulation and identifies the robotics integration layer as the central open challenge.
-
ViPE: Video Pose Engine for 3D Geometric Perception
ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.
-
Reconstructing 4D Spatial Intelligence: A Survey
A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.
-
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
Combining joint, limb, RGB, Taylor-video, optical-flow, and depth streams with two video backbones and a validation-tuned weighted ensemble reaches 73.213% top-1 accuracy on iMiGUE, the best MiGA challenge result to date.
Discussion (0). Sign in to comment.