REVIEW 17 cited by
Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce Metric3D v2, a geometric foundation model for zero-shot metric depth and surface normal estimation from a single image, which is crucial for metric 3D recovery. While depth and normal are geometrically related and highly complimentary, they present distinct challenges. SoTA monocular depth methods achieve zero-shot generalization by learning affine-invariant depths, which cannot recover real-world metrics. Meanwhile, SoTA normal estimation methods have limited zero-shot performance due to the lack of large-scale labeled data. To tackle these issues, we propose solutions for both metric depth estimation and surface normal estimation. For metric depth estimation, we show that the key to a zero-shot single-view model lies in resolving the metric ambiguity from various camera models and large-scale data training. We propose a canonical camera space transformation module, which explicitly addresses the ambiguity problem and can be effortlessly plugged into existing monocular models. For surface normal estimation, we propose a joint depth-normal optimization module to distill diverse data knowledge from metric depth, enabling normal estimators to learn beyond normal labels. Equipped with these modules, our depth-normal models can be stably trained with over 16 million of images from thousands of camera models with different-type annotations, resulting in zero-shot generalization to in-the-wild images with unseen camera settings. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. Our project page is at https://JUGGHM.github.io/Metric3Dv2.
Forward citations
Cited by 17 Pith papers
-
DepthMaster: Unified Monocular Depth Estimation for Perspective and Panoramic Images
DepthMaster unifies metric monocular depth estimation for perspective and panoramic images by patching panoramas into perspective views, adding a consistency loss and virtual cameras, and training mostly on perspectiv...
-
DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
DVPSFormer runs depth-aware video panoptic segmentation online, using segmentation queries as a scene-discretization prior for a single-pass metric depth head, and beats prior DVPQ scores on Cityscapes-DVPS and SemKITTI-DVPS.
-
SeeSE3: Emergence of 3D Space in Vision Features
Self-supervised vision features, especially DINOv2, contain a subspace that a small trained adapter can map to 3D camera motion, enabling pose estimation and latent-space navigation without explicit 3D reconstruction.
-
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras
X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.
-
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras
A 0.04B-parameter feed-forward model estimates metric depth from variable calibrated fisheye and pinhole views using calibration tokens and Jacobian distortion bias, with a new multi-view synthetic dataset.
-
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
Qwen-RobotWorld is a language-conditioned video world model using Double-Stream MMDiT, an 8.6M-frame embodied corpus, and progressive curriculum training that ranks first on EWMBench and DreamGen Bench.
-
Towards Consistent Video Geometry Estimation
One transformer, trained with random-sized temporal attention chunks, unifies offline, streaming, and long-video depth, normal, and point-map estimation and reports new best numbers on five public benchmarks.
-
PRISM-SLAM: Probabilistic Ray-Grounded Inference for Scale-aware Metric SLAM
PRISM-SLAM adds a Plücker Ray-Distance Factor and dynamic uncertainty gating to a VFM-augmented factor graph to deliver scale-consistent metric SLAM at 30 FPS from monocular RGB.
-
PRISM-SLAM: Probabilistic Ray-Grounded Inference for Scale-aware Metric SLAM
PRISM-SLAM achieves scale-aware metric SLAM from RGB input by anchoring VFM depth priors with Plücker ray-distance factors in a factor graph and using dynamic scene uncertainty gating, producing metric trajectories wh...
-
UAVScenes: A Multi-Modal Dataset for UAVs
UAVScenes adds frame-wise image and LiDAR semantic labels, reconstructed 6-DoF poses, and 3D maps to 120k frames of the MARS-LVIG dataset, with six benchmark tasks.
-
Depth Anything V2
Depth Anything V2 delivers finer, more robust monocular depth predictions by replacing real labeled images with synthetic data, scaling the teacher model, and using large-scale pseudo-labeled real images for student training.
-
Towards Consistent Video Geometry Estimation
ViGeo is a feed-forward transformer for video geometry that introduces dynamic chunking attention and a completion-based data refinement framework to achieve SOTA on depth, normals, and point map estimation.
-
Qwen-Image Technical Report
Qwen-Image is a foundation model that reaches state-of-the-art results in image generation and editing by combining a large-scale text-focused data pipeline with curriculum learning and dual semantic-reconstructive en...
-
MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
MoGe-2 recovers metric-scale 3D point maps with fine details from single images via data refinement and extension of affine-invariant predictions.
-
UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler
UniDepthV2 predicts metric 3D points directly from single images using a self-promptable camera module, pseudo-spherical representation, and new losses for improved cross-domain generalization.
-
Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction
NDF treats a fixed-image depth estimator as an implicit field and optimizes it on observed depth at test time, improving inpainting accuracy and cross-view consistency.
-
LTM: Large-scale Terrain Model for Wildfire-prone Landscapes
A ray-tracing pipeline aligns ground-level image pixels to outdated DEM rasters for real-time 3D terrain reconstruction in wildfire zones, validated primarily through a custom simulator.
Discussion (0). Sign in to comment.