Pith. sign in

REVIEW 17 cited by

Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.15506 v4 pith:SYRXEC6L submitted 2024-03-22 cs.CV

classification cs.CV
keywords metricnormaldepthestimationzero-shotcameramodelssurface
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Metric3D v2, a geometric foundation model for zero-shot metric depth and surface normal estimation from a single image, which is crucial for metric 3D recovery. While depth and normal are geometrically related and highly complimentary, they present distinct challenges. SoTA monocular depth methods achieve zero-shot generalization by learning affine-invariant depths, which cannot recover real-world metrics. Meanwhile, SoTA normal estimation methods have limited zero-shot performance due to the lack of large-scale labeled data. To tackle these issues, we propose solutions for both metric depth estimation and surface normal estimation. For metric depth estimation, we show that the key to a zero-shot single-view model lies in resolving the metric ambiguity from various camera models and large-scale data training. We propose a canonical camera space transformation module, which explicitly addresses the ambiguity problem and can be effortlessly plugged into existing monocular models. For surface normal estimation, we propose a joint depth-normal optimization module to distill diverse data knowledge from metric depth, enabling normal estimators to learn beyond normal labels. Equipped with these modules, our depth-normal models can be stably trained with over 16 million of images from thousands of camera models with different-type annotations, resulting in zero-shot generalization to in-the-wild images with unseen camera settings. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. Our project page is at https://JUGGHM.github.io/Metric3Dv2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DepthMaster: Unified Monocular Depth Estimation for Perspective and Panoramic Images

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    DepthMaster unifies metric monocular depth estimation for perspective and panoramic images by patching panoramas into perspective views, adding a consistency loss and virtual cameras, and training mostly on perspectiv...

  2. DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0 of 10

    DVPSFormer runs depth-aware video panoptic segmentation online, using segmentation queries as a scene-discretization prior for a single-pass metric depth head, and beats prior DVPQ scores on Cityscapes-DVPS and SemKITTI-DVPS.

  3. SeeSE3: Emergence of 3D Space in Vision Features

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Self-supervised vision features, especially DINOv2, contain a subspace that a small trained adapter can map to 3D camera motion, enabling pose estimation and latent-space navigation without explicit 3D reconstruction.

  4. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 conditional novelty 6.0 of 10

    X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.

  5. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A 0.04B-parameter feed-forward model estimates metric depth from variable calibrated fisheye and pinhole views using calibration tokens and Jacobian distortion bias, with a new multi-view synthetic dataset.

  6. Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Qwen-RobotWorld is a language-conditioned video world model using Double-Stream MMDiT, an 8.6M-frame embodied corpus, and progressive curriculum training that ranks first on EWMBench and DreamGen Bench.

  7. Towards Consistent Video Geometry Estimation

    cs.CV 2026-05 conditional novelty 6.0 of 10

    One transformer, trained with random-sized temporal attention chunks, unifies offline, streaming, and long-video depth, normal, and point-map estimation and reports new best numbers on five public benchmarks.

  8. PRISM-SLAM: Probabilistic Ray-Grounded Inference for Scale-aware Metric SLAM

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    PRISM-SLAM adds a Plücker Ray-Distance Factor and dynamic uncertainty gating to a VFM-augmented factor graph to deliver scale-consistent metric SLAM at 30 FPS from monocular RGB.

  9. PRISM-SLAM: Probabilistic Ray-Grounded Inference for Scale-aware Metric SLAM

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    PRISM-SLAM achieves scale-aware metric SLAM from RGB input by anchoring VFM depth priors with Plücker ray-distance factors in a factor graph and using dynamic scene uncertainty gating, producing metric trajectories wh...

  10. UAVScenes: A Multi-Modal Dataset for UAVs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    UAVScenes adds frame-wise image and LiDAR semantic labels, reconstructed 6-DoF poses, and 3D maps to 120k frames of the MARS-LVIG dataset, with six benchmark tasks.

  11. Depth Anything V2

    cs.CV 2024-06 unverdicted novelty 6.0 of 10

    Depth Anything V2 delivers finer, more robust monocular depth predictions by replacing real labeled images with synthetic data, scaling the teacher model, and using large-scale pseudo-labeled real images for student training.

  12. Towards Consistent Video Geometry Estimation

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    ViGeo is a feed-forward transformer for video geometry that introduces dynamic chunking attention and a completion-based data refinement framework to achieve SOTA on depth, normals, and point map estimation.

  13. Qwen-Image Technical Report

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Qwen-Image is a foundation model that reaches state-of-the-art results in image generation and editing by combining a large-scale text-focused data pipeline with curriculum learning and dual semantic-reconstructive en...

  14. MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    MoGe-2 recovers metric-scale 3D point maps with fine details from single images via data refinement and extension of affine-invariant predictions.

  15. UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

    cs.CV 2025-02 conditional novelty 5.0 of 10

    UniDepthV2 predicts metric 3D points directly from single images using a self-promptable camera module, pseudo-spherical representation, and new losses for improved cross-domain generalization.

  16. Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction

    cs.CV 2026-07 conditional novelty 4.0 of 10

    NDF treats a fixed-image depth estimator as an implicit field and optimizes it on observed depth at test time, improving inpainting accuracy and cross-view consistency.

  17. LTM: Large-scale Terrain Model for Wildfire-prone Landscapes

    cs.CV 2026-07 reject novelty 4.0 of 10

    A ray-tracing pipeline aligns ground-level image pixels to outdated DEM rasters for real-time 3D terrain reconstruction in wildfire zones, validated primarily through a custom simulator.

Pith tools