Pith. sign in

REVIEW 25 cited by

MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19115 v3 pith:WKOY7H5W submitted 2024-10-24 cs.CV

classification cs.CV
keywords geometrymodelgloballocalmonocularpointsupervisionaccurate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present MoGe, a powerful model for recovering 3D geometry from monocular open-domain images. Given a single image, our model directly predicts a 3D point map of the captured scene with an affine-invariant representation, which is agnostic to true global scale and shift. This new representation precludes ambiguous supervision in training and facilitate effective geometry learning. Furthermore, we propose a set of novel global and local geometry supervisions that empower the model to learn high-quality geometry. These include a robust, optimal, and efficient point cloud alignment solver for accurate global shape learning, and a multi-scale local geometry loss promoting precise local geometry supervision. We train our model on a large, mixed dataset and demonstrate its strong generalizability and high accuracy. In our comprehensive evaluation on diverse unseen datasets, our model significantly outperforms state-of-the-art methods across all tasks, including monocular estimation of 3D point map, depth map, and camera field of view. Code and models can be found on our project page.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Iterative sparse-3D-convolution refinement in a log-depth voxel shell, instead of 2D image-plane refinement, sharply improves fine-detail geometry in monocular point maps and sets state of the art on local fine-detail...

  2. Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots

    cs.RO 2025-09 conditional novelty 7.0 of 10

    A learned plug-in that denoises consumer depth cameras to simulation-like metric depth enables zero-shot sim-to-real transfer of depth-only manipulation policies trained on raw simulated depth.

  3. Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A procedural-generation benchmark (PDE) shows that depth models are surprisingly vulnerable to camera changes and occlusion, while resisting lighting changes.

  4. Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Rig3R conditions learned 3D reconstruction on optional rig metadata and predicts rig-relative raymaps, enabling state-of-the-art pose estimation and rig calibration discovery from images.

  5. MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation

    cs.CV 2025-02 conditional novelty 7.0 of 10

    A controllable image-to-video method that jointly drives camera and object motion by translating scene-space user designs into DCT-coded point trajectories and color-coded bounding boxes for a DiT-based diffusion model.

  6. Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Qwen-3D feeds the full visual-language representation of a 3D scene into a mask decoder, beating prior 3D LMMs on grounding and segmentation benchmarks.

  7. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.

  8. VEGA: Learning Navigation VLAs from In-the-Wild Egocentric Video with Geometric Trajectory Supervision

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    VEGA reconstructs local geometry from monocular egocentric video to create supervised trajectories that train a flow-matching VLA policy, yielding lower collision rates on a new benchmark and in real-world tests.

  9. Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A diffusion model guided by hand-object interaction and geometric cues reconstructs 3D hand-held object geometry from monocular RGB images.

  10. BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    BenchDepth evaluates eight depth foundation models by their performance on five downstream tasks, finding Depth Anything V2's relative version to be the most practically useful.

  11. Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Puzzles synthesizes posed video-depth clips from single images and keyframes, letting 3D reconstruction models match full-data accuracy using only 10% of the data.

  12. E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    E3D-Bench compares 16 3D geometric foundation models on depth, reconstruction, pose, and view-synthesis tasks with a unified evaluation toolkit.

  13. UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Fine-tuning a pretrained video diffusion transformer to predict geometry in one shared global frame produces consistent, camera-free surface normals and coordinates across entire video clips.

  14. Large-scale visual SLAM for in-the-wild videos

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A visual SLAM pipeline integrating deep flow, semantic masking, monocular depth regularization, and loop closure produces one continuous 3D trajectory from casual 15-minute videos where COLMAP and GLOMAP fragment or distort.

  15. LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A single feed-forward network predicts layered ray-surface intersection point maps and a ray stopping index, enabling efficient single-view reconstruction of visible and occluded geometry.

  16. SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A neural system that reconstructs dense 3D scenes from monocular RGB video at 20+ FPS by regressing local pointmaps and incrementally registering them into one global model without explicit pose optimization.

  17. Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Reloc3r trains a symmetric, scale-free relative pose regression transformer on 8M image pairs and uses motion averaging for absolute poses, outperforming prior regression methods on six localization benchmarks.

  18. DepthCues: Evaluating Monocular Depth Perception in Large Vision Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A six-task benchmark shows newer large vision models encode human-like monocular depth cues, and cue understanding strongly correlates with their depth estimation performance.

  19. Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues

    cs.CV 2026-08 conditional novelty 5.0 of 10

    Harris eigenvalues and spatiotemporal density values from event cameras encode motion direction and, when added to an optical flow network, improve accuracy in data-scarce settings.

  20. ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A 6.1M-parameter monocular depth network, distilled from Depth Anything v2-Large over 14.1M multi-domain images, achieves the best zero-shot accuracy–efficiency trade-off among lightweight models across five benchmark...

  21. Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction

    cs.CV 2025-04 conditional novelty 5.0 of 10

    Mono3R adds a monocular-guided refinement module to DUSt3R, aligning frozen MoGe pointmaps and features with pairwise predictions and iteratively updating them, yielding better pose estimates and denser point clouds i...

  22. Exploring Representation-Aligned Latent Space for Better Generation

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Aligning VAE latents with DINOv2 semantic features improves latent diffusion model image generation on ImageNet by about 15% FID.

  23. Dance Style Recognition Using Laban Movement Analysis

    cs.CV 2025-04 conditional novelty 4.0 of 10

    A sliding-window Laban Movement Analysis pipeline from 3D pose and mesh estimates reports 99%+ accuracy on AIST++ dance styles, though the cross-validation split and a non-temporal baseline are not clearly documented.

  24. The Fourth Monocular Depth Estimation Challenge

    cs.CV 2025-04 conditional novelty 4.0 of 10

    The fourth MDEC winner, HRI, achieved a 3D F-score of 23.05% on SYNS-Patches, a small improvement over the previous best of 22.58%, under a new two-degree-of-freedom alignment protocol.

  25. The First WARA Robotics Mobile Manipulation Challenge -- Lessons Learned

    cs.RO 2025-05 unverdicted novelty 2.0 of 10

    Four university teams developed mobile manipulation systems for lab glassware transport and dishwasher loading; the paper reports their approaches and lessons for future challenge editions.

Pith tools