Pith. sign in

Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

citation-role summary

background 1 baseline 1

citation-polarity summary

fields

cs.CV 7

representative citing papers

Modality Forcing for Scalable Spatial Generation

cs.CV · 2026-06-11 · unverdicted · novelty 6.0

Modality Forcing lets a single DiT produce image and depth outputs in any order after training on sparse real-world depth, with larger image-pretrained models yielding better depth accuracy and a 57% AbsRel reduction versus prior joint generative baselines.

Depth Anything V2

cs.CV · 2024-06-13 · unverdicted · novelty 6.0

Depth Anything V2 delivers finer, more robust monocular depth predictions by replacing real labeled images with synthetic data, scaling the teacher model, and using large-scale pseudo-labeled real images for student training.

SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry

cs.CV · 2026-06-19 · unverdicted · novelty 5.0

SCOPE uses affine-invariant 3D point maps with shared parameters and three consistency innovations to estimate 3D geometry from extended monocular videos, reporting 24.2% and 34.9% error reductions on ScanNet.

citing papers explorer

Showing 7 of 7 citing papers.

  • DepthMaster: Unified Monocular Depth Estimation for Perspective and Panoramic Images cs.CV · 2026-06-10 · unverdicted · none · ref 37

    DepthMaster unifies metric monocular depth estimation for perspective and panoramic images by patching panoramas into perspective views, adding a consistency loss and virtual cameras, and training mostly on perspective data to reach SOTA zero-shot results on 13 datasets.

  • PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation cs.CV · 2026-07-02 · conditional · none · ref 10

    PixGS is a single-stage pixel-space diffusion model that directly produces high-quality 3D Gaussian Splats from text or images in ~1s, outperforming multi-stage latent methods on standard benchmarks.

  • Modality Forcing for Scalable Spatial Generation cs.CV · 2026-06-11 · unverdicted · none · ref 14

    Modality Forcing lets a single DiT produce image and depth outputs in any order after training on sparse real-world depth, with larger image-pretrained models yielding better depth accuracy and a 57% AbsRel reduction versus prior joint generative baselines.

  • Depth Anything V2 cs.CV · 2024-06-13 · unverdicted · none · ref 20

    Depth Anything V2 delivers finer, more robust monocular depth predictions by replacing real labeled images with synthetic data, scaling the teacher model, and using large-scale pseudo-labeled real images for student training.

  • SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry cs.CV · 2026-06-19 · unverdicted · none · ref 34

    SCOPE uses affine-invariant 3D point maps with shared parameters and three consistency innovations to estimate 3D geometry from extended monocular videos, reporting 24.2% and 34.9% error reductions on ScanNet.

  • DecoRec: Decomposed 3D Scene Reconstruction from Single-View Images via Object-Level Diffusion cs.CV · 2026-05-16 · unverdicted · none · ref 80

    DecoRec decomposes single-view 3D scene reconstruction into per-object diffusion reconstructions followed by a differentiable rendering and diffusion-guided merging pipeline.

  • MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details cs.CV · 2025-07-03 · unverdicted · none · ref 16

    MoGe-2 recovers metric-scale 3D point maps with fine details from single images via data refinement and extension of affine-invariant predictions.