Pith. sign in

REVIEW 6 cited by

Monocular Depth Estimation using Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.14816 v1 pith:FDFAB6ST submitted 2023-02-28 cs.CV

classification cs.CV
keywords depthdiffusiontrainingdatadatasetdenoisingdepthgenestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We formulate monocular depth estimation using denoising diffusion models, inspired by their recent successes in high fidelity image generation. To that end, we introduce innovations to address problems arising due to noisy, incomplete depth maps in training data, including step-unrolled denoising diffusion, an $L_1$ loss, and depth infilling during training. To cope with the limited availability of data for supervised training, we leverage pre-training on self-supervised image-to-image translation tasks. Despite the simplicity of the approach, with a generic loss and architecture, our DepthGen model achieves SOTA performance on the indoor NYU dataset, and near SOTA results on the outdoor KITTI dataset. Further, with a multimodal posterior, DepthGen naturally represents depth ambiguity (e.g., from transparent surfaces), and its zero-shot performance combined with depth imputation, enable a simple but effective text-to-3D pipeline. Project page: https://depth-gen.github.io

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal Score Rescaling for Temperature Sampling in Diffusion and Flow Models

    cs.LG 2025-10 conditional novelty 6.0 of 10

    TSR multiplies a pre-trained diffusion/flow model's score by r_t=(η_tσ²+1)/(η_tσ²/k+1), exactly sharpening Gaussian data by 1/k and empirically steering sample diversity on five tasks without retraining.

  2. Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage diffusion model generates realistic depth noise on synthetic CAD data, and pretraining 3D networks on the resulting data improves few-shot real-world 3D tasks.

  3. DreamCube: 3D Panorama Generation via Multi-plane Synchronization

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A synchronized multi-plane adaptation of 2D diffusion operators enables seam-consistent cubemap generation, and DreamCube extends this to joint RGB-D panorama generation and 3D scene lifting.

  4. FMOcc: TPV-Driven Flow Matching for 3D Occupancy Prediction with Selective State Space Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FMOcc uses flow matching with tri-perspective view and selective state space layers to improve 3D occupancy prediction from two camera frames.

  5. Depth Anything at Any Condition

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A fine-tuned Depth Anything V2 model using perturbation consistency and spatial distance constraints improves monocular depth estimation under adverse conditions without any labeled data.

  6. Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A reparameterization recipe that lets pre-trained Stable Diffusion checkpoints be finetuned as flow matching models, giving faster convergence and better performance under parameter-efficient constraints.

Pith tools