Pith. sign in

REVIEW 27 cited by

From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.10326 v6 pith:ZD3C52JA submitted 2019-07-24 cs.CV

From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation

classification cs.CV
keywords depthguidancenetworkschallengingconvolutionaldensedesiredeffective
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Estimating accurate depth from a single image is challenging because it is an ill-posed problem as infinitely many 3D scenes can be projected to the same 2D scene. However, recent works based on deep convolutional neural networks show great progress with plausible results. The convolutional neural networks are generally composed of two parts: an encoder for dense feature extraction and a decoder for predicting the desired depth. In the encoder-decoder schemes, repeated strided convolution and spatial pooling layers lower the spatial resolution of transitional outputs, and several techniques such as skip connections or multi-layer deconvolutional networks are adopted to recover the original resolution for effective dense prediction. In this paper, for more effective guidance of densely encoded features to the desired depth prediction, we propose a network architecture that utilizes novel local planar guidance layers located at multiple stages in the decoding phase. We show that the proposed method outperforms the state-of-the-art works with significant margin evaluating on challenging benchmarks. We also provide results from an ablation study to validate the effectiveness of the proposed method.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation

    cs.CV 2026-06 unverdicted novelty 7.0

    DrivingDepth achieves SOTA metric depth on nuScenes by residual pixel-wise scale correction on frozen foundation models using sparse LiDAR prompts, preserving geometric consistency.

  2. SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping

    cs.CV 2026-05 unverdicted novelty 7.0

    SeeGroup formulates per-pixel multi-layer depth as a point process with permutation-invariant likelihood to support arbitrary groupings, raising quadruplet relative depth accuracy from 61.34% to 70.09% on the LayeredD...

  3. VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

    cs.CL 2026-05 unverdicted novelty 7.0

    VITA-QinYu is the first expressive end-to-end spoken language model supporting role-playing and singing alongside conversation, trained on 15.8K hours of data and outperforming prior models on expressiveness and conve...

  4. VDPP: Video Depth Post-Processing for Speed and Scalability

    cs.CV 2026-04 unverdicted novelty 7.0

    VDPP is an RGB-free video depth post-processor that achieves over 43 FPS on Jetson Orin Nano by refining geometry at low resolution rather than reconstructing full scenes.

  5. LiftFormer: Lifting and Frame Theory Based Monocular Depth Estimation Using Depth and Edge Oriented Subspace Representation

    cs.CV 2026-04 unverdicted novelty 7.0

    LiftFormer transforms monocular depth prediction into depth-oriented geometric and edge-aware subspace representations via lifting and frame theory, achieving state-of-the-art results on standard datasets.

  6. UniDAC: Universal Metric Depth Estimation for Any Camera

    cs.CV 2026-03 unverdicted novelty 7.0

    UniDAC achieves universal metric depth estimation across camera types by decoupling relative depth prediction from spatially varying scale estimation using a depth-guided module and distortion-aware positional embedding.

  7. Efficient Test-Time Optimization for Depth Completion via Low-Rank Decoder Adaptation

    cs.CV 2026-03 unverdicted novelty 7.0

    Low-rank decoder adaptation enables efficient test-time optimization for zero-shot depth completion by updating only the subspace containing depth-relevant information.

  8. RAD: Retrieval-Augmented Monocular Metric Depth Estimation for Underrepresented Classes

    cs.CV 2026-02 unverdicted novelty 7.0

    RAD retrieves semantically similar RGB-D context samples for low-confidence regions and fuses them via matched cross-attention to cut relative absolute depth error by 29.2% on NYU Depth v2 underrepresented classes whi...

  9. ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

    cs.CV 2023-02 accept novelty 7.0

    ZoeDepth combines relative depth pre-training on many datasets with metric depth fine-tuning and automatic head routing to achieve strong zero-shot generalization while preserving metric scale.

  10. The Multipath Blind Spot: $K$-Agnostic Robust Calibration for Sparse-Anchor Metric Depth from Frozen Foundations

    cs.CV 2026-07 accept novelty 6.5

    MRAC gates sparse anchors via Theil–Sen + MAD consistency with a frozen foundation's relative depth, repairing multipath outliers that collapse residual-on-CFA and blind VI-Depth while winning 84% of same-backbone cells.

  11. Twins: Learn to Predict Unified Representations with Focal Loss

    cs.CV 2026-07 conditional novelty 6.0

    Channel-wise concatenation of SigLIP2 and Flux VAE features into one token, trained with a focal-style flow-matching loss, yields a unified representation with 1.59 gFID on ImageNet 256 and VAE-level reconstruction.

  12. Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0

    Uncertainty-weighted multi-teacher distillation plus dense bird's-eye-view radar fusion improves self-supervised depth estimation under adverse weather, cutting night absRel by ~23% on nuScenes.

  13. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 conditional novelty 6.0

    X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.

  14. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 unverdicted novelty 6.0

    A 0.04B-parameter feed-forward model estimates metric depth from variable calibrated fisheye and pinhole views using calibration tokens and Jacobian distortion bias, with a new multi-view synthetic dataset.

  15. HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

    cs.CV 2026-06 unverdicted novelty 6.0

    HYDRA-X presents the first unified multimodal model using a single ViT for holistic image-video tokenization, with ablations on attention and compression plus a latent-level editing improvement.

  16. Efficient Hybrid CNN-GNN Architecture for Monocular Depth Estimation

    cs.CV 2026-05 unverdicted novelty 6.0

    GraphDepth integrates multi-scale GraphSAGE layers into a ResNet-101 U-Net with attention-gated skips and uncertainty estimation to deliver competitive monocular depth accuracy at lower computational cost than transformers.

  17. Last-Layer-Centric Feature Recombination: Unleashing 3D Geometric Knowledge in DINOv3 for Monocular Depth Estimation

    cs.CV 2026-04 unverdicted novelty 6.0

    Layer analysis of DINOv3 shows non-uniform 3D geometric knowledge concentrated in deeper layers, enabling a last-layer-centric recombination module that improves monocular depth estimation accuracy to state-of-the-art levels.

  18. G3Splat: Geometrically Consistent Generalizable Gaussian Splatting

    cs.CV 2025-12 conditional novelty 6.0

    Adding ray-alignment and local-normal orientation losses to generalizable Gaussian splatting fixes geometrically degenerate splats and improves zero-shot depth, mesh, and pose estimation.

  19. Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

    cs.CV 2025-12 conditional novelty 6.0

    Across 69 monocular depth estimators, human-likeness of error patterns peaks near human-level accuracy and declines for the most accurate models: accuracy does not guarantee human-like depth perception.

  20. XD-RCDepth: Lightweight Radar-Camera Depth Estimation with Explainability-Aligned and Distribution-Aware Distillation

    cs.CV 2025-10 conditional novelty 6.0

    XD-RCDepth uses explainability-aligned and depth-distribution distillation to shrink a radar-camera depth model by 29.7% parameters while improving MAE by about 8%.

  21. Radar-Guided Polynomial Fitting for Metric Depth Estimation

    cs.CV 2025-03 unverdicted novelty 6.0

    POLAR converts scaleless monocular depth maps to metric scale via radar-guided polynomial fitting and first-derivative regularization, claiming 24.9% MAE and 33.2% RMSE gains over prior methods on three datasets.

  22. DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV

    cs.CV 2026-07 conditional novelty 5.0

    DAPM jointly estimates depth and camera pose from single drone images by injecting an ideal-ground-plane depth prior and progressive per-pixel depth bins, trained on a new 42k simulated UAV dataset.

  23. VolFill: Single-View Amodal 3D Scene Reconstruction with Volumetric Flow Matching

    cs.CV 2026-05 unverdicted novelty 5.0

    VolFill uses a hybrid 3D VAE to compress sparse truncated unsigned distance function grids into latent space and a latent Diffusion Transformer to denoise complete scenes, conditioned on geometry foundation models, ou...

  24. Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction

    cs.CV 2026-04 unverdicted novelty 5.0

    A multilevel perceptual CRF model using Swin Transformer, HPF fusion, HA adapters, and dynamic scaling attention achieves state-of-the-art monocular depth estimation on NYU Depth v2, KITTI, and MatterPort3D with reduc...

  25. ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving

    cs.CV 2025-08 unverdicted novelty 5.0

    ROVR is a new diverse depth dataset for autonomous driving with 200K frames, released pipelines, and ablations showing sparse ground truth supports model training.

  26. UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

    cs.CV 2025-02 conditional novelty 5.0

    UniDepthV2 predicts metric 3D points directly from single images using a self-promptable camera module, pseudo-spherical representation, and new losses for improved cross-domain generalization.

  27. Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

    cs.RO 2025-08 conditional novelty 4.0

    The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.