Pith. sign in

REVIEW 9 cited by

BinsFormer: Revisiting Adaptive Bins for Monocular Depth Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.00987 v1 pith:M6YPIT6U submitted 2022-04-03 cs.CV

classification cs.CV
keywords depthbinsestimationtaskbinsformermonocularadaptivedecoder
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Monocular depth estimation is a fundamental task in computer vision and has drawn increasing attention. Recently, some methods reformulate it as a classification-regression task to boost the model performance, where continuous depth is estimated via a linear combination of predicted probability distributions and discrete bins. In this paper, we present a novel framework called BinsFormer, tailored for the classification-regression-based depth estimation. It mainly focuses on two crucial components in the specific task: 1) proper generation of adaptive bins and 2) sufficient interaction between probability distribution and bins predictions. To specify, we employ the Transformer decoder to generate bins, novelly viewing it as a direct set-to-set prediction problem. We further integrate a multi-scale decoder structure to achieve a comprehensive understanding of spatial geometry information and estimate depth maps in a coarse-to-fine manner. Moreover, an extra scene understanding query is proposed to improve the estimation accuracy, which turns out that models can implicitly learn useful information from an auxiliary environment classification task. Extensive experiments on the KITTI, NYU, and SUN RGB-D datasets demonstrate that BinsFormer surpasses state-of-the-art monocular depth estimation methods with prominent margins. Code and pretrained models will be made publicly available at \url{https://github.com/zhyever/Monocular-Depth-Estimation-Toolbox}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    BenchDepth evaluates eight depth foundation models by their performance on five downstream tasks, finding Depth Anything V2's relative version to be the most practically useful.

  2. Enhancing Monocular Depth Estimation with Multi-Source Auxiliary Tasks

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Training a frozen DINOv2 backbone with a shared decoder on auxiliary multi-label dense classification improves monocular depth accuracy by about 11 percent on in-domain benchmarks.

  3. LiRCDepth: Lightweight Radar-Camera Depth Estimation via Knowledge Distillation and Uncertainty Guidance

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LiRCDepth shows that a MobileNetV2-based radar-camera depth estimator can recover much of the accuracy of a ResNet-based teacher via feature, structure, and uncertainty-weighted depth distillation.

  4. Amodal Depth Anything: Amodal Depth Estimation in the Wild

    cs.CV 2024-12 conditional novelty 6.0 of 10

    The paper introduces ADIW, a 564K-image pseudo-labeled dataset for relative amodal depth, and two fine-tuned models (Amodal-DAV2 and Amodal-DepthFM) that predict occluded-object depth from an image, observed depth, an...

  5. Video Depth without Video Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A single-image latent diffusion model extended with cross-frame attention and global scale-shift alignment produces state-of-the-art video depth without a video diffusion model.

  6. Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Ego Scene Augmentation boosts egocentric VQA accuracy by 8.14% (indoor) and 8.72% (outdoor) by injecting a Depth-Anything-derived object/depth/text scene graph into the MLLM prompt.

  7. Region-aware Depth Scale Adaptation with Sparse Measurements

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A non-learning method segments an image and gives each region its own scale and shift, fitted to a few sparse depth points, to turn relative monocular depth predictions into metric depth more accurately than a single ...

  8. CoL3D: Collaborative Learning of Single-view Depth and Camera Intrinsics for Metric 3D Shape Recovery

    cs.CV 2025-02 conditional novelty 5.0 of 10

    CoL3D jointly learns monocular depth and camera intrinsics with a canonical incidence field and a point-cloud shape loss, recovering metric 3D shape without intrinsics at inference.

  9. PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A lightweight refiner with a coarse-to-fine denoising module, noise-based pretraining, and a scale-shift invariant gradient-matching loss achieves state-of-the-art high-resolution metric depth with up to 10x faster inference.

Pith tools