Pith. sign in

REVIEW 4 cited by

DEFOM-Stereo: Depth Foundation Model Based Stereo Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09466 v3 pith:PK6TPCWI submitted 2025-01-16 cs.CV

classification cs.CV
keywords depthmatchingdefom-stereodisparityestimationfoundationmodelmonocular
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Recently, monocular relative depth estimation has shown remarkable generalization using vision foundation models. Thus, to facilitate robust stereo matching with monocular depth cues, we incorporate a robust monocular relative depth model into the recurrent stereo-matching framework, building a new framework for depth foundation model-based stereo-matching, DEFOM-Stereo. In the feature extraction stage, we construct the combined context and matching feature encoder by integrating features from conventional CNNs and DEFOM. In the update stage, we use the depth predicted by DEFOM to initialize the recurrent disparity and introduce a scale update module to refine the disparity at the correct scale. DEFOM-Stereo is verified to have much stronger zero-shot generalization compared with SOTA methods. Moreover, DEFOM-Stereo achieves top performance on the KITTI 2012, KITTI 2015, Middlebury, and ETH3D benchmarks, ranking $1^{st}$ on many metrics. In the joint evaluation under the robust vision challenge, our model simultaneously outperforms previous models on the individual benchmarks, further demonstrating its outstanding capabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    BenchDepth evaluates eight depth foundation models by their performance on five downstream tasks, finding Depth Anything V2's relative version to be the most practically useful.

  2. PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    One trained model produces competitive stereo, optical flow, feature correspondences, and depth under one checkpoint by recasting all matching tasks as 2D pixel displacement on frozen DINOv2 features.

  3. Performance of universal machine-learned potentials with explicit long-range interactions in biomolecular simulations

    physics.chem-ph 2025-08 unverdicted novelty 5.0 of 10

    The abstract claims a systematic benchmark of universal machine-learned potentials on biomolecular simulations, but the body text supplied is an unrelated stereo-vision paper, leaving the claim unverifiable.

  4. BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A single network that iteratively aligns monocular features with stereo hypotheses reduces zero-shot stereo depth error by over 40% on Middlebury and ETH3D.

Pith tools