Pith. sign in

REVIEW 7 cited by

IGEV++: Iterative Multi-range Geometry Encoding Volumes for Stereo Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00638 v3 pith:BJOI7G6X submitted 2024-09-01 cs.CV

classification cs.CV
keywords igevgeometrymatchinglargeregionsachievesdisparitiesdisparity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Stereo matching is a core component in many computer vision and robotics systems. Despite significant advances over the last decade, handling matching ambiguities in ill-posed regions and large disparities remains an open challenge. In this paper, we propose a new deep network architecture, called IGEV++, for stereo matching. The proposed IGEV++ constructs Multi-range Geometry Encoding Volumes (MGEV), which encode coarse-grained geometry information for ill-posed regions and large disparities, while preserving fine-grained geometry information for details and small disparities. To construct MGEV, we introduce an adaptive patch matching module that efficiently and effectively computes matching costs for large disparity ranges and/or ill-posed regions. We further propose a selective geometry feature fusion module to adaptively fuse multi-range and multi-granularity geometry features in MGEV. Then, we input the fused geometry features into ConvGRUs to iteratively update the disparity map. MGEV allows to efficiently handle large disparities and ill-posed regions, such as occlusions and textureless regions, and enjoys rapid convergence during iterations. Our IGEV++ achieves the best performance on the Scene Flow test set across all disparity ranges, up to 768px. Our IGEV++ also achieves state-of-the-art accuracy on the Middlebury, ETH3D, KITTI 2012, and 2015 benchmarks. Specifically, IGEV++ achieves a 3.23\% 2-pixel outlier rate (Bad 2.0) on the large disparity benchmark, Middlebury, representing error reductions of 31.9\% and 54.8\% compared to RAFT-Stereo and GMStereo, respectively. We also present a real-time version of IGEV++ that achieves the best performance among all published real-time methods on the KITTI benchmarks. The code is publicly available at https://github.com/gangweix/IGEV and https://github.com/gangweix/IGEV-plusplus.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GeoStereo: A Unified Stereo Geometry Estimation Framework for Disparity and Surface Normal

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A unified stereo framework couples feed-forward disparity matching with a diffusion-based normal estimator through disparity-to-normal initialization and warped right-view conditioning, claiming zero-shot SOTA on seve...

  2. WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching

    cs.CV 2026-03 accept novelty 6.0 of 10

    Warping alone, with a classification head before iterative regression, matches or beats cost-volume stereo methods on ETH3D, KITTI and Middlebury at higher speed.

  3. Iterative Volume Fusion for Asymmetric Stereo Matching

    cs.CV 2025-08 conditional novelty 6.0 of 10

    IVF-AStereo fuses correlation and concatenation cost volumes in two phases to keep stereo disparity accurate when camera views differ in resolution or color.

  4. STEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow Matching

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A hybrid stereo-matching model uses a cascade matching network to propose disparities and a diffusion transformer to refine ambiguous regions; it claims state-of-the-art benchmark results.

  5. BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A single network that iteratively aligns monocular features with stereo hypotheses reduces zero-shot stereo depth error by over 40% on Middlebury and ETH3D.

  6. Stereo 3D Gaussian Splatting SLAM for Outdoor Urban Scenes

    cs.RO 2025-07 conditional novelty 5.0 of 10

    BGS-SLAM combines ORB-SLAM2 tracking with 3D Gaussian splatting mapping supervised by pretrained stereo depth networks, and reports improved outdoor mapping and tracking over the tested 3DGS-SLAM baselines.

  7. ESMStereo: Enhanced ShuffleMixer Disparity Upsampling for Real-Time and Accurate Stereo Matching

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A real-time stereo matching architecture whose Enhanced ShuffleMixer upsampler fuses disparity and image features to recover detail lost by compact cost volumes, reaching state-of-the-art speed-accuracy trade-offs.

Pith tools