REVIEW 4 cited by
FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth Estimators
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Matching cross-modality features between images and point clouds is a fundamental problem for image-to-point cloud registration. However, due to the modality difference between images and points, it is difficult to learn robust and discriminative cross-modality features by existing metric learning methods for feature matching. Instead of applying metric learning on cross-modality data, we propose to unify the modality between images and point clouds by pretrained large-scale models first, and then establish robust correspondence within the same modality. We show that the intermediate features, called diffusion features, extracted by depth-to-image diffusion models are semantically consistent between images and point clouds, which enables the building of coarse but robust cross-modality correspondences. We further extract geometric features on depth maps produced by the monocular depth estimator. By matching such geometric features, we significantly improve the accuracy of the coarse correspondences produced by diffusion features. Extensive experiments demonstrate that without any task-specific training, direct utilization of both features produces accurate image-to-point cloud registration. On three public indoor and outdoor benchmarks, the proposed method averagely achieves a 20.6 percent improvement in Inlier Ratio, a three-fold higher Inlier Number, and a 48.6 percent improvement in Registration Recall than existing state-of-the-arts.
Forward citations
Cited by 4 Pith papers
-
Image-to-Point Cloud Registration Made Easy with Rectified Flow-based LiDAR Upsampling
A rectified-flow upsampler turns one sparse LiDAR scan into a dense intensity image that stock feature matchers align to camera images, yielding 6-DoF poses (4.89°/1.63 m mean error on R3LIVE) without training on the ...
-
Diff$^2$I2P: Differentiable Image-to-Point Cloud Registration with Diffusion Prior
Diff2I2P improves cross-modal registration recall to 83.0% on 7-Scenes, up from 75.8% for the 2D3D-MATR baseline, by distilling a diffusion prior into feature learning via score distillation.
-
CA-I2P: Channel-Adaptive Registration Network with Global Optimal Selection
CA-I2P combines channel-adaptive feature adjustment and optimal-transport-based global selection to improve image-to-point cloud registration, reporting SOTA recall on RGB-D Scenes V2 (63.3% RR) and 7-Scenes (79.5% RR).
-
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.
Discussion (0). Sign in to comment.