REVIEW 2 cited by
Scene Coordinate Reconstruction: Posing of Image Collections via Incremental Learning of a Relocalizer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We address the task of estimating camera parameters from a set of images depicting a scene. Popular feature-based structure-from-motion (SfM) tools solve this task by incremental reconstruction: they repeat triangulation of sparse 3D points and registration of more camera views to the sparse point cloud. We re-interpret incremental structure-from-motion as an iterated application and refinement of a visual relocalizer, that is, of a method that registers new views to the current state of the reconstruction. This perspective allows us to investigate alternative visual relocalizers that are not rooted in local feature matching. We show that scene coordinate regression, a learning-based relocalization approach, allows us to build implicit, neural scene representations from unposed images. Different from other learning-based reconstruction methods, we do not require pose priors nor sequential inputs, and we optimize efficiently over thousands of images. In many cases, our method, ACE0, estimates camera poses with an accuracy close to feature-based SfM, as demonstrated by novel view synthesis. Project page: https://nianticlabs.github.io/acezero/
Forward citations
Cited by 2 Pith papers
-
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
A diffusion model trained on 1 million 360-degree videos synthesizes novel views with camera translation and enables 3D reconstruction from a single image.
-
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
A deep visual SLAM pipeline, augmented with monocular depth priors, learned motion probability maps, and uncertainty-aware bundle adjustment, estimates camera poses and consistent depths from casual monocular videos o...
Discussion (0). Continue with ORCID to comment.