Pith. sign in

REVIEW 14 cited by

MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.19152 v1 pith:QOPCMQA3 submitted 2024-09-27 cs.CV

classification cs.CV
keywords imagesfoundationlocalmethodsoverallpipelinereconstructionssettings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Structure-from-Motion (SfM), a task aiming at jointly recovering camera poses and 3D geometry of a scene given a set of images, remains a hard problem with still many open challenges despite decades of significant progress. The traditional solution for SfM consists of a complex pipeline of minimal solvers which tends to propagate errors and fails when images do not sufficiently overlap, have too little motion, etc. Recent methods have attempted to revisit this paradigm, but we empirically show that they fall short of fixing these core issues. In this paper, we propose instead to build upon a recently released foundation model for 3D vision that can robustly produce local 3D reconstructions and accurate matches. We introduce a low-memory approach to accurately align these local reconstructions in a global coordinate system. We further show that such foundation models can serve as efficient image retrievers without any overhead, reducing the overall complexity from quadratic to linear. Overall, our novel SfM pipeline is simple, scalable, fast and truly unconstrained, i.e. it can handle any collection of images, ordered or not. Extensive experiments on multiple benchmarks show that our method provides steady performance across diverse settings, especially outperforming existing methods in small- and medium-scale settings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ActMVS: Active Scene Reconstruction with Monocular Multi-View Stereo

    cs.RO 2026-05 unverdicted novelty 7.0 of 10

    ActMVS is the first monocular framework for active scene reconstruction that combines view factor graph construction with global depth optimization to generate online, globally consistent dense depth maps competitive ...

  2. WaterSplat-SLAM: Photorealistic Monocular SLAM in Underwater Environment

    cs.RO 2026-04 unverdicted novelty 7.0 of 10

    WaterSplat-SLAM achieves robust camera tracking and high-fidelity rendering in underwater environments by coupling semantic medium filtering into two-view reconstruction and using an online medium-aware Gaussian map.

  3. A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features

    cs.CV 2025-10 unverdicted novelty 7.0 of 10

    FastForward represents scenes as collections of 3D-anchored image features and performs camera pose estimation via feed-forward correspondence prediction, achieving competitive accuracy with minimal mapping time.

  4. VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold

    cs.CV 2025-05 unverdicted novelty 7.0 of 10

    VGGT-SLAM aligns VGGT submaps via SL(4) manifold optimization of 15-DoF homographies to enable consistent dense RGB SLAM on long uncalibrated monocular videos.

  5. Glob3R: Global Structure-from-Motion with 3D Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A frozen Pi3X backbone plus dense warping tracks and keyframe sliding-window global optimization yields more accurate, scalable SfM than feed-forward or classical baselines alone.

  6. Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Framework for unpaired RGB-thermal novel view synthesis via VGGT-based independent pose estimation, Procrustes alignment with cross-modal matcher, multi-modal 3D Gaussian Splatting, and a new benchmarking framework fo...

  7. D\'ej\`a View: Looping Transformers for Multi-View 3D Reconstruction

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    DéjàView applies a single transformer block recurrently for K refinement steps, matching or exceeding larger feed-forward models on five multi-view 3D benchmarks with fewer parameters and comparable compute.

  8. Self-Improving 4D Perception via Self-Distillation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    SelfEvo enables pretrained 4D perception models to self-improve on unlabeled videos via self-distillation, delivering up to 36.5% relative gains in video depth estimation and 20.1% in camera estimation across eight be...

  9. Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A patternless joint optimization of camera-to-robot calibration and metric-scaled 3D reconstruction, built on MASt3R pointmaps and per-camera scale factors.

  10. Efficient 3D Content Reconstruction and Generation

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Presents Instant3D for rapid text/image-to-3D generation via multi-view diffusion plus feed-forward reconstruction, and FastMap for 10x faster structure-from-motion with comparable accuracy.

  11. Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers

    cs.CV 2025-09 unverdicted novelty 5.0 of 10

    Block-sparse global attention accelerates multi-view reconstruction transformers by over 3x by exploiting concentrated attention on cross-view correspondences.

  12. ViPE: Video Pose Engine for 3D Geometric Perception

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.

  13. UGOD: Uncertainty-Guided Differentiable Opacity and Soft Dropout for Enhanced Sparse-View 3DGS

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A sparse-view 3D Gaussian Splatting method that learns per-Gaussian uncertainty to modulate opacity and apply differentiable soft dropout, reporting small PSNR gains over three baselines.

  14. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

Pith tools