Pith. sign in

REVIEW 28 cited by

Grounding Image Matching in 3D with MASt3R

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09756 v1 pith:D3VTHCBY submitted 2024-06-14 cs.CV

Grounding Image Matching in 3D with MASt3R

classification cs.CV
keywords matchingapproachdensedust3rimagemast3rproblempropose
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Image Matching is a core component of all best-performing algorithms and pipelines in 3D vision. Yet despite matching being fundamentally a 3D problem, intrinsically linked to camera pose and scene geometry, it is typically treated as a 2D problem. This makes sense as the goal of matching is to establish correspondences between 2D pixel fields, but also seems like a potentially hazardous choice. In this work, we take a different stance and propose to cast matching as a 3D task with DUSt3R, a recent and powerful 3D reconstruction framework based on Transformers. Based on pointmaps regression, this method displayed impressive robustness in matching views with extreme viewpoint changes, yet with limited accuracy. We aim here to improve the matching capabilities of such an approach while preserving its robustness. We thus propose to augment the DUSt3R network with a new head that outputs dense local features, trained with an additional matching loss. We further address the issue of quadratic complexity of dense matching, which becomes prohibitively slow for downstream applications if not carefully treated. We introduce a fast reciprocal matching scheme that not only accelerates matching by orders of magnitude, but also comes with theoretical guarantees and, lastly, yields improved results. Extensive experiments show that our approach, coined MASt3R, significantly outperforms the state of the art on multiple matching tasks. In particular, it beats the best published methods by 30% (absolute improvement) in VCRE AUC on the extremely challenging Map-free localization dataset.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Geo-Align: Video Generation Alignment via Metric Geometry Reward

    cs.CV 2026-05 unverdicted novelty 7.0

    Geo-Align applies RL with a perceptual reward derived from 3D camera trajectory estimation to improve controllability and fidelity in video generation without paired training data.

  2. GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 7.0

    GenRecon lifts object-level generative priors to scene-scale reconstruction by chunking scenes and using projection-based conditioning on multi-view features, claiming 16% better results than prior methods.

  3. No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos

    cs.CV 2026-05 unverdicted novelty 7.0

    NoPo4D is the first feed-forward system for dynamic 4D Gaussian splatting from unposed multi-view videos, using velocity decomposition supervised by optical flow and a bidirectional motion encoder.

  4. EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

    cs.RO 2026-05 conditional novelty 7.0

    EvoScene-VLA maintains an action-updated scene prior across control chunks in VLA policies, raising success rates on RoboTwin tasks from 87.2% to 89.1% fixed and 86.1% to 88.5% randomized while outperforming baselines...

  5. CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

    cs.CV 2026-05 unverdicted novelty 7.0

    CRePE supplies depth-aware positional distributions along curved rays for stable unified-camera control in frozen video DiT models.

  6. WildSplatter: Feed-forward 3D Gaussian Splatting with Appearance Control from Unconstrained Images

    cs.CV 2026-04 unverdicted novelty 7.0

    WildSplatter jointly learns 3D Gaussians and appearance embeddings from unconstrained photo collections to enable fast feed-forward reconstruction and flexible lighting control in 3D Gaussian Splatting.

  7. LuMon: A Comprehensive Benchmark and Development Suite with Novel Datasets for Lunar Monocular Depth Estimation

    cs.CV 2026-04 unverdicted novelty 7.0

    A new benchmark with real lunar stereo ground truth and analog data shows that sim-to-real fine-tuned monocular depth models achieve large in-domain gains but minimal generalization to actual lunar images.

  8. Affostruction: 3D Affordance Grounding with Generative Reconstruction

    cs.CV 2026-01 unverdicted novelty 7.0

    Affostruction reconstructs full 3D object geometry from partial RGBD views and grounds text-based affordances on both visible and unobserved surfaces, reporting large gains over prior methods.

  9. A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features

    cs.CV 2025-10 unverdicted novelty 7.0

    FastForward represents scenes as collections of 3D-anchored image features and performs camera pose estimation via feed-forward correspondence prediction, achieving competitive accuracy with minimal mapping time.

  10. MACRO: Training-free Multi-plane Attention for Closeup Render Optimization

    cs.CV 2026-07 conditional novelty 6.0

    Training-free multi-plane attention with image-space scale-matched reference crops restores correct close-up detail from 3DGS without retraining the enhancer.

  11. Modality Forcing for Scalable Spatial Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    Modality Forcing lets a single DiT produce image and depth outputs in any order after training on sparse real-world depth, with larger image-pretrained models yielding better depth accuracy and a 57% AbsRel reduction ...

  12. Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    Robust Dreamer uses Latent Gaussian Memory anchored to diffusion latents and Deviation Learning with a Dynamic Deviation Archive to reduce drift in long-horizon action-controlled image-to-video generation, reporting S...

  13. SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data

    cs.CV 2026-05 unverdicted novelty 6.0

    SADGE is a new fused similarity metric combining DINOv3 appearance and MASt3R geometry via constrained bilinear interaction that correlates with downstream synthetic-to-real performance at Pearson r=0.88 across multip...

  14. SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs

    cs.CV 2026-05 unverdicted novelty 6.0

    SpaceMind++ adds an explicit voxelized allocentric cognitive map and coordinate-guided fusion to video MLLMs, claiming SOTA on VSI-Bench and improved out-of-distribution generalization on three other 3D benchmarks.

  15. FluSplat: Sparse-View 3D Editing without Test-Time Optimization

    cs.CV 2026-04 unverdicted novelty 6.0

    FluSplat trains a model with geometric alignment constraints on multi-view edits to produce consistent 3D scene edits from sparse views in a single forward pass without test-time optimization.

  16. VGGT-HPE: Reframing Head Pose Estimation as Relative Pose Prediction

    cs.CV 2026-04 conditional novelty 6.0

    Reframing head pose estimation as relative pose prediction between image pairs enables a synthetic-only trained model to outperform absolute regression methods on real benchmarks.

  17. Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models

    cs.CV 2025-11 unverdicted novelty 6.0

    A feed-forward video latent transformer that predicts time-varying 3D Gaussian primitives from one image to produce controllable 4D scenes with appearance, geometry, and motion.

  18. Streaming 4D Visual Geometry Transformer

    cs.CV 2025-07 unverdicted novelty 6.0

    A causal transformer with key-value caching and distillation from a bidirectional VGGT model enables efficient online 4D geometry reconstruction from videos.

  19. RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos

    cs.CV 2024-12 unverdicted novelty 6.0

    RoDyGS separates static and dynamic elements in monocular videos using Gaussian splatting with regularization and introduces the Kubric-MRig benchmark for pose-free dynamic novel view synthesis.

  20. Bundle Adjustment in the Eager Mode

    cs.RO 2024-09 unverdicted novelty 6.0

    Introduces an eager-mode PyTorch BA library with GPU-accelerated sparse ops claiming 18.5-23x speedups over GTSAM, g2o, and Ceres.

  21. ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis

    cs.CV 2024-09 unverdicted novelty 6.0

    ViewCrafter tames video diffusion models with point-based 3D guidance and iterative trajectory planning to produce high-fidelity novel views from single or sparse images.

  22. VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

    cs.CV 2026-03 conditional novelty 5.5

    A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.

  23. VGP-Nav: Metric-Aware Visual Geometric Perception for Robot Navigation

    cs.RO 2026-06 unverdicted novelty 5.0

    VGP-Nav resolves monocular scale ambiguity online via ground-plane geometry to produce metric localization and obstacle maps from RGB alone for robot navigation.

  24. SAVMap: Structure-Aided Visual Mapping of Large-Scale 2.5D Manhattan Wireframes from Panoramic Video

    cs.CV 2026-06 unverdicted novelty 5.0

    SAVMap extracts semantic structure points from panoramic video, tracks them across rectified views, and recovers 3D wireframe maps via Manhattan-constrained SfM, reporting 4.8 cm aggregate MAE over 5000 shelf elements...

  25. Feature-Optimized Vision for Adaptive 3D Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 5.0

    An adaptive multi-criteria feature selection policy for 3D scene reconstruction that outperforms random, texture-only, and uniform-grid baselines on synthetic multi-view tests for completeness and RMSE.

  26. PatchPoison: Poisoning Multi-View Datasets to Degrade 3D Reconstruction

    cs.CV 2026-04 unverdicted novelty 5.0

    PatchPoison injects 12x12 pixel checkerboard patches into multi-view images to disrupt SfM feature matching, causing 3DGS reconstructions to diverge with 6.8x higher LPIPS error on NeRF-Synthetic while remaining unobtrusive.

  27. UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

    cs.CV 2025-02 conditional novelty 5.0

    UniDepthV2 predicts metric 3D points directly from single images using a self-promptable camera module, pseudo-spherical representation, and new losses for improved cross-domain generalization.

  28. Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs

    cs.CV 2024-08 unverdicted novelty 5.0

    Splatt3R is a feed-forward network that predicts 3D Gaussian splats directly from uncalibrated stereo image pairs by extending MASt3R with appearance attributes and a two-stage training procedure.