ActMVS is the first monocular framework for active scene reconstruction that combines view factor graph construction with global depth optimization to generate online, globally consistent dense depth maps competitive with RGB-D methods on Replica datasets.
Mast3r- sfm: a fully-integrated solution for unconstrained structure- from-motion
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 10roles
background 1polarities
background 1representative citing papers
WaterSplat-SLAM achieves robust camera tracking and high-fidelity rendering in underwater environments by coupling semantic medium filtering into two-view reconstruction and using an online medium-aware Gaussian map.
FastForward represents scenes as collections of 3D-anchored image features and performs camera pose estimation via feed-forward correspondence prediction, achieving competitive accuracy with minimal mapping time.
VGGT-SLAM aligns VGGT submaps via SL(4) manifold optimization of 15-DoF homographies to enable consistent dense RGB SLAM on long uncalibrated monocular videos.
Framework for unpaired RGB-thermal novel view synthesis via VGGT-based independent pose estimation, Procrustes alignment with cross-modal matcher, multi-modal 3D Gaussian Splatting, and a new benchmarking framework for cross-modal consistency.
DéjàView applies a single transformer block recurrently for K refinement steps, matching or exceeding larger feed-forward models on five multi-view 3D benchmarks with fewer parameters and comparable compute.
SelfEvo enables pretrained 4D perception models to self-improve on unlabeled videos via self-distillation, delivering up to 36.5% relative gains in video depth estimation and 20.1% in camera estimation across eight benchmarks.
Presents Instant3D for rapid text/image-to-3D generation via multi-view diffusion plus feed-forward reconstruction, and FastMap for 10x faster structure-from-motion with comparable accuracy.
Block-sparse global attention accelerates multi-view reconstruction transformers by over 3x by exploiting concentrated attention on cross-view correspondences.
ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.
citing papers explorer
-
ActMVS: Active Scene Reconstruction with Monocular Multi-View Stereo
ActMVS is the first monocular framework for active scene reconstruction that combines view factor graph construction with global depth optimization to generate online, globally consistent dense depth maps competitive with RGB-D methods on Replica datasets.
-
WaterSplat-SLAM: Photorealistic Monocular SLAM in Underwater Environment
WaterSplat-SLAM achieves robust camera tracking and high-fidelity rendering in underwater environments by coupling semantic medium filtering into two-view reconstruction and using an online medium-aware Gaussian map.
-
A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
FastForward represents scenes as collections of 3D-anchored image features and performs camera pose estimation via feed-forward correspondence prediction, achieving competitive accuracy with minimal mapping time.
-
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
VGGT-SLAM aligns VGGT submaps via SL(4) manifold optimization of 15-DoF homographies to enable consistent dense RGB SLAM on long uncalibrated monocular videos.
-
Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers
Framework for unpaired RGB-thermal novel view synthesis via VGGT-based independent pose estimation, Procrustes alignment with cross-modal matcher, multi-modal 3D Gaussian Splatting, and a new benchmarking framework for cross-modal consistency.
-
D\'ej\`a View: Looping Transformers for Multi-View 3D Reconstruction
DéjàView applies a single transformer block recurrently for K refinement steps, matching or exceeding larger feed-forward models on five multi-view 3D benchmarks with fewer parameters and comparable compute.
-
Self-Improving 4D Perception via Self-Distillation
SelfEvo enables pretrained 4D perception models to self-improve on unlabeled videos via self-distillation, delivering up to 36.5% relative gains in video depth estimation and 20.1% in camera estimation across eight benchmarks.
-
Efficient 3D Content Reconstruction and Generation
Presents Instant3D for rapid text/image-to-3D generation via multi-view diffusion plus feed-forward reconstruction, and FastMap for 10x faster structure-from-motion with comparable accuracy.
-
Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
Block-sparse global attention accelerates multi-view reconstruction transformers by over 3x by exploiting concentrated attention on cross-view correspondences.
-
ViPE: Video Pose Engine for 3D Geometric Perception
ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.