REVIEW 16 cited by
BA-Net: Dense Bundle Adjustment Network
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper introduces a network architecture to solve the structure-from-motion (SfM) problem via feature-metric bundle adjustment (BA), which explicitly enforces multi-view geometry constraints in the form of feature-metric error. The whole pipeline is differentiable so that the network can learn suitable features that make the BA problem more tractable. Furthermore, this work introduces a novel depth parameterization to recover dense per-pixel depth. The network first generates several basis depth maps according to the input image and optimizes the final depth as a linear combination of these basis depth maps via feature-metric BA. The basis depth maps generator is also learned via end-to-end training. The whole system nicely combines domain knowledge (i.e. hard-coded multi-view geometry constraints) and deep learning (i.e. feature learning and basis depth maps learning) to address the challenging dense SfM problem. Experiments on large scale real data prove the success of the proposed method.
Forward citations
Cited by 16 Pith papers
-
Glob3R: Global Structure-from-Motion with 3D Foundation Models
A frozen Pi3X backbone plus dense warping tracks and keyframe sliding-window global optimization yields more accurate, scalable SfM than feed-forward or classical baselines alone.
-
Learning Adaptive Solvers for Distributed Factor Graph Optimization on Matrix Lie Groups
A learned feedback policy replaces manual parameter tuning in distributed Riemannian optimization over matrix Lie groups, achieving lower objective values on multi-robot mapping benchmarks.
-
PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception
PAGE-4D is a feedforward extension of VGGT that uses a dynamics-aware aggregator and mask to disentangle pose estimation from geometry reconstruction in videos with moving objects.
-
SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views
SPFSplatV2 reconstructs 3D Gaussian scenes and camera poses from sparse unposed images with a single shared transformer backbone, using masked attention and a reprojection loss, and reports state-of-the-art results wi...
-
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
A decoder-only Transformer with causal attention and cached past-frame features performs incremental 3D reconstruction from streaming images, beating the RNN-based CUT3R on several benchmark metrics.
-
ZeroVO: Visual Odometry with Minimal Assumptions
A two-frame visual odometry model using estimated depth, language priors, and semi-supervised pseudo-label filtering achieves zero-shot metric-scale pose estimation across multiple driving datasets.
-
Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion
MVGD jointly generates novel-view images and scale-consistent depth maps with a pixel-level diffusion model, reporting state-of-the-art scores on several view synthesis and depth benchmarks.
-
Continuous 3D Perception Model with Persistent State
A recurrent transformer with a persistent state performs online metric 3D reconstruction from image streams and can infer unseen scene geometry from virtual camera queries.
-
MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing
MambaVO improves deep visual odometry by adding Mamba-based matching refinement and a smoothed training objective, achieving state-of-the-art absolute trajectory error on EuRoC, TUM-RGBD, KITTI, and TartanAir.
-
Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
STVO improves visual odometry by combining temporal motion propagation and depth-based spatial attention to make multi-frame optical flow matching more consistent, setting state-of-the-art ATE on TUM-RGBD, EuRoC, ETH3...
-
Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
A pipeline turns internet VR180 stereo videos into 110k dynamic 3D point cloud clips, and a DUSt3R variant trained on these clips predicts 3D motion and structure from image pairs better than a model trained on synthe...
-
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
A deep visual SLAM pipeline, augmented with monocular depth priors, learned motion probability maps, and uncertainty-aware bundle adjustment, estimates camera poses and consistent depths from casual monocular videos o...
-
Rolling Shutter Camera Self-Calibration
A target-free bundle adjustment method jointly estimates camera intrinsics and the row readout time ratio for rolling shutter cameras by combining continuous-time trajectory estimation with correction-field rectification.
-
3D Plant Root Skeleton Detection and Extraction
A multi-view computer vision pipeline detects and matches lateral roots in images, triangulates them, and refines the result with bundle adjustment to reconstruct 3D root skeletons from a few views, evaluated on a cus...
-
Who is a Better Player: LLM against LLM
The abstract claims an LLM-versus-LLM board game benchmark with new stability and sentiment metrics, but the body text is a mismatched manuscript about SEM 3D reconstruction.
-
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.
Discussion (0). Continue with ORCID to comment.