REVIEW 11 cited by
BA-Net: Dense Bundle Adjustment Network
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
BA-Net: Dense Bundle Adjustment Network
read the original abstract
This paper introduces a network architecture to solve the structure-from-motion (SfM) problem via feature-metric bundle adjustment (BA), which explicitly enforces multi-view geometry constraints in the form of feature-metric error. The whole pipeline is differentiable so that the network can learn suitable features that make the BA problem more tractable. Furthermore, this work introduces a novel depth parameterization to recover dense per-pixel depth. The network first generates several basis depth maps according to the input image and optimizes the final depth as a linear combination of these basis depth maps via feature-metric BA. The basis depth maps generator is also learned via end-to-end training. The whole system nicely combines domain knowledge (i.e. hard-coded multi-view geometry constraints) and deep learning (i.e. feature learning and basis depth maps learning) to address the challenging dense SfM problem. Experiments on large scale real data prove the success of the proposed method.
Forward citations
Cited by 11 Pith papers
-
Accelerating Transformer-Based Monocular SLAM via Geometric Utility Scoring
LeanGate is a lightweight feed-forward network that predicts geometric utility scores to skip over 90% of redundant frames in GFM-based monocular SLAM, reducing tracking FLOPs by 85% and achieving 5x speedup while mai...
-
Glob3R: Global Structure-from-Motion with 3D Foundation Models
A frozen Pi3X backbone plus dense warping tracks and keyframe sliding-window global optimization yields more accurate, scalable SfM than feed-forward or classical baselines alone.
-
Learning Adaptive Solvers for Distributed Factor Graph Optimization on Matrix Lie Groups
A learned feedback policy replaces manual parameter tuning in distributed Riemannian optimization over matrix Lie groups, achieving lower objective values on multi-robot mapping benchmarks.
-
Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence
Tango3D unifies dense pixel-to-point 2D-3D alignment and global retrieval in one shared space using a geometry-aware 2D backbone, 3D VAE tokens, and three-stage progressive training.
-
Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
Scal3R achieves better accuracy and consistency in large-scale 3D scene reconstruction by maintaining a compressed global context through test-time adaptation of lightweight neural networks on long video sequences.
-
PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception
A fine-tuned VGGT with a learned dynamics mask improves camera pose, depth, and point-cloud reconstruction on dynamic-scene benchmarks over the original static-scene model.
-
PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception
PAGE-4D is a feedforward extension of VGGT that uses a dynamics-aware aggregator and mask to disentangle pose estimation from geometry reconstruction in videos with moving objects.
-
Efficient 3D Content Reconstruction and Generation
Presents Instant3D for rapid text/image-to-3D generation via multi-view diffusion plus feed-forward reconstruction, and FastMap for 10x faster structure-from-motion with comparable accuracy.
-
TTT3R: 3D Reconstruction as Test-Time Training
TTT3R derives a closed-form learning rate from memory-observation alignment confidence to boost length generalization in RNN-based 3D reconstruction by 2x in global pose estimation.
-
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
By fine-tuning DUST3R to output per-timestep pointmaps on scarce dynamic video datasets, MonST3R achieves stronger video depth and pose estimation without explicit motion modeling.
-
VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
VGGT-Long extends VGGT with chunking, overlap alignment, and loop closure to produce consistent kilometer-scale 3D reconstructions from monocular RGB sequences without retraining or extra supervision.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.