Pith. sign in

REVIEW 8 cited by

FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.15259 v3 pith:RXCN3P3C submitted 2024-04-23 cs.CV

classification cs.CV
keywords methoddepthcameraintrinsicsdifferentiablegradient-descentposesdegree
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces FlowMap, an end-to-end differentiable method that solves for precise camera poses, camera intrinsics, and per-frame dense depth of a video sequence. Our method performs per-video gradient-descent minimization of a simple least-squares objective that compares the optical flow induced by depth, intrinsics, and poses against correspondences obtained via off-the-shelf optical flow and point tracking. Alongside the use of point tracks to encourage long-term geometric consistency, we introduce differentiable re-parameterizations of depth, intrinsics, and pose that are amenable to first-order optimization. We empirically show that camera parameters and dense depth recovered by our method enable photo-realistic novel view synthesis on 360-degree trajectories using Gaussian Splatting. Our method not only far outperforms prior gradient-descent based bundle adjustment methods, but surprisingly performs on par with COLMAP, the state-of-the-art SfM method, on the downstream task of 360-degree novel view synthesis (even though our method is purely gradient-descent based, fully differentiable, and presents a complete departure from conventional SfM).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Light3R-SfM: Towards Feed-forward Structure-from-Motion

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Light3R-SfM replaces global bundle adjustment in Structure-from-Motion with a learned attention module and a sparse tree of image pairs, cutting runtime by up to two orders of magnitude while keeping competitive pose ...

  2. Glob3R: Global Structure-from-Motion with 3D Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A frozen Pi3X backbone plus dense warping tracks and keyframe sliding-window global optimization yields more accurate, scalable SfM than feed-forward or classical baselines alone.

  3. CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    CGGS generates viewpoint-consistent, text-aligned ego-centric 3D scenes via consistency-augmented multi-view diffusion, flow-guided layout initialization, and mutual-information depth-refined Gaussian optimization.

  4. 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    4D-Animal fits SMAL animal models to video using silhouette, part, pixel, and tracking losses from off-the-shelf 2D models, removing the need for sparse keypoint annotations.

  5. E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    E3D-Bench compares 16 3D geometric foundation models on depth, reconstruction, pose, and view-synthesis tasks with a unified evaluation toolkit.

  6. Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single-pass transformer generalizes DUSt3R's pointmap regression from two views to all-to-all multi-view attention, reconstructing 1000+ images and estimating camera poses in one forward pass.

  7. MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 700K-scene procedural, non-semantic synthetic dataset improves large reconstruction models by 1.2 to 1.8 dB PSNR when combined with real data.

  8. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

Pith tools