Pith. sign in

Structure from Motion for Panorama-Style Videos

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We present a novel Structure from Motion pipeline that is capable of reconstructing accurate camera poses for panorama-style video capture without prior camera intrinsic calibration. While panorama-style capture is common and convenient, previous reconstruction methods fail to obtain accurate reconstructions due to the rotation-dominant motion and small baseline between views. Our method is built on the assumption that the camera motion approximately corresponds to motion on a sphere, and we introduce three novel relative pose methods to estimate the fundamental matrix and camera distortion for spherical motion. These solvers are efficient and robust, and provide an excellent initialization for bundle adjustment. A soft prior on the camera poses is used to discourage large deviations from the spherical motion assumption when performing bundle adjustment, which allows cameras to remain properly constrained for optimization in the absence of well-triangulated 3D points. To validate the effectiveness of the proposed method we evaluate our approach on both synthetic and real-world data, and demonstrate that camera poses are accurate enough for multiview stereo.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2024 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

cs.CV · 2024-12-12 · conditional · novelty 6.0

A pipeline turns internet VR180 stereo videos into 110k dynamic 3D point cloud clips, and a DUSt3R variant trained on these clips predicts 3D motion and structure from image pairs better than a model trained on synthetic data.

citing papers explorer

Showing 1 of 1 citing paper.

  • Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos cs.CV · 2024-12-12 · conditional · none · ref 87 · internal anchor

    A pipeline turns internet VR180 stereo videos into 110k dynamic 3D point cloud clips, and a DUSt3R variant trained on these clips predicts 3D motion and structure from image pairs better than a model trained on synthetic data.