Pith. sign in

REVIEW 3 cited by

RoMo: Robust Motion Segmentation Improves Structure from Motion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.18650 v1 pith:ZUNDKLCM submitted 2024-11-27 cs.CV

classification cs.CV
keywords segmentationmotionvideobaselinescameradynamicposesproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion (SfM) often depends on robustly separating static from dynamic parts of a video. The lack of a robust solution to this problem limits the performance of SfM camera-calibration pipelines. We propose a novel approach to video-based motion segmentation to identify the components of a scene that are moving w.r.t. a fixed world frame. Our simple but effective iterative method, RoMo, combines optical flow and epipolar cues with a pre-trained video segmentation model. It outperforms unsupervised baselines for motion segmentation as well as supervised baselines trained from synthetic data. More importantly, the combination of an off-the-shelf SfM pipeline with our segmentation masks establishes a new state-of-the-art on camera calibration for scenes with dynamic content, outperforming existing methods by a substantial margin.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    RiGS decomposes scenes into static, rigid, and transient 4D Gaussians with an object-wise dynamic mask and scene flow guidance to model multi-scale motions and achieve SOTA novel view synthesis.

  2. Robust Multimodal Dynamic Object Segmentation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A multimodal trajectory-classification network plus a point-query SAM refinement step yields better dynamic masks and static reconstructions than DAS3R-style baselines on DAVIS.

  3. ViPE: Video Pose Engine for 3D Geometric Perception

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.

Pith tools