Pith. sign in

REVIEW 1 cited by

RoMo: Robust Motion Segmentation Improves Structure from Motion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.18650 v1 pith:ZUNDKLCM submitted 2024-11-27 cs.CV

classification cs.CV
keywords segmentationmotionvideobaselinescameradynamicposesproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There has been extensive progress in the reconstruction and generation of 4D scenes from monocular casually-captured video. While these tasks rely heavily on known camera poses, the problem of finding such poses using structure-from-motion (SfM) often depends on robustly separating static from dynamic parts of a video. The lack of a robust solution to this problem limits the performance of SfM camera-calibration pipelines. We propose a novel approach to video-based motion segmentation to identify the components of a scene that are moving w.r.t. a fixed world frame. Our simple but effective iterative method, RoMo, combines optical flow and epipolar cues with a pre-trained video segmentation model. It outperforms unsupervised baselines for motion segmentation as well as supervised baselines trained from synthetic data. More importantly, the combination of an off-the-shelf SfM pipeline with our segmentation masks establishes a new state-of-the-art on camera calibration for scenes with dynamic content, outperforming existing methods by a substantial margin.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robust Multimodal Dynamic Object Segmentation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A multimodal trajectory-classification network plus a point-query SAM refinement step yields better dynamic masks and static reconstructions than DAS3R-style baselines on DAVIS.

Pith tools