Pith. sign in

REVIEW 2 cited by

Guess What Moves: Unsupervised Video and Image Segmentation by Anticipating Motion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.07844 v2 pith:ITL6JKI4 submitted 2022-05-16 cs.CV

classification cs.CV
keywords segmentationimagemotionobjectsunsupervisedapproachflowimages
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Motion, measured via optical flow, provides a powerful cue to discover and learn objects in images and videos. However, compared to using appearance, it has some blind spots, such as the fact that objects become invisible if they do not move. In this work, we propose an approach that combines the strengths of motion-based and appearance-based segmentation. We propose to supervise an image segmentation network with the pretext task of predicting regions that are likely to contain simple motion patterns, and thus likely to correspond to objects. As the model only uses a single image as input, we can apply it in two settings: unsupervised video segmentation, and unsupervised image segmentation. We achieve state-of-the-art results for videos, and demonstrate the viability of our approach on still images containing novel objects. Additionally we experiment with different motion models and optical flow backbones and find the method to be robust to these change. Project page and code available at https://www.robots.ox.ac.uk/~vgg/research/gwm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching

    cs.CV 2025-05 conditional novelty 6.0 of 10

    FlowCut creates pseudo-labels from real videos using DINO features plus optical flow, filters them by IoU matching, and trains a video segmentation model that reaches state-of-the-art on YouTubeVIS and DAVIS.

  2. On Moving Object Segmentation from Monocular Video with Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    M3Former fuses appearance and motion streams in a Mask2Former-style transformer and, when trained on a diverse mix that includes KITTI and DAVIS train splits, reaches state-of-the-art motion segmentation scores on tho...

Pith tools