FlowCut creates pseudo-labels from real videos using DINO features plus optical flow, filters them by IoU matching, and trains a video segmentation model that reaches state-of-the-art on YouTubeVIS and DAVIS.
Unsupervised Semantic Segmentation with Self-supervised Object-centric Representations
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this paper, we show that recent advances in self-supervised feature learning enable unsupervised object discovery and semantic segmentation with a performance that matches the state of the field on supervised semantic segmentation 10 years ago. We propose a methodology based on unsupervised saliency masks and self-supervised feature clustering to kickstart object discovery followed by training a semantic segmentation network on pseudo-labels to bootstrap the system on images with multiple objects. We present results on PASCAL VOC that go far beyond the current state of the art (50.0 mIoU), and we report for the first time results on MS COCO for the whole set of 81 classes: our method discovers 34 categories with more than $20\%$ IoU, while obtaining an average IoU of 19.6 for all 81 categories.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching
FlowCut creates pseudo-labels from real videos using DINO features plus optical flow, filters them by IoU matching, and trains a video segmentation model that reaches state-of-the-art on YouTubeVIS and DAVIS.