Pith. sign in

FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in videos. We formulate this task as a structured prediction problem and design a two-stream fully convolutional neural network which fuses together motion and appearance in a unified framework. Since large-scale video datasets with pixel level segmentations are problematic, we show how to bootstrap weakly annotated videos together with existing image recognition datasets for training. Through experiments on three challenging video segmentation benchmarks, our method substantially improves the state-of-the-art for segmenting generic (unseen) objects. Code and pre-trained models are available on the project website.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Learning segmentation from point trajectories

cs.CV · 2025-01-21 · conditional · novelty 7.0

A self-supervised low-rank trajectory loss, combined with optical flow, gives state-of-the-art unsupervised video object segmentation on three benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Learning segmentation from point trajectories cs.CV · 2025-01-21 · conditional · none · ref 25 · internal anchor

    A self-supervised low-rank trajectory loss, combined with optical flow, gives state-of-the-art unsupervised video object segmentation on three benchmarks.