Pith. sign in

REVIEW 3 cited by

FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1701.05384 v2 pith:TVK4TVPF submitted 2017-01-19 cs.CV

classification cs.CV
keywords objectsvideosappearancegenericmotioncombinedatasetsframework
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in videos. We formulate this task as a structured prediction problem and design a two-stream fully convolutional neural network which fuses together motion and appearance in a unified framework. Since large-scale video datasets with pixel level segmentations are problematic, we show how to bootstrap weakly annotated videos together with existing image recognition datasets for training. Through experiments on three challenging video segmentation benchmarks, our method substantially improves the state-of-the-art for segmenting generic (unseen) objects. Code and pre-trained models are available on the project website.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning segmentation from point trajectories

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A self-supervised low-rank trajectory loss, combined with optical flow, gives state-of-the-art unsupervised video object segmentation on three benchmarks.

  2. FisheyeMODNet: Moving Object detection on Surround-view Cameras for Autonomous Driving

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A lightweight two-stream CNN trained on a new fisheye surround-view dataset detects moving vehicles and pedestrians, reaching about 40% moving-object IoU versus 10% when trained on rectilinear KITTI data.

  3. Exploiting Temporality for Semi-Supervised Video Segmentation

    cs.CV 2019-08 conditional novelty 5.0 of 10

    Placing temporal modules inside the encoder of a U-Net, and propagating their outputs to the next convolutional blocks, improves semi-supervised video segmentation on CityScapes by about 6 mIoU points over a frame-by-...

Pith tools