Pith. sign in

REVIEW

What is Point Supervision Worth in Video Instance Segmentation?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01990 v1 pith:A5F2HWCM submitted 2024-04-01 cs.CV

classification cs.CV
keywords objectpointvideoannotationsfullyinstancemethodsproposed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video instance segmentation (VIS) is a challenging vision task that aims to detect, segment, and track objects in videos. Conventional VIS methods rely on densely-annotated object masks which are expensive. We reduce the human annotations to only one point for each object in a video frame during training, and obtain high-quality mask predictions close to fully supervised models. Our proposed training method consists of a class-agnostic proposal generation module to provide rich negative samples and a spatio-temporal point-based matcher to match the object queries with the provided point annotations. Comprehensive experiments on three VIS benchmarks demonstrate competitive performance of the proposed framework, nearly matching fully supervised methods.

Discussion (0). Sign in to comment.

Pith tools