Pith. sign in

REVIEW 2 cited by

NOVIS: A Case for End-to-End Near-Online Video Instance Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.15266 v2 pith:NOJSJIJR submitted 2023-08-29 cs.CV

classification cs.CV
keywords instancenear-onlinevideomethodsnovissegmentationbeliefclips
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Until recently, the Video Instance Segmentation (VIS) community operated under the common belief that offline methods are generally superior to a frame by frame online processing. However, the recent success of online methods questions this belief, in particular, for challenging and long video sequences. We understand this work as a rebuttal of those recent observations and an appeal to the community to focus on dedicated near-online VIS approaches. To support our argument, we present a detailed analysis on different processing paradigms and the new end-to-end trainable NOVIS (Near-Online Video Instance Segmentation) method. Our transformer-based model directly predicts spatio-temporal mask volumes for clips of frames and performs instance tracking between clips via overlap embeddings. NOVIS represents the first near-online VIS approach which avoids any handcrafted tracking heuristics. We outperform all existing VIS methods by large margins and provide new state-of-the-art results on both YouTube-VIS (2019/2021) and the OVIS benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LOMM achieves 54.0 AP on YouTube-VIS 2022 (offline) and 48.2 AP online, via foreground-probability-weighted memory and occupancy-guided decoupled association.

  2. QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Enriching image-trained Mask2Former queries with appearance and center losses, plus training-free propagation, lifts MinVIS by up to +6.7 AP and nears video-supervised VIS.

Pith tools