Pith. sign in

REVIEW 14 cited by

Mask2Former for Video Instance Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.10764 v1 pith:N5DUEMWQ submitted 2021-12-20 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords segmentationvideomask2formerimagestate-of-the-artarchitecturesinstanceuniversal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We find Mask2Former also achieves state-of-the-art performance on video instance segmentation without modifying the architecture, the loss or even the training pipeline. In this report, we show universal image segmentation architectures trivially generalize to video segmentation by directly predicting 3D segmentation volumes. Specifically, Mask2Former sets a new state-of-the-art of 60.4 AP on YouTubeVIS-2019 and 52.6 AP on YouTubeVIS-2021. We believe Mask2Former is also capable of handling video semantic and panoptic segmentation, given its versatility in image segmentation. We hope this will make state-of-the-art video segmentation research more accessible and bring more attention to designing universal image and video segmentation architectures.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CObL: Toward Zero-Shot Ordinal Layering without User Prompting

    cs.CV 2025-08 conditional novelty 8.0 of 10

    CObL uses multiple linked Stable Diffusion models to decompose an image into occlusion-ordered object layers, guided at inference so the layers reproduce the input.

  2. GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences

    cs.CV 2025-07 conditional novelty 7.0 of 10

    GVCCS is the first open dataset of ground-based visible all-sky camera video with instance-level contrail masks, temporal tracking, and flight IDs, plus Mask2Former baselines.

  3. Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation

    cs.MM 2026-08 conditional novelty 6.0 of 10

    A new audio-visual instance segmentation architecture, using audio separation and an audio-modulated Mamba, reaches 48.54 mAP on AVISeg with a COCO-pretrained ResNet50.

  4. SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A modified video segmentation architecture decouples processing latency from target count, enabling real-time (>36 FPS) tracking of 10+ objects simultaneously while preserving individual identities.

  5. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified model jointly segments, tracks, and captions all objects in videos, trained with VLM-generated synthetic captions, achieving SOTA on VidSTG, VLN, and BenSMOT.

  6. AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A self-training loop with a mask-quality filter, learned on synthetic data, raises unsupervised video instance segmentation on YouTubeVIS-2019 to 52.6 AP50, above VideoCutLER's 48.2.

  7. Temporal Cluster Assignment for Efficient Real-Time Video Segmentation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A temporal clustering strategy that reuses cluster assignments across neighboring frames improves the speed-accuracy trade-off of token-clustering video segmentation models.

  8. Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LOMM achieves 54.0 AP on YouTube-VIS 2022 (offline) and 48.2 AP online, via foreground-probability-weighted memory and occupancy-guided decoupled association.

  9. GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A camera-only Gaussian-surfel pipeline reconstructs full Waymo scenes, converts them to binary occupancy labels, and trains CVT-Occ to generalize on Occ3D-Waymo and Occ3D-nuScenes at a level close to or above LiDAR-la...

  10. GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A remote sensing vision-language model that uses task-aware resolution adjustment and attention-based cropping to perform pixel-level segmentation alongside image- and region-level tasks.

  11. Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OpenBench, a new benchmark with categories semantically far from the COCO training space, shows that fine-tuning CLIP hurts open-vocabulary segmentation, and the proposed OVSNet method achieves state-of-the-art on bot...

  12. OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OneIG-Bench introduces a 2,440-prompt, six-dimension benchmark with automated metrics for text-to-image models, covering alignment, text, reasoning, style, and diversity in English and Chinese.

  13. Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Concatenating monocular depth maps as an extra input channel improves video instance segmentation and reaches 56.2 AP, a new state of the art on OVIS.

  14. Interleaved Transceiver Design for a Continuous- Transmission MIMO-OFDM ISAC System

    eess.SP 2025-08 unverdicted novelty 4.0 of 10

    The abstract advertises a MIMO-OFDM ISAC transceiver design with a claimed first ADPM convergence proof, but the full text is a different paper on continual video instance segmentation, making the ISAC claims unreviewable.

Pith tools