REVIEW 14 cited by
Mask2Former for Video Instance Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We find Mask2Former also achieves state-of-the-art performance on video instance segmentation without modifying the architecture, the loss or even the training pipeline. In this report, we show universal image segmentation architectures trivially generalize to video segmentation by directly predicting 3D segmentation volumes. Specifically, Mask2Former sets a new state-of-the-art of 60.4 AP on YouTubeVIS-2019 and 52.6 AP on YouTubeVIS-2021. We believe Mask2Former is also capable of handling video semantic and panoptic segmentation, given its versatility in image segmentation. We hope this will make state-of-the-art video segmentation research more accessible and bring more attention to designing universal image and video segmentation architectures.
Forward citations
Cited by 14 Pith papers
-
CObL: Toward Zero-Shot Ordinal Layering without User Prompting
CObL uses multiple linked Stable Diffusion models to decompose an image into occlusion-ordered object layers, guided at inference so the layers reproduce the input.
-
GVCCS: A Dataset for Contrail Identification and Tracking on Visible Whole Sky Camera Sequences
GVCCS is the first open dataset of ground-based visible all-sky camera video with instance-level contrail masks, temporal tracking, and flight IDs, plus Mask2Former baselines.
-
Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation
A new audio-visual instance segmentation architecture, using audio separation and an audio-modulated Mamba, reaches 48.54 mAP on AVISeg with a COCO-pretrained ResNet50.
-
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
A modified video segmentation architecture decouples processing latency from target count, enabling real-time (>36 FPS) tracking of 10+ objects simultaneously while preserving individual identities.
-
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
A unified model jointly segments, tracks, and captions all objects in videos, trained with VLM-generated synthetic captions, achieving SOTA on VidSTG, VLN, and BenSMOT.
-
AutoQ-VIS: Improving Unsupervised Video Instance Segmentation via Automatic Quality Assessment
A self-training loop with a mask-quality filter, learned on synthetic data, raises unsupervised video instance segmentation on YouTubeVIS-2019 to 52.6 AP50, above VideoCutLER's 48.2.
-
Temporal Cluster Assignment for Efficient Real-Time Video Segmentation
A temporal clustering strategy that reuses cluster assignments across neighboring frames improves the speed-accuracy trade-off of token-clustering video segmentation models.
-
Latest Object Memory Management for Temporally Consistent Video Instance Segmentation
LOMM achieves 54.0 AP on YouTube-VIS 2022 (offline) and 48.2 AP online, via foreground-probability-weighted memory and occupancy-guided decoupled association.
-
GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting
A camera-only Gaussian-surfel pipeline reconstructs full Waymo scenes, converts them to binary occupancy labels, and trains CVT-Occ to generalize on Occ3D-Waymo and Occ3D-nuScenes at a level close to or above LiDAR-la...
-
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
A remote sensing vision-language model that uses task-aware resolution adjustment and attention-based cropping to perform pixel-level segmentation alongside image- and region-level tasks.
-
Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
OpenBench, a new benchmark with categories semantically far from the COCO training space, shows that fine-tuning CLIP hurts open-vocabulary segmentation, and the proposed OVSNet method achieves state-of-the-art on bot...
-
OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
OneIG-Bench introduces a 2,440-prompt, six-dimension benchmark with automated metrics for text-to-image models, covering alignment, text, reasoning, style, and diversity in English and Chinese.
-
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
Concatenating monocular depth maps as an extra input channel improves video instance segmentation and reaches 56.2 AP, a new state of the art on OVIS.
-
Interleaved Transceiver Design for a Continuous- Transmission MIMO-OFDM ISAC System
The abstract advertises a MIMO-OFDM ISAC transceiver design with a claimed first ADPM convergence proof, but the full text is a different paper on continual video instance segmentation, making the ISAC claims unreviewable.
Discussion (0). Sign in to comment.