Pith. sign in

REVIEW 18 cited by

Masked-attention Mask Transformer for Universal Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.01527 v3 pith:2ES3ABWK submitted 2021-12-02 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords segmentationtaskimageinstancemasksemanticsarchitecturescoco
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Image segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task. We present Masked-attention Mask Transformer (Mask2Former), a new architecture capable of addressing any image segmentation task (panoptic, instance or semantic). Its key components include masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best specialized architectures by a significant margin on four popular datasets. Most notably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO) and semantic segmentation (57.7 mIoU on ADE20K).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 130 citations worldwide. Full citation record

  1. FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics

    cs.LG 2025-08 conditional novelty 7.0 of 10

    A 188M-parameter Mamba model pretrained on 11M+ simulated sPHENIX events with a new serialization and neighbor-prediction task beats task-specific baselines on three downstream detector tasks when frozen and paired wi...

  2. Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring

    cs.RO 2026-07 conditional novelty 6.0 of 10

    An object-centric, training-free pipeline using CLIP-derived room-probability vectors to score frontiers improves zero-shot ObjectNav success by a relative 3% over an image-based baseline on HM3D.

  3. Leak-Free Cross-Validated Stacking with Per-Architecture Calibration for Sand-Boil Segmentation in Earthen Levees

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A leak-free stacking protocol and mask-conditioned synthetic generation improve sand-boil segmentation to 0.707 IoU, but stacking underperforms the best single model and synthetic gains come only from label-fidelity f...

  4. Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A self-supervised multimodal encoder trained with vision, proprioception, and force yields a vision-only latent that recovers end-effector state and force above vision baselines on RH20T, with modest absolute force accuracy.

  5. Ego-Human Motion Prediction with 3D-Aware LLM

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Ego3DLM jointly predicts past and future 3D body pose and motion descriptions in a single autoregressive pass, conditioned on egocentric video, 3D scene features, and three-point tracking, achieving state-of-the-art o...

  6. Hilti-Trimble-Oxford Dataset: 360 Visual-Inertial Benchmark with Floor Plan Priors for SLAM and Localization

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A 30-sequence 360-degree visual-inertial dataset from an active construction site with LiDAR ground truth and floor plans, benchmarked via an open challenge with 84 teams.

  7. GLOW: A Unified Particle Flow Transformer

    hep-ex 2025-08 conditional novelty 6.0 of 10

    GLOW combines masked-attention transformer decoding with energy-fraction incidence supervision, improving simulated CLIC jet energy resolution by about 15% over HGPflow.

  8. 3D Can Be Explored In 2D: Pseudo-Label Generation for LiDAR Point Clouds Using Sensor-Intensity-Based 2D Semantic Segmentation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A pipeline segments aligned LiDAR point clouds by rendering intensity-colored 2D views, applying a camera-domain 2D segmentation model, and voting back-projected labels, producing competitive pseudo-labels for unsuper...

  9. ZenSVI: An Open-Source Software for the Integrated Acquisition, Processing and Analysis of Street View Imagery Towards Scalable Urban Science

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ZenSVI provides an integrated, documented Python pipeline for acquiring, cleaning, analyzing, and visualizing street view imagery for urban science.

  10. Examining the Associations between Visual and Non-Visual Elements and Cyclists' Route Choices for Various Trip Purposes

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Montreal cyclists' deviations from shortest routes vary by trip purpose and are associated with greenery, motorization, and active-mobility presence in street-view imagery, but the regression analyses predict trip pur...

  11. TravelAgent: Generative Agents in the Built Environment

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A new LLM-driven agent platform navigates virtual urban environments with multimodal sensory inputs, achieving a 76% self-reported task completion rate across 100 simulations.

  12. Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance

    cs.LG 2025-08 conditional novelty 4.0 of 10

    Blending a base diffusion model with its RL-finetuned version at sampling time lets users dial alignment strength, with the blend weight corresponding to the KL-regularization coefficient beta/w.

  13. Leveraging Pathology Foundation Models for Panoptic Segmentation of Melanoma in H&E Images

    eess.IV 2025-07 conditional novelty 4.0 of 10

    A network combining a pathology foundation model (Virchow2) with an Efficient-UNet segments melanoma tissue types and won the PUMA challenge tissue segmentation task.

  14. Fighting Fires from Space: Leveraging Vision Transformers for Enhanced Wildfire Detection and Characterization

    cs.CV 2025-04 conditional novelty 4.0 of 10

    On the Pereira Landsat-8 wildfire dataset, a Swin-Unet achieves 89.93% IoU (0.93% above the best published CNN baseline), while a custom CNN U-Net achieves 93.58% IoU.

  15. Segmentation of arbitrary features in very high resolution remote sensing imagery

    cs.CV 2024-12 reject novelty 4.0 of 10

    A new automated remote sensing segmentation pipeline, EcoMapper, plus an empirical index relating achievable segmentation quality to feature size and image resolution.

  16. Is Semantic SLAM Ready for Embedded Systems ? A Comparative Survey

    cs.RO 2025-05 conditional novelty 3.0 of 10

    A comparative survey and embedded benchmark concluding that semantic geometric SLAM is more practical for real-time deployment than NeRF- or Gaussian-splatting-based semantic SLAM.

  17. A Survey on Training-free Open-Vocabulary Semantic Segmentation

    cs.CV 2025-05 conditional novelty 2.0 of 10

    A structured review of over 30 training-free open-vocabulary semantic segmentation methods, organized by whether they rely on CLIP alone, auxiliary visual foundation models, or generative models.

  18. Image Segmentation with transformers: An Overview, Challenges and Future

    cs.CV 2025-01 reject

    A high-level review of transformer-based image segmentation that restates known models and challenges without new results, and contains citation and metric errors.

Pith tools