Pith. sign in

REVIEW 2 cited by

AISFormer: Amodal Instance Segmentation with Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.06323 v4 pith:DZI5FNHH submitted 2022-10-12 cs.CV

classification cs.CV
keywords aisformeramodalmaskvisiblecoherenceinstanceinvisiblemasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Amodal Instance Segmentation (AIS) aims to segment the region of both visible and possible occluded parts of an object instance. While Mask R-CNN-based AIS approaches have shown promising results, they are unable to model high-level features coherence due to the limited receptive field. The most recent transformer-based models show impressive performance on vision tasks, even better than Convolution Neural Networks (CNN). In this work, we present AISFormer, an AIS framework, with a Transformer-based mask head. AISFormer explicitly models the complex coherence between occluder, visible, amodal, and invisible masks within an object's regions of interest by treating them as learnable queries. Specifically, AISFormer contains four modules: (i) feature encoding: extract ROI and learn both short-range and long-range visual features. (ii) mask transformer decoding: generate the occluder, visible, and amodal mask query embeddings by a transformer decoder (iii) invisible mask embedding: model the coherence between the amodal and visible masks, and (iv) mask predicting: estimate output masks including occluder, visible, amodal and invisible. We conduct extensive experiments and ablation studies on three challenging benchmarks i.e. KINS, D2SA, and COCOA-cls to evaluate the effectiveness of AISFormer. The code is available at: https://github.com/UARK-AICV/AISFormer

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A multi-grained network that recognizes seen actions with a supervised classifier and unseen actions via video-level coarse filtering plus proposal-level matching achieves state-of-the-art open-vocabulary temporal act...

  2. A2VIS: Amodal-Aware Approach to Video Instance Segmentation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A2VIS integrates amodal, full-shape masks into video instance segmentation via global prototypes and a spatiotemporal-prior mask head, improving occlusion-robust tracking on synthetic benchmarks.

Pith tools