Pith. sign in

REVIEW 1 cited by

MacFormer: Semantic Segmentation with Fine Object Boundaries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05699 v1 pith:Y2K33TOX submitted 2024-08-11 cs.CV

classification cs.CV
keywords featuressegmentationsemanticboundariesmacformerobjectagentcomponents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semantic segmentation involves assigning a specific category to each pixel in an image. While Vision Transformer-based models have made significant progress, current semantic segmentation methods often struggle with precise predictions in localized areas like object boundaries. To tackle this challenge, we introduce a new semantic segmentation architecture, ``MacFormer'', which features two key components. Firstly, using learnable agent tokens, a Mutual Agent Cross-Attention (MACA) mechanism effectively facilitates the bidirectional integration of features across encoder and decoder layers. This enables better preservation of low-level features, such as elementary edges, during decoding. Secondly, a Frequency Enhancement Module (FEM) in the decoder leverages high-frequency and low-frequency components to boost features in the frequency domain, benefiting object boundaries with minimal computational complexity increase. MacFormer is demonstrated to be compatible with various network architectures and outperforms existing methods in both accuracy and efficiency on benchmark datasets ADE20K and Cityscapes under different computational constraints.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SCASeg: Strip Cross-Attention for Efficient Semantic Segmentation

    cs.CV 2024-11 unverdicted novelty 5.0 of 10

    SCASeg proposes a strip cross-attention decoder with lateral connections and a cross-layer block to efficiently capture global-local context, reporting competitive or superior results on ADE20K, Cityscapes, COCO-Stuff...

Pith tools