Pith. sign in

REVIEW 1 cited by

AerialFormer: Multi-resolution Transformer for Aerial Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.06842 v2 pith:FMHTYL2Y submitted 2023-06-12 cs.CV

classification cs.CV
keywords aerialformersegmentationaerialimagemd-cnnspathtransformertransformers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Aerial Image Segmentation is a top-down perspective semantic segmentation and has several challenging characteristics such as strong imbalance in the foreground-background distribution, complex background, intra-class heterogeneity, inter-class homogeneity, and tiny objects. To handle these problems, we inherit the advantages of Transformers and propose AerialFormer, which unifies Transformers at the contracting path with lightweight Multi-Dilated Convolutional Neural Networks (MD-CNNs) at the expanding path. Our AerialFormer is designed as a hierarchical structure, in which Transformer encoder outputs multi-scale features and MD-CNNs decoder aggregates information from the multi-scales. Thus, it takes both local and global contexts into consideration to render powerful representations and high-resolution segmentation. We have benchmarked AerialFormer on three common datasets including iSAID, LoveDA, and Potsdam. Comprehensive experiments and extensive ablation studies show that our proposed AerialFormer outperforms previous state-of-the-art methods with remarkable performance. Our source code will be publicly available upon acceptance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new 40-class remote sensing benchmark and a CLIP-plus-DINO dual-stream model are proposed for segmenting arbitrary semantic classes in satellite and aerial images.

Pith tools