Pith. sign in

REVIEW 3 cited by

Hierarchical Multi-Scale Attention for Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.10821 v1 pith:LM436D4O submitted 2020-05-21 cs.CV

classification cs.CV
keywords cityscapesmulti-scalepredictionsresultsscalesapproachattentionbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-scale inference is commonly used to improve the results of semantic segmentation. Multiple images scales are passed through a network and then the results are combined with averaging or max pooling. In this work, we present an attention-based approach to combining multi-scale predictions. We show that predictions at certain scales are better at resolving particular failures modes, and that the network learns to favor those scales for such cases in order to generate better predictions. Our attention mechanism is hierarchical, which enables it to be roughly 4x more memory efficient to train than other recent approaches. In addition to enabling faster training, this allows us to train with larger crop sizes which leads to greater model accuracy. We demonstrate the result of our method on two datasets: Cityscapes and Mapillary Vistas. For Cityscapes, which has a large number of weakly labelled images, we also leverage auto-labelling to improve generalization. Using our approach we achieve a new state-of-the-art results in both Mapillary (61.1 IOU val) and Cityscapes (85.1 IOU test).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Separate memory banks for shallow pixel details and deep semantic cues, merged by asymmetric cross-attention modules, improve unsupervised video object segmentation on DAVIS-16, FBMS, and YouTube-Objects.

  2. Solving Scene Understanding for Autonomous Navigation in Unstructured Environments

    cs.CV 2025-07 conditional novelty 2.0 of 10

    On the Indian Driving Dataset, a ResNet50-backed U-Net reaches 0.6496 test MIoU, outperforming plain U-Net, DeepLabV3, PSPNet, and SegNet.

  3. Graph Collaborative Attention Network for Link Prediction in Knowledge Graphs

    cs.LG 2025-07 reject novelty 2.0 of 10

    GCAT is presented as a new graph attention model for knowledge graph link prediction, but its equations are those of KBGAT and its reported benchmark numbers do not support the stated performance claims.

Pith tools