Pith. sign in

REVIEW 2 cited by

Global Self-Attention Networks for Image Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.03019 v2 pith:DHPO2T3J submitted 2020-10-06 cs.CV cs.LG

classification cs.CVcs.LG
keywords modulenetworksattentionglobalnetworkself-attentiondeeplayer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, a series of works in computer vision have shown promising results on various image and video understanding tasks using self-attention. However, due to the quadratic computational and memory complexities of self-attention, these works either apply attention only to low-resolution feature maps in later stages of a deep network or restrict the receptive field of attention in each layer to a small local region. To overcome these limitations, this work introduces a new global self-attention module, referred to as the GSA module, which is efficient enough to serve as the backbone component of a deep network. This module consists of two parallel layers: a content attention layer that attends to pixels based only on their content and a positional attention layer that attends to pixels based on their spatial locations. The output of this module is the sum of the outputs of the two layers. Based on the proposed GSA module, we introduce new standalone global attention-based deep networks that use GSA modules instead of convolutions to model pixel interactions. Due to the global extent of the proposed GSA module, a GSA network has the ability to model long-range pixel interactions throughout the network. Our experimental results show that GSA networks outperform the corresponding convolution-based networks significantly on the CIFAR-100 and ImageNet datasets while using less parameters and computations. The proposed GSA networks also outperform various existing attention-based networks on the ImageNet dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ensemble-Based Survival Models with the Self-Attended Beran Estimator Predictions

    cs.LG 2025-06 reject novelty 6.0 of 10

    SurvBESA applies self-attention to predicted survival functions from bagged Beran estimators and reports improved ranking performance on benchmark survival datasets.

  2. MDD-Net: Multimodal Depression Detection through Mutual Transformer

    cs.CV 2025-08 conditional novelty 4.0 of 10

    MDD-Net fuses acoustic and visual features with mutual transformers and reports 0.7707 F1 on the D-Vlog depression dataset.

Pith tools