Pith. sign in

REVIEW 1 cited by

Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.02358 v2 pith:3XLAU2UO submitted 2021-05-05 cs.CV

classification cs.CV
keywords attentionexternalself-attentionimagelayerslinearmechanismclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Attention mechanisms, especially self-attention, have played an increasingly important role in deep feature representation for visual tasks. Self-attention updates the feature at each position by computing a weighted sum of features using pair-wise affinities across all positions to capture the long-range dependency within a single sample. However, self-attention has quadratic complexity and ignores potential correlation between different samples. This paper proposes a novel attention mechanism which we call external attention, based on two external, small, learnable, shared memories, which can be implemented easily by simply using two cascaded linear layers and two normalization layers; it conveniently replaces self-attention in existing popular architectures. External attention has linear complexity and implicitly considers the correlations between all data samples. We further incorporate the multi-head mechanism into external attention to provide an all-MLP architecture, external attention MLP (EAMLP), for image classification. Extensive experiments on image classification, object detection, semantic segmentation, instance segmentation, image generation, and point cloud analysis reveal that our method provides results comparable or superior to the self-attention mechanism and some of its variants, with much lower computational and memory costs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ensemble-Based Survival Models with the Self-Attended Beran Estimator Predictions

    cs.LG 2025-06 reject novelty 6.0 of 10

    SurvBESA applies self-attention to predicted survival functions from bagged Beran estimators and reports improved ranking performance on benchmark survival datasets.

Pith tools