Pith. sign in

REVIEW 1 cited by

ClusTR: Exploring Efficient Self-attention via Clustering for Vision Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.13138 v1 pith:D7ONUHWV submitted 2022-08-28 cs.CV

classification cs.CV
keywords clustrcomputationaldenseachievesattentioncomplexitycontent-basedcost
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although Transformers have successfully transitioned from their language modelling origins to image-based applications, their quadratic computational complexity remains a challenge, particularly for dense prediction. In this paper we propose a content-based sparse attention method, as an alternative to dense self-attention, aiming to reduce the computation complexity while retaining the ability to model long-range dependencies. Specifically, we cluster and then aggregate key and value tokens, as a content-based method of reducing the total token count. The resulting clustered-token sequence retains the semantic diversity of the original signal, but can be processed at a lower computational cost. Besides, we further extend the clustering-guided attention from single-scale to multi-scale, which is conducive to dense prediction tasks. We label the proposed Transformer architecture ClusTR, and demonstrate that it achieves state-of-the-art performance on various vision tasks but at lower computational cost and with fewer parameters. For instance, our ClusTR small model with 22.7M parameters achieves 83.2\% Top-1 accuracy on ImageNet. Source code and ImageNet models will be made publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting the Integration of Convolution and Attention for Vision Backbone

    cs.CV 2024-11 conditional novelty 6.0 of 10

    By splitting local and global feature processing across different granularities, GLNet matches state-of-the-art vision backbones with fewer FLOPs and higher throughput.

Pith tools