Pith. sign in

REVIEW

MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.15168 v3 pith:D3NX4YGO submitted 2021-05-31 cs.CV cs.LG

classification cs.CVcs.LG
keywords msg-transformertransformersvisualcomputationalinformationmanipulatingmessengernetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers have offered a new methodology of designing neural networks for visual recognition. Compared to convolutional networks, Transformers enjoy the ability of referring to global features at each stage, yet the attention module brings higher computational overhead that obstructs the application of Transformers to process high-resolution visual data. This paper aims to alleviate the conflict between efficiency and flexibility, for which we propose a specialized token for each region that serves as a messenger (MSG). Hence, by manipulating these MSG tokens, one can flexibly exchange visual information across regions and the computational complexity is reduced. We then integrate the MSG token into a multi-scale architecture named MSG-Transformer. In standard image classification and object detection, MSG-Transformer achieves competitive performance and the inference on both GPU and CPU is accelerated. Code is available at https://github.com/hustvl/MSG-Transformer.

Discussion (0). Sign in to comment.

Pith tools