Pith. sign in

REVIEW 2 cited by

Pay Attention to MLPs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.08050 v2 pith:I55MIEJI submitted 2021-05-17 cs.LG cs.CLcs.CV

classification cs.LGcs.CLcs.CV
keywords transformersgmlpmlpsmodeltasksvisionwellaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers have become one of the most important architectural innovations in deep learning and have enabled many breakthroughs over the past few years. Here we propose a simple network architecture, gMLP, based on MLPs with gating, and show that it can perform as well as Transformers in key language and vision applications. Our comparisons show that self-attention is not critical for Vision Transformers, as gMLP can achieve the same accuracy. For BERT, our model achieves parity with Transformers on pretraining perplexity and is better on some downstream NLP tasks. On finetuning tasks where gMLP performs worse, making the gMLP model substantially larger can close the gap with Transformers. In general, our experiments show that gMLP can scale as well as Transformers over increased data and compute.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity

    cs.LG 2026-07 conditional novelty 6.0 of 10

    SDM sparsifies the Gated DeltaNet update rule to enable 1000x larger recurrent memory states at iso-FLOP, improving long-context recall and short-context reasoning over GDN and matching full attention at 8B scale.

  2. A Lightweight Foundation Model for Collider Physics with Multi-Domain Adaptation

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A lightweight autoencoder pre-trained on LHC track data transfers to collider and out-of-domain scientific tasks, matching a transformer within ~2% at about 46x lower per-epoch training cost.

Pith tools