Pith. sign in

REVIEW 1 cited by

Foundation Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.06423 v2 pith:HJK7R5VC submitted 2022-10-12 cs.LG cs.CLcs.CV

classification cs.LGcs.CLcs.CV
keywords transformertransformersvisionbertbetterfoundationlanguagemodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different implementations for better performance, e.g., Post-LayerNorm for BERT, and Pre-LayerNorm for GPT and vision Transformers. We call for the development of Foundation Transformer for true general-purpose modeling, which serves as a go-to architecture for various tasks and modalities with guaranteed training stability. In this work, we introduce a Transformer variant, named Magneto, to fulfill the goal. Specifically, we propose Sub-LayerNorm for good expressivity, and the initialization strategy theoretically derived from DeepNet for stable scaling up. Extensive experiments demonstrate its superior performance and better stability than the de facto Transformer variants designed for various applications, including language modeling (i.e., BERT, and GPT), machine translation, vision pretraining (i.e., BEiT), speech recognition, and multimodal pretraining (i.e., BEiT-3).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement

    cs.CV 2025-07 conditional novelty 5.0 of 10

    CSCI uses color histograms as self-supervised targets and a two-step S2A self-attention to reduce clothing-color bias, improving CC-ReID on four benchmarks.

Pith tools