Pith. sign in

REVIEW 2 cited by

Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.00234 v3 pith:EW5LM7YT submitted 2021-01-01 cs.CL cs.LG

classification cs.CLcs.LG
keywords parametersharinggenerativesubformertransformersmethodsmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Transformers have shown improved performance when compared to previous architectures for sequence processing such as RNNs. Despite their sizeable performance gains, as recently suggested, the model is computationally expensive to train and with a high parameter budget. In light of this, we explore parameter-sharing methods in Transformers with a specific focus on generative models. We perform an analysis of different parameter sharing/reduction methods and develop the Subformer. Our model combines sandwich-style parameter sharing, which overcomes naive cross-layer parameter sharing in generative models, and self-attentive embedding factorization (SAFE). Experiments on machine translation, abstractive summarization and language modeling show that the Subformer can outperform the Transformer even when using significantly fewer parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mobius Learning: Cyclic Depth Folding in Transformers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Möbius Learning, which cyclically shifts block order across data streams, achieves lower validation loss than fixed-order looped training at loop depths 6, 10, and 15 in a 124M-parameter GPT-2 experiment.

  2. FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A block-pruning and model-extension method for LLMs that replaces pruned blocks with weight-shared blocks plus low-rank adapters, reporting state-of-the-art recovery on several benchmarks.

Pith tools