Experimental study finds Recursive-Transformer for ASR encoders achieves comparable performance with 66% fewer parameters when limited recursion is applied in the latent space.
Subformer: Exploring weight sharing for parameter efficiency in generative transformers
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4roles
background 1polarities
background 1representative citing papers
A 53K-parameter weight-shared transformer generates novel valid SMILES at 95% rate on ZINC-250K and resolves constraints hierarchically via bracket, ring, and valence stages as shown by probing and ablation.
BWTA achieves near full-precision accuracy on BERT and LLMs using binary weights and ternary activations, with 16-24x kernel speedups via specialized CUDA kernels.
citing papers explorer
-
Rethinking Depth: A study of the Recursive-Transformer for Speech Recognition
Experimental study finds Recursive-Transformer for ASR encoders achieves comparable performance with 66% fewer parameters when limited recursion is applied in the latent space.
-
SMolLM: Small Language Models Learn Small Molecular Grammar
A 53K-parameter weight-shared transformer generates novel valid SMILES at 95% rate on ZINC-250K and resolves constraints hierarchically via bracket, ring, and valence stages as shown by probing and ablation.
-
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
BWTA achieves near full-precision accuracy on BERT and LLMs using binary weights and ternary activations, with 16-24x kernel speedups via specialized CUDA kernels.
- Do Transformers Need Three Projections? Systematic Study of QKV Variants