This paper introduces Zamba2, a suite of 1.2B, 2.7B, and 7.4B hybrid Mamba2-transformer models that claims state-of-the-art small-model quality and 30-50% lower time-to-first-token, with open weights and a 5T-token pretraining dataset.
The Zyphra Cookbook
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Zamba2 Suite: Technical Report
This paper introduces Zamba2, a suite of 1.2B, 2.7B, and 7.4B hybrid Mamba2-transformer models that claims state-of-the-art small-model quality and 30-50% lower time-to-first-token, with open weights and a 5T-token pretraining dataset.