A natively 1.58-bit, 2B-parameter model trained on 4T tokens roughly matches 1-2B full-precision open LLMs on average across 16 benchmarks while using far less memory and energy.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
BitNet b1.58 2B4T Technical Report
A natively 1.58-bit, 2B-parameter model trained on 4T tokens roughly matches 1-2B full-precision open LLMs on average across 16 benchmarks while using far less memory and energy.