Stochastically rounding a single pretrained model into several low-precision copies yields a training-free ensemble that improves NLL and calibration on large models.
The evaluation of MMLU was conducted using the template provided in the official repository2, and the computation was based on a micro-average
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Ex Uno Pluria: Insights on Ensembling in Low Precision Number Systems
Stochastically rounding a single pretrained model into several low-precision copies yields a training-free ensemble that improves NLL and calibration on large models.