Stochastically rounding a single pretrained model into several low-precision copies yields a training-free ensemble that improves NLL and calibration on large models.
Table 5 summarizes our experimental results using the Adam optimizer (Kingma and Ba, 2015)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Ex Uno Pluria: Insights on Ensembling in Low Precision Number Systems
Stochastically rounding a single pretrained model into several low-precision copies yields a training-free ensemble that improves NLL and calibration on large models.