← back to paper
arxiv: 2607.10611 · 2 revisions
M+Adam: Low-Precision Training via Additive-Multiplicative Optimization