← back to paper
arxiv: 2508.20577 · 2 revisions
MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training