← back to paper
arxiv: 2605.22297 · 2 revisions
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs