On the large learning rate dynamics of SGD

Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, Guy Gur-Ari · 2006 · arXiv 2006.10265

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

representative citing papers

Phantom transitions in language model fine-tuning

cs.CL · 2026-05-25 · accept · novelty 7.0

Apparent phase transitions during fine-tuning on near-synonym tasks are phantoms originating in the softmax readout; an order parameter isolates kinematic and structural failure modes and a few dimensionless quantities predict critical learning rates across architectures via blind test.

citing papers explorer

Showing 1 of 1 citing paper after filters.

Phantom transitions in language model fine-tuning cs.CL · 2026-05-25 · accept · none · ref 20
Apparent phase transitions during fine-tuning on near-synonym tasks are phantoms originating in the softmax readout; an order parameter isolates kinematic and structural failure modes and a few dimensionless quantities predict critical learning rates across architectures via blind test.

On the large learning rate dynamics of SGD

fields

years

verdicts

representative citing papers

citing papers explorer