In one-dimensional denoising score matching with two-layer ReLU networks, a large SGD learning rate provably prevents the learned score from getting close to the empirical optimal score, mitigating memorization.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
In one-dimensional denoising score matching with two-layer ReLU networks, a large SGD learning rate provably prevents the learned score from getting close to the empirical optimal score, mitigating memorization.