Temporal correlations from lazy random walks enable efficient SGD learning of k-juntas via temporal-difference loss on ReLU networks, achieving linear sample complexity in d.
arXiv preprint arXiv:2205.15809 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
ReWA uses reparameterization with weight decay and adaptive rates to create a stable optimization landscape connected to ℓ_p regularization, yielding higher sparsity than ℓ1 on ResNets for CIFAR-10 and ImageNet while preserving accuracy.
citing papers explorer
-
The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently
Temporal correlations from lazy random walks enable efficient SGD learning of k-juntas via temporal-difference loss on ReLU networks, achieving linear sample complexity in d.
-
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
ReWA uses reparameterization with weight decay and adaptive rates to create a stable optimization landscape connected to ℓ_p regularization, yielding higher sparsity than ℓ1 on ResNets for CIFAR-10 and ImageNet while preserving accuracy.