Making network size itself a trainable parameter via an auxiliary weight or a controller mask lets small networks grow during gradient descent and outperform equivalent fixed-size networks on toy tasks.
Both approaches require a smooth transition function ψ(x) that approximates a step function while maintain- ing differentiability
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Growing Neural Networks: Dynamic Evolution through Gradient Descent
Making network size itself a trainable parameter via an auxiliary weight or a controller mask lets small networks grow during gradient descent and outperform equivalent fixed-size networks on toy tasks.