A multi-branch CNN fed with the trEFM frequency trace and cantilever parameters extracts bi-exponential kinetic parameters (τ1, τ2, A) more accurately and noise-robustly than the prior single-exponential neural network.
Channel Normalization in Convolutional Neural Network avoids Vanishing Gradients
1 Pith paper cite this work, alongside 18 external citations. Polarity classification is still indexing.
abstract
Normalization layers are widely used in deep neural networks to stabilize training. In this paper, we consider the training of convolutional neural networks with gradient descent on a single training example. This optimization problem arises in recent approaches for solving inverse problems such as the deep image prior or the deep decoder. We show that for this setup, channel normalization, which centers and normalizes each channel individually, avoids vanishing gradients, whereas, without normalization, gradients vanish which prevents efficient optimization. This effect prevails in deep single-channel linear convolutional networks, and we show that without channel normalization, gradient descent takes at least exponentially many steps to come close to an optimum. Contrary, with channel normalization, the gradients remain bounded, thus avoiding exploding gradients.
fields
cond-mat.mtrl-sci 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-Output Convolutional Neural Network for Improved Parameter Extraction in Time-Resolved Electrostatic Force Microscopy Data
A multi-branch CNN fed with the trEFM frequency trace and cantilever parameters extracts bi-exponential kinetic parameters (τ1, τ2, A) more accurately and noise-robustly than the prior single-exponential neural network.