REVIEW 4 cited by
Correlations Are Ruining Your Gradient Descent
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Herein the topics of (natural) gradient descent, data decorrelation, and approximate methods for backpropagation are brought into a common discussion. Natural gradient descent illuminates how gradient vectors, pointing at directions of steepest descent, can be improved by considering the local curvature of loss landscapes. We extend this perspective and show that to fully solve the problem illuminated by natural gradients in neural networks, one must recognise that correlations in the data at any linear transformation, including node responses at every layer of a neural network, cause a non-orthonormal relationship between the model's parameters. To solve this requires a method for decorrelating inputs at each individual layer of a neural network. We describe a range of methods which have been proposed for decorrelation and whitening of node output, and expand on these to provide a novel method specifically useful for distributed computing and computational neuroscience. Implementing decorrelation within multi-layer neural networks, we can show that not only is training via backpropagation sped up significantly but also existing approximations of backpropagation, which have failed catastrophically in the past, benefit significantly in their accuracy and convergence speed. This has the potential to provide a route forward for approximate gradient descent methods which have previously been discarded, training approaches for analogue and neuromorphic hardware, and potentially insights as to the efficacy and utility of decorrelation processes in the brain.
Forward citations
Cited by 4 Pith papers
-
Conditioned Direct Feedback Alignment via Activity and Error Geometry
Conditioning the activity and error factors of the direct feedback alignment update with damped inverse second moments improves DFA accuracy in nuisance-dominated regimes and clean confirmations.
-
Mistake gating leads to energy and memory efficient continual learning
Mistake-gated plasticity reduces neural network updates by 50-80% by gating changes on classification errors, improving efficiency for continual learning without added hyperparameters.
-
White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
WARM builds few-shot segmentation prototypes by whitening support features before cross-attention and restoring their statistics afterward, reducing random-seed dependence and improving accuracy on S3DIS and ScanNet.
-
Real-Time Decorrelation-Based Anomaly Detection for Multivariate Time Series
DAD detects multivariate time-series anomalies by measuring the change in an online-learned decorrelation matrix, achieving the best mean AUC (0.8027) among 15 methods on 50 datasets.
Discussion (0). Continue with ORCID to comment.