REVIEW 7 cited by
Learning Activation Functions to Improve Deep Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Artificial neural networks typically have a fixed, non-linear activation function at each neuron. We have designed a novel form of piecewise linear activation function that is learned independently for each neuron using gradient descent. With this adaptive activation function, we are able to improve upon deep neural network architectures composed of static rectified linear units, achieving state-of-the-art performance on CIFAR-10 (7.51%), CIFAR-100 (30.83%), and a benchmark from high-energy physics involving Higgs boson decay modes.
Forward citations
Cited by 7 Pith papers
-
GNet: A scalable and flexible Gaussian process network with nonparametric neurons
GNet models nonparametric 1D GP activations and uses a jointly inverse Kalman filter plus closed-form gradients to train and predict without forming covariance matrices, matching or beating baselines at large n.
-
Rethinking Neural Nonlinearity as Gating
Common activations and softmax are instances of a single input-conditioned Threshold Gating primitive with few branches, enabling lossless conversion and a unified analog implementation path.
-
More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations
Mixture of Activations mixes activation functions token-adaptively in FFNs via lightweight gates, strictly more expressive than fixed or learnable activations, and yields lower pretraining loss from 0.12B to 2B models.
-
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
On VoiceBank-DEMAND, replacing dense layers or ReLU activations with GR-KAN layers in MP-SENet and Demucs improved PESQ by up to 0.1 and cut parameters by up to 4x in one comparison.
-
Quantum Variational Activation Functions Empower Kolmogorov-Arnold Networks
QKANs show strong empirical performance on regression, vision, and language tasks, but the claimed exponential parameter reduction is not rigorously established.
-
Graph Interpolating Activation Improves Both Natural and Robust Accuracies in Data-Efficient Deep Learning
Graph Laplacian interpolating activation replaces softmax in DNNs and improves natural accuracy, robust accuracy, and data efficiency.
-
Hybrid activation functions for deep neural networks: S3 and S4 -- a novel approach to gradient flow optimization
A new parameterized activation function, S4, that blends sigmoid and softsign through a smooth sigmoid-weighted transition is claimed to improve accuracy and convergence on small neural network benchmarks.
Discussion (0). Continue with ORCID to comment.