Pith. sign in

REVIEW 7 cited by

Learning Activation Functions to Improve Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1412.6830 v3 pith:5BGH6R37 submitted 2014-12-21 cs.NE cs.CVcs.LGstat.ML

classification cs.NEcs.CVcs.LGstat.ML
keywords activationfunctionneuraldeepimprovelinearnetworksneuron
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Artificial neural networks typically have a fixed, non-linear activation function at each neuron. We have designed a novel form of piecewise linear activation function that is learned independently for each neuron using gradient descent. With this adaptive activation function, we are able to improve upon deep neural network architectures composed of static rectified linear units, achieving state-of-the-art performance on CIFAR-10 (7.51%), CIFAR-100 (30.83%), and a benchmark from high-energy physics involving Higgs boson decay modes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GNet: A scalable and flexible Gaussian process network with nonparametric neurons

    stat.ME 2026-07 conditional novelty 6.5 of 10

    GNet models nonparametric 1D GP activations and uses a jointly inverse Kalman filter plus closed-form gradients to train and predict without forming covariance matrices, matching or beating baselines at large n.

  2. Rethinking Neural Nonlinearity as Gating

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Common activations and softmax are instances of a single input-conditioned Threshold Gating primitive with few branches, enabling lossless conversion and a unified analog implementation path.

  3. More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Mixture of Activations mixes activation functions token-adaptively in FFNs via lightweight gates, strictly more expressive than fixed or learnable activations, and yields lower pretraining loss from 0.12B to 2B models.

  4. From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology

    eess.AS 2024-12 conditional novelty 6.0 of 10

    On VoiceBank-DEMAND, replacing dense layers or ReLU activations with GR-KAN layers in MP-SENet and Demucs improved PESQ by up to 0.1 and cut parameters by up to 4x in one comparison.

  5. Quantum Variational Activation Functions Empower Kolmogorov-Arnold Networks

    quant-ph 2025-09 reject novelty 5.0 of 10

    QKANs show strong empirical performance on regression, vision, and language tasks, but the claimed exponential parameter reduction is not rigorously established.

  6. Graph Interpolating Activation Improves Both Natural and Robust Accuracies in Data-Efficient Deep Learning

    cs.LG 2019-07 unverdicted novelty 5.0 of 10

    Graph Laplacian interpolating activation replaces softmax in DNNs and improves natural accuracy, robust accuracy, and data efficiency.

  7. Hybrid activation functions for deep neural networks: S3 and S4 -- a novel approach to gradient flow optimization

    cs.LG 2025-07 reject novelty 3.0 of 10

    A new parameterized activation function, S4, that blends sigmoid and softsign through a smooth sigmoid-weighted transition is claimed to improve accuracy and convergence on small neural network benchmarks.

Pith tools