Pith. sign in

REVIEW 5 cited by

Pad\'e Activation Units: End-to-end Learning of Flexible Activation Functions in Deep Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.06732 v3 pith:42TDTCI7 submitted 2019-07-15 cs.LG cs.NE

classification cs.LGcs.NE
keywords activationfunctionsdeeplearningpauschoicedependsend-to-end
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The performance of deep network learning strongly depends on the choice of the non-linear activation function associated with each neuron. However, deciding on the best activation is non-trivial, and the choice depends on the architecture, hyper-parameters, and even on the dataset. Typically these activations are fixed by hand before training. Here, we demonstrate how to eliminate the reliance on first picking fixed activation functions by using flexible parametric rational functions instead. The resulting Pad\'e Activation Units (PAUs) can both approximate common activation functions and also learn new ones while providing compact representations. Our empirical evidence shows that end-to-end learning deep networks with PAUs can increase the predictive performance. Moreover, PAUs pave the way to approximations with provable robustness. https://github.com/ml-research/pau

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 16 citations worldwide. Full citation record

  1. Statistics of correlations in nonlinear recurrent neural networks

    q-bio.NC 2025-10 unverdicted novelty 7.0 of 10

    Derives exact correlation statistics for nonlinear RNNs in the large-N limit with Gaussian quenched disorder using path integrals, generalizing linear results and adding 1/N corrections.

  2. Balancing Expressivity and Robustness: Constrained Rational Activations for Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A constrained rational activation with denominator degree one larger than numerator and no constant term stabilizes high-UTD continuous control, while trading off long-term plasticity.

  3. Human Body Weight Estimation Through Music-Induced Bed Vibrations

    eess.SP 2025-09 conditional novelty 5.0 of 10

    Body weight can be estimated from music-induced bed vibration transfer functions with 1.55 kg (wooden bed) and 4.36 kg (steel bed) mean absolute error, using a physics-inspired rational-activation network.

  4. MNO : A Multi-modal Neural Operator for Parametric Nonlinear BVPs

    cs.CE 2025-07 conditional novelty 5.0 of 10

    The paper introduces MNO, an FMM-inspired neural operator that jointly maps PDE coefficients, source terms, and boundary conditions to the solution, and shows it works on 1D Poisson, Darcy flow, and a nonlinear BVP.

  5. A parametric activation function based on Wendland RBF

    cs.LG 2025-06 reject novelty 5.0 of 10

    A Wendland-RBF-based activation with linear and exponential terms is reported to outperform ReLU on Fashion-MNIST, but the evidence in the preprint is insufficient to verify the result.

Pith tools