Pith. sign in

REVIEW 2 cited by

Revise Saturated Activation Functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1602.05980 v2 pith:DT3VTURT submitted 2016-02-18 cs.LG

classification cs.LG
keywords functionsactivationdeeplogistictanhcomparablefunctionnetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we revise two commonly used saturated functions, the logistic sigmoid and the hyperbolic tangent (tanh). We point out that, besides the well-known non-zero centered property, slope of the activation function near the origin is another possible reason making training deep networks with the logistic function difficult to train. We demonstrate that, with proper rescaling, the logistic sigmoid achieves comparable results with tanh. Then following the same argument, we improve tahn by penalizing in the negative part. We show that "penalized tanh" is comparable and even outperforms the state-of-the-art non-saturated functions including ReLU and leaky ReLU on deep convolution neural networks. Our results contradict to the conclusion of previous works that the saturation property causes the slow convergence. It suggests further investigation is necessary to better understand activation functions in deep architectures.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scaling Up Thermodynamic AI Models

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    A backpropagation training method for deep conv nets enables thermodynamic inference on Ising hardware with reported CIFAR accuracies plus theory bounding inference cost versus accuracy.

  2. A parametric activation function based on Wendland RBF

    cs.LG 2025-06 reject novelty 5.0 of 10

    A Wendland-RBF-based activation with linear and exponential terms is reported to outperform ReLU on Fashion-MNIST, but the evidence in the preprint is insufficient to verify the result.

Pith tools