Pith. sign in

REVIEW 1 cited by

Saturated Non-Monotonic Activation Functions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07537 v2 pith:QACDKC7S submitted 2023-05-12 cs.NE cs.LG

classification cs.NEcs.LG
keywords functionsactivationnon-monotonicreludeeplearningportionpositive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Activation functions are essential to deep learning networks. Popular and versatile activation functions are mostly monotonic functions, some non-monotonic activation functions are being explored and show promising performance. But by introducing non-monotonicity, they also alter the positive input, which is proved to be unnecessary by the success of ReLU and its variants. In this paper, we double down on the non-monotonic activation functions' development and propose the Saturated Gaussian Error Linear Units by combining the characteristics of ReLU and non-monotonic activation functions. We present three new activation functions built with our proposed method: SGELU, SSiLU, and SMish, which are composed of the negative portion of GELU, SiLU, and Mish, respectively, and ReLU's positive portion. The results of image classification experiments on CIFAR-100 indicate that our proposed activation functions are highly effective and outperform state-of-the-art baselines across multiple deep learning architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A parametric activation function based on Wendland RBF

    cs.LG 2025-06 reject novelty 5.0 of 10

    A Wendland-RBF-based activation with linear and exponential terms is reported to outperform ReLU on Fashion-MNIST, but the evidence in the preprint is insufficient to verify the result.

Pith tools