Pith. sign in

REVIEW 1 cited by

Comparison of non-linear activation functions for deep neural networks on MNIST classification task

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.02763 v1 pith:ER24PTAH submitted 2018-04-08 cs.LG stat.ML

Comparison of non-linear activation functions for deep neural networks on MNIST classification task

classification cs.LG stat.ML
keywords networksactivationfunctionsneuralwillanalysedinfluenceperformances
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Activation functions play a key role in neural networks so it becomes fundamental to understand their advantages and disadvantages in order to achieve better performances. This paper will first introduce common types of non linear activation functions that are alternative to the well known sigmoid function and then evaluate their characteristics. Moreover deeper neural networks will be analysed because they positively influence the final performances compared to shallower networks. They also strictly depend on the weight initialisation hence the effect of drawing weights from Gaussian and uniform distribution will be analysed making particular attention on how the number of incoming and outgoing connection to a node influence the whole network.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FlexAct: Why Learn when you can Pick?

    cs.LG 2026-01 reject novelty 2.0

    A Gumbel-Softmax router that discretely selects among five fixed activation functions, plus a gradient-norm regularizer, recovers the generating activation on toy regression tasks but never beats the matching fixed ac...