Pith. sign in

hub

Mish: A self regularized non-monotonic activation function

21 Pith papers cite this work. Polarity classification is still indexing.

21 Pith papers citing it
abstract

We propose $\textit{Mish}$, a novel self-regularized non-monotonic activation function which can be mathematically defined as: $f(x)=x\tanh(softplus(x))$. As activation functions play a crucial role in the performance and training dynamics in neural networks, we validated experimentally on several well-known benchmarks against the best combinations of architectures and activation functions. We also observe that data augmentation techniques have a favorable effect on benchmarks like ImageNet-1k and MS-COCO across multiple architectures. For example, Mish outperformed Leaky ReLU on YOLOv4 with a CSP-DarkNet-53 backbone on average precision ($AP_{50}^{val}$) by 2.1$\%$ in MS-COCO object detection and ReLU on ResNet-50 on ImageNet-1k in Top-1 accuracy by $\approx$1$\%$ while keeping all other network parameters and hyperparameters constant. Furthermore, we explore the mathematical formulation of Mish in relation with the Swish family of functions and propose an intuitive understanding on how the first derivative behavior may be acting as a regularizer helping the optimization of deep neural networks. Code is publicly available at https://github.com/digantamisra98/Mish.

hub tools

citation-role summary

background 1 method 1

citation-polarity summary

representative citing papers

Empowering Multi-Robot Cooperation via Sequential World Models

cs.RO · 2025-09-16 · unverdicted · novelty 6.0

SeqWM introduces sequential autoregressive agent-wise world models for multi-robot MBRL, outperforming baselines in performance and sample efficiency on Bi-DexHands and Multi-Quadruped tasks with physical robot deployment.

YOLOv4: Optimal Speed and Accuracy of Object Detection

cs.CV · 2020-04-23 · unverdicted · novelty 5.0

YOLOv4 achieves 43.5% AP (65.7% AP50) on MS COCO at ~65 FPS on Tesla V100 by integrating WRC, CSP, CmBN, SAT, Mish activation, Mosaic augmentation, DropBlock, and CIoU loss.

Performance of the Eos detector with water

hep-ex · 2026-06-08 · unverdicted · novelty 4.0

First calibration results from the Eos four-tonne water Cherenkov detector validate simulations and reconstruction methods using deployed optical and radioactive sources.

Voice Biomarkers for Depression and Anxiety

cs.LG · 2026-05-11 · unverdicted · novelty 4.0

Deep learning models extract content-agnostic voice biomarkers for depression and anxiety from a ~65k-utterance proprietary dataset, achieving 71% sensitivity and specificity when combined with lexical features.

citing papers explorer

Showing 21 of 21 citing papers.