Pith. sign in

REVIEW 1 cited by

Adaptive Friction in Deep Learning: Enhancing Optimizers with Sigmoid and Tanh Function

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11839 v1 pith:67LWSHFP submitted 2024-08-07 cs.LG cs.AI

classification cs.LGcs.AI
keywords optimizersadaptivefrictiondeepalgorithmscoefficientsconvergenceexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adaptive optimizers are pivotal in guiding the weight updates of deep neural networks, yet they often face challenges such as poor generalization and oscillation issues. To counter these, we introduce sigSignGrad and tanhSignGrad, two novel optimizers that integrate adaptive friction coefficients based on the Sigmoid and Tanh functions, respectively. These algorithms leverage short-term gradient information, a feature overlooked in traditional Adam variants like diffGrad and AngularGrad, to enhance parameter updates and convergence.Our theoretical analysis demonstrates the wide-ranging adjustment capability of the friction coefficient S, which aligns with targeted parameter update strategies and outperforms existing methods in both optimization trajectory smoothness and convergence rate. Extensive experiments on CIFAR-10, CIFAR-100, and Mini-ImageNet datasets using ResNet50 and ViT architectures confirm the superior performance of our proposed optimizers, showcasing improved accuracy and reduced training time. The innovative approach of integrating adaptive friction coefficients as plug-ins into existing optimizers, exemplified by the sigSignAdamW and sigSignAdamP variants, presents a promising strategy for boosting the optimization performance of established algorithms. The findings of this study contribute to the advancement of optimizer design in deep learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Few-Shot Learning with Adaptive Weight Masking in Conditional GANs

    cs.CV 2024-12 reject novelty 2.0 of 10

    A CGAN with residual generator blocks and a heuristically computed weight mask on the discriminator reports better IS/FID on MNIST, but the paper lacks downstream few-shot evaluation and any code or training details.

Pith tools