Pith. sign in

REVIEW 2 cited by

SSN: Learning Sparse Switchable Normalization via SparsestMax

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.03793 v1 pith:PADC7PSZ submitted 2019-03-09 cs.CV

classification cs.CV
keywords normalizationsparseoptimizationswitchablevariouscomputationsconstraineddifferent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Normalization methods improve both optimization and generalization of ConvNets. To further boost performance, the recently-proposed switchable normalization (SN) provides a new perspective for deep learning: it learns to select different normalizers for different convolution layers of a ConvNet. However, SN uses softmax function to learn importance ratios to combine normalizers, leading to redundant computations compared to a single normalizer. This work addresses this issue by presenting Sparse Switchable Normalization (SSN) where the importance ratios are constrained to be sparse. Unlike $\ell_1$ and $\ell_0$ constraints that impose difficulties in optimization, we turn this constrained optimization problem into feed-forward computation by proposing SparsestMax, which is a sparse version of softmax. SSN has several appealing properties. (1) It inherits all benefits from SN such as applicability in various tasks and robustness to a wide range of batch sizes. (2) It is guaranteed to select only one normalizer for each normalization layer, avoiding redundant computations. (3) SSN can be transferred to various tasks in an end-to-end manner. Extensive experiments show that SSN outperforms its counterparts on various challenging benchmarks such as ImageNet, Cityscapes, ADE20K, and Kinetics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptively Sparse Transformers

    cs.CL 2019-08 conditional novelty 7.0 of 10

    An adaptively sparse Transformer with per-head learned α-entmax attention yields sparser, more confident attention heads and slight BLEU gains over softmax Transformers on four machine translation datasets.

  2. Attentive Normalization

    cs.CV 2019-08 conditional novelty 6.0 of 10

    An attention-weighted mixture of K affine transforms inside normalization improves ImageNet top-1 accuracy by 0.5-2.7% and COCO AP by up to 1.8%/2.2% across four architectures.

Pith tools