REVIEW 2 cited by
0/1 Deep Neural Networks via Block Coordinate Descent
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
The step function is one of the simplest and most natural activation functions for deep neural networks (DNNs). As it counts 1 for positive variables and 0 for others, its intrinsic characteristics (e.g., discontinuity and no viable information of subgradients) impede its development for several decades. Even if there is an impressive body of work on designing DNNs with continuous activation functions that can be deemed as surrogates of the step function, it is still in the possession of some advantageous properties, such as complete robustness to outliers and being capable of attaining the best learning-theoretic guarantee of predictive accuracy. Hence, in this paper, we aim to train DNNs with the step function used as an activation function (dubbed as 0/1 DNNs). We first reformulate 0/1 DNNs as an unconstrained optimization problem and then solve it by a block coordinate descend (BCD) method. Moreover, we acquire closed-form solutions for sub-problems of BCD as well as its convergence properties. Furthermore, we also integrate $\ell_{2,0}$-regularization into 0/1 DNN to accelerate the training process and compress the network scale. As a result, the proposed algorithm has a high performance on classifying MNIST and Fashion-MNIST datasets. As a result, the proposed algorithm has a desirable performance on classifying MNIST, FashionMNIST, Cifar10, and Cifar100 datasets.
Forward citations
Cited by 2 Pith papers
-
Beyond Low-rank Decomposition: A Shortcut Approach for Efficient On-Device Learning
ASI compresses training activations with a single warm-started subspace iteration and a once-per-model rank selection, cutting on-device training memory by up to 120x and FLOPs by up to 1.86x on standard benchmarks.
-
KPL: Training-Free Medical Knowledge Mining of Vision-Language Models
KPL combines LLM-generated class descriptions, visual retrieval, and a log-space Greenkhorn algorithm to boost CLIP zero-shot accuracy on medical and natural image datasets.
Discussion (0). Continue with ORCID to comment.