REVIEW 5 cited by
Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Major winning Convolutional Neural Networks (CNNs), such as AlexNet, VGGNet, ResNet, GoogleNet, include tens to hundreds of millions of parameters, which impose considerable computation and memory overhead. This limits their practical use for training, optimization and memory efficiency. On the contrary, light-weight architectures, being proposed to address this issue, mainly suffer from low accuracy. These inefficiencies mostly stem from following an ad hoc procedure. We propose a simple architecture, called SimpleNet, based on a set of designing principles, with which we empirically show, a well-crafted yet simple and reasonably deep architecture can perform on par with deeper and more complex architectures. SimpleNet provides a good tradeoff between the computation/memory efficiency and the accuracy. Our simple 13-layer architecture outperforms most of the deeper and complex architectures to date such as VGGNet, ResNet, and GoogleNet on several well-known benchmarks while having 2 to 25 times fewer number of parameters and operations. This makes it very handy for embedded systems or systems with computational and memory limitations. We achieved state-of-the-art result on CIFAR10 outperforming several heavier architectures, near state of the art on MNIST and competitive results on CIFAR100 and SVHN. We also outperformed the much larger and deeper architectures such as VGGNet and popular variants of ResNets among others on the ImageNet dataset. Models are made available at: https://github.com/Coderx7/SimpleNet
Forward citations
Cited by 5 Pith papers
-
Solving MNIST with a globally trained Mixture of Quantum Experts
A globally trained mixture of 16 quantum experts classifies full-resolution MNIST parity with 97.5% test accuracy using 10 qubits, and joint training improves compute-efficiency until saturation.
-
Mish: A Self Regularized Non-Monotonic Activation Function
Mish, a smooth non-monotonic activation function, is proposed and shown to often match or beat ReLU, Swish, and Leaky ReLU on CIFAR-10, ImageNet, and MS-COCO benchmarks.
-
DEX: Data Channel Extension for Efficient CNN Inference on Tiny AI Accelerators
DEX raises TinyML classification accuracy by an average of 3.5 percentage points at no extra inference latency by stacking evenly sampled image patches into spare input channels.
-
Improved Background Estimation for Gas Plume Identification in Hyperspectral Images
Per-plume tuned background estimators, especially K-Nearest Segments, substantially raise neural network confidence for gas plume identification on 640 simulated LWIR images, but the evaluation uses oracle hyperparame...
-
Machine-learning enables Image Reconstruction and Classification in a "see-through" camera
A U-net reconstructs MNIST, EMNIST, and Kanji49 images from raw sensor data of a see-through lensless camera, but classification benefits are inconsistent and the manuscript is incomplete.
Discussion (0). Continue with ORCID to comment.