REVIEW 2 cited by
Hardware-aware training for large-scale and diverse deep learning inference workloads using in-memory computing-based accelerators
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Analog in-memory computing (AIMC) -- a promising approach for energy-efficient acceleration of deep learning workloads -- computes matrix-vector multiplications (MVMs) but only approximately, due to nonidealities that often are non-deterministic or nonlinear. This can adversely impact the achievable deep neural network (DNN) inference accuracy as compared to a conventional floating point (FP) implementation. While retraining has previously been suggested to improve robustness, prior work has explored only a few DNN topologies, using disparate and overly simplified AIMC hardware models. Here, we use hardware-aware (HWA) training to systematically examine the accuracy of AIMC for multiple common artificial intelligence (AI) workloads across multiple DNN topologies, and investigate sensitivity and robustness to a broad set of nonidealities. By introducing a new and highly realistic AIMC crossbar-model, we improve significantly on earlier retraining approaches. We show that many large-scale DNNs of various topologies, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and transformers, can in fact be successfully retrained to show iso-accuracy on AIMC. Our results further suggest that AIMC nonidealities that add noise to the inputs or outputs, not the weights, have the largest impact on DNN accuracy, and that RNNs are particularly robust to all nonidealities.
Forward citations
Cited by 2 Pith papers
-
AnalogNAS-Bench: A NAS Benchmark for Analog In-Memory Computing
AnalogNAS-Bench extends NAS-Bench-201 with analog in-memory computing metrics, revealing that architecture rankings under quantization do not transfer to analog noise, and that 3x3 convolutions, pooling, and skip conn...
-
Rapid yet accurate Tile-circuit and device modeling for Analog In-Memory Computing
A python Tile-circuit model reproduces analog matrix-vector multiply outputs from circuit simulation to 99.999% R², and reveals that Gaussian-noise hardware-aware training is insufficient against instantaneous-current...
Discussion (0). Continue with ORCID to comment.