REVIEW 4 major objections 5 minor 101 references
SPRING: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference
T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read SPRING claims that a binary-mask sparsity scheme plus stochastic-rounded fixed-point arithmetic can make CNN training and inference far faster and more energy-efficient, reporting 15.6x training speedup over a GTX 1080 Ti.
desk verdict A plausible new architecture combining sparsity, reduced-precision training, and monolithic 3D memory, but the headline speedups rest on an unjustified 50% uniform sparsity assumption that conflates activation with weight sparsity, making the numbers conditional estimates rather than verified results. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the binary-mask sparsity scheme. A mask bit accompanies each activation or weight element; zero-valued elements are removed for storage, and before each multiply-accumulate the activation mask and weight mask are ANDed to identify common nonzero positions. A sequential scanning filter drops 'dangling' nonzeros, those where a nonzero activation aligns with a zero weight or vice versa, and a zero-collapsing shifter presents dense zero-free vectors to the MAC lanes. A post-compute sparsity module recompresses outputs after the activation function. Alongside this, each MAC lane embeds stochastic rounding, which rounds fixed-point results up or down with probability proportional to the discarded fraction, so training can run at 20-bit fixed-point precision without, per the paper's argument, extra convergence epochs. The third leg is a monolithic 3D RRAM interface with 1 KB row-wide buses and decoupled read/write interconnects to keep the PEs fed during memory-intensive training.
What would settle it
Measure per-layer activation and weight sparsity during training and inference on all seven CNNs, then rerun the paper's cycle-accurate simulator with those measured densities; if the average sparsity is below 50%, the claimed 15.6x and 15.5x speedups do not follow.
Extended reading notes
Core claim
The paper argues that sparsity and reduced precision are not just inference techniques: they can be carried through the entire training process. SPRING compresses activations and weights by dropping zero elements and keeping a binary mask that records where the zeros were. Before each multiply-accumulate, the activation mask and weight mask are ANDed so that only positions where both operands are nonzero are sent to the MAC lanes, and a post-compute module compresses newly zeroed outputs after the activation function. To make low-precision training work, each MAC lane embeds stochastic rounding, which rounds a fixed-point result up or down with probability proportional to the discarded fraction, allowing 4 integer bits and 16 fractional bits to train without, the paper argues, extra convergence epochs. A monolithic 3D nonvolatile-memory interface with row-wide 1 KB buses and decoupled read/write paths supplies the bandwidth that training demands. On seven ImageNet CNNs, the paper reports geometric-mean speedups of 15.6x for training and 15.5x for inference over a GTX 1080 Ti, with larger gains on lightweight networks such as MobileNet V2 and smaller gains on memory-bound networks such as VGG-19.
Load-bearing premise
Every reported speedup and energy gain assumes the seven CNNs are uniformly 50% sparse; since SPRING's gains come from skipping zero entries, real sparsity below 50% would shrink all the headline ratios roughly in proportion.
Editorial extensions
If this is right
- Sparsity-aware training removes the usual wall between the training phase and the deployment phase, so a single chip could handle both without sacrificing the speedups that zero-skipping provides.
- Stochastic rounding makes low-precision fixed-point arithmetic viable for backpropagation, which would let future accelerators use simpler, more energy-efficient MAC units than the FP32 units GPUs rely on.
- The reported gains are larger on lightweight CNNs, showing that the monolithic 3D memory interface relieves the bandwidth bottleneck most effectively when the working set fits on chip.
- If the reported batch-level numbers hold, SPRING would reduce training energy by roughly two orders of magnitude relative to a GTX 1080 Ti, making on-device and edge training far more practical.
Reading between the lines
- The headline numbers lean on an assumed uniform 50% sparsity applied to all seven CNNs, but real networks have per-layer and per-phase sparsity that varies; a measured sparsity profile would likely spread the speedups from roughly 5x to over 50x instead of a single 15x average.
- Because stochastic rounding injects pseudo-random noise into every rounding decision, the convergence guarantee depends on the quality of the random source and the exact integer/fraction bit split; sweeping IL and FL bit widths during training would map the accuracy-versus-precision tradeoff directly.
- The pipelined sequential mask filter preserves batch throughput but adds per-image latency, so for single-image edge inference the architecture may be less attractive than the batch-level speedup numbers suggest.
- The paper's architectural ideas are separable: a designer could adopt the binary-mask sparsity scheme without the 3D RRAM stack, or use stochastic rounding alone to improve an existing fixed-point training pipeline, and still capture part of the claimed benefit.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents SPRING, a sparsity-aware reduced-precision monolithic 3D CNN accelerator architecture that targets both training and inference. The design uses binary masks to compress zero-valued activations and weights, a pre-compute sparsity module to skip ineffectual MACs, stochastic rounding for fixed-point training, and a monolithic 3D RRAM memory interface to increase bandwidth. The authors evaluate SPRING on seven ImageNet CNNs with a custom cycle-accurate simulator, and report large geometric-mean improvements relative to a GTX 1080 Ti: 15.6x performance, 4.2x power, and 66.0x energy for training, and 15.5x, 4.5x, and 69.1x for inference. The paper includes RTL synthesis of the accelerator and uses standard tools such as Design Compiler, NVSim, NVMain, and FinCACTI.
Significance. If the reported results are validated, SPRING would be a meaningful contribution to the sparse CNN accelerator literature, especially because it addresses training as well as inference and combines sparsity, reduced precision, and monolithic 3D NVM in a single architecture. The use of a custom cycle-accurate simulator, RTL synthesis, and standard circuit-level memory tools is a strength. However, the central quantitative claims rest on an unverified blanket sparsity assumption and an undocumented GPU baseline, so the specific speedup and energy numbers are not yet established. The architecture itself, with its binary-mask datapath and stochastic-rounding MAC lanes, is plausible and worth further study.
major comments (4)
- [Section 4 (Simulation Methodology)] The blanket assumption "The sparsity of the CNNs are assumed to be 50%" is load-bearing and unsupported. The seven TensorFlow-Slim CNNs are dense, unpruned models, so their exact-zero weight sparsity is negligible; the cited reference [32] reports activation sparsity during training for AlexNet, VGG, and Inception, not weight sparsity. Since SPRING's pre-compute sparsity module skips MACs only when an activation or a weight is exactly zero, the effective fraction of skipped MACs for dense weights is roughly the activation sparsity, not the ~75% that would result from independent 50% activation and 50% weight sparsity. The same issue applies to the memory-compression benefit: with dense weights, weight masks are almost all ones and weight data cannot be compressed. Because skipped-MAC count and compressed-memory traffic feed directly into the cycle-accurate simulator, the reported 15.6x/15.5x speedups and 66.0x/69.1x energy improvements are systematically inflated unless the authors actually prune each network to 50% weight sparsity. Please provide per-layer activation and weight sparsity measurements for each benchmark, or train/prune the models so that the sparsity input to the simulator is realized.
- [Section 4 and Section 5] The GTX 1080 Ti baseline is not described with enough detail to make the comparison reproducible. The paper states only the GPU's peak TFLOPS, memory bandwidth, and die size; it does not explain how execution time, power, or energy are estimated (e.g., which profiler, which cycle-level or analytical model, whether board power or chip power is used, and whether the GPU model accounts for sparsity). Without this information, the claimed GPU-normalized improvements cannot be audited, and the 66.0x energy-efficiency ratio in particular cannot be interpreted. Please add a complete description of the baseline estimation methodology, including the source of all power and timing numbers.
- [Section 6 (Discussions and Limitations) and Section 5] The training results are batch-level simulations, not end-to-end training. The paper acknowledges this in Section 6 and relies on the reference [45] to argue that stochastic rounding with 16 FL bits converges in a similar number of epochs as FP32 training. However, no convergence experiment is run on any of the seven CNNs, so the headline "training speedup" is actually a per-batch throughput improvement with an unverified convergence assumption. Please either run a convergence experiment (even on a reduced dataset or a smaller proxy network) or reword the claims to state explicitly that the results are per-batch, not wall-clock training time to target accuracy.
- [Section 3.1 (Sparsity-aware acceleration), Algorithm 1] The cost of the pre-compute sparsity module is not modeled in sufficient detail to guarantee that the MAC lanes are not stalled. The paper states that the sequential scanning and filtering scheme is pipelined and that throughput is unaffected, but it provides no cycle counts, hardware area, or energy numbers for the mask-generation, dangling-data filtering, and zero-collapsing shifter stages. Since the performance advantages of SPRING come precisely from this module, the cycle-accurate simulator should incorporate its latency and throughput (including the variable-length compression behavior) rather than assume it runs in shadow of the MAC lanes. Please provide implementation data and simulator modeling details for the pre-compute sparsity module.
minor comments (5)
- [Section 2.2] The sentence "the sparsity levels of CNN weights typically range from 20% to 80% [48], [49]" should clarify that these numbers are for pruned/compressed networks, not dense networks, to avoid misleading readers about natural weight sparsity.
- [Figures 15 and 16] The GTX 1080 Ti bars are invisible because the energy-efficiency scale is dominated by SPRING's values; consider a log scale or a table with exact normalized numbers.
- [Equation (4)] The probability expressions should be parenthesized: "with probability (floor(x) + epsilon - x)/epsilon" and "with probability (x - floor(x))/epsilon" to avoid ambiguity.
- [Table 1] The parameter "tBU RST" should be "tBURST" (write burst time).
- [Section 6] The "at most 5%" mask-overhead claim holds relative to uncompressed data but not relative to the compressed data stream; clarify the denominator to avoid confusion.
Circularity Check
No definitional or constructional circularity; headline gains are conditional on an externally-sourced 50% sparsity assumption and author-supplied simulation parameters, but do not reduce by construction to the paper's inputs.
full rationale
The paper's derivation chain is not circular. The central claims are accelerator speedups computed by a cycle-accurate simulator from the architecture in Section 3 and the RTL/NVSim/NVMain/Capo flow of Section 4, compared against a GTX 1080 Ti baseline. The 50% sparsity rate is an explicit input ('The sparsity of the CNNs are assumed to be 50%, as it is shown in [32]...'), not a quantity inferred from SPRING's outputs; the reported 15.6x/66.0x numbers are conditional on that input and would shrink if actual weight/activation sparsity is lower, but they are not algebraically equal to the assumption. This is an assumption-validity and external-benchmark concern, not a circular reduction. The stochastic-rounding convergence assumption is attributed to an external paper [45], and Section 6 openly discloses that training results are batch-level and that monolithic-3D process degradation is not modeled. There are self-citations to prior work by the same group ([69], [70] for the monolithic 3D RRAM interface, [77] for floorplanning, and [84] for the design-space exploration used to set Table 1), and the evaluation infrastructure is the authors' own. However, no load-bearing argument reduces to an unverified self-citation: the binary-mask sparsity skip, reduced-precision MAC with stochastic rounding, and benchmark comparisons are described concretely and are externally comprehensible. No equation (e.g., Eqs. 1-4, mask/AND filtering in Section 3.1) has an output that is identical to an input by construction, and no uniqueness theorem is imported from the authors' prior work. I therefore find no circular steps; score 2 reflects minor non-load-bearing self-citation in the simulation/methodology, not circularity.
Assumptions & free parameters
free parameters (3)
- Uniform 50% sparsity assumption =
50%
- Fixed-point precision selection =
4 IL bits, 16 FL bits
- Accelerator configuration =
700 MHz; 64 PEs; 72 MAC lanes/PE; 16 multipliers/MAC; 24 MB weight buffer; 12 MB activation buffer; 4 MB mask buffer…
assumptions (3)
- domain assumption Stochastic rounding with 16 fractional bits trains CNNs to accuracy comparable to FP32 without extra epochs.
- domain assumption Monolithic 3D integration process-induced device and interconnect degradation is negligible.
- domain assumption The GTX 1080 Ti baseline is modeled correctly from its peak TFLOPS and memory bandwidth.
Cite this review
Pith. "Pith review of SPRING: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference." pith.science (2026). https://pith.science/paper/QWL3WDSM
@misc{pith2026190900557,
author = {Pith},
title = {Pith review of: SPRING: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/QWL3WDSM}},
note = {Machine review of arXiv:1909.00557}
}
read the original abstract
CNNs outperform traditional machine learning algorithms across a wide range of applications. However, their computational complexity makes it necessary to design efficient hardware accelerators. Most CNN accelerators focus on exploring dataflow styles that exploit computational parallelism. However, potential performance speedup from sparsity has not been adequately addressed. The computation and memory footprint of CNNs can be significantly reduced if sparsity is exploited in network evaluations. To take advantage of sparsity, some accelerator designs explore sparsity encoding and evaluation on CNN accelerators. However, sparsity encoding is just performed on activation or weight and only in inference. It has been shown that activation and weight also have high sparsity levels during training. Hence, sparsity-aware computation should also be considered in training. To further improve performance and energy efficiency, some accelerators evaluate CNNs with limited precision. However, this is limited to the inference since reduced precision sacrifices network accuracy if used in training. In addition, CNN evaluation is usually memory-intensive, especially in training. In this paper, we propose SPRING, a SParsity-aware Reduced-precision Monolithic 3D CNN accelerator for trainING and inference. SPRING supports both CNN training and inference. It uses a binary mask scheme to encode sparsities in activation and weight. It uses the stochastic rounding algorithm to train CNNs with reduced precision without accuracy loss. To alleviate the memory bottleneck in CNN evaluation, especially in training, SPRING uses an efficient monolithic 3D NVM interface to increase memory bandwidth. Compared to GTX 1080 Ti, SPRING achieves 15.6X, 4.2X and 66.0X improvements in performance, power reduction, and energy efficiency, respectively, for CNN training, and 15.5X, 4.5X and 69.1X improvements for inference.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[32]
Compressing DMA engine: Leveraging activation sparsity for training deep neural networks,
M. Rhu, M. O’Connor, N. Chatterjee, J. Pool, Y. Kwon, and S. W. Keckler, “Compressing DMA engine: Leveraging activation sparsity for training deep neural networks,” in Proc. IEEE Int. Symp. High Performance Computer Architecture , Feb. 2018, pp. 78– 91
work page 2018
-
[45]
Deep learning with limited numerical precision,
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P . Narayanan, “Deep learning with limited numerical precision,” in Proc. Int. Conf. Machine Learning, July 2015, pp. 1737–1746. 13
work page 2015
-
[1]
DaDianNao: A machine-learning supercomputer,
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “DaDianNao: A machine-learning supercomputer,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Dec. 2014, pp. 609–622
2014
-
[2]
Scaledeep: A scalable compute architecture for learning and evaluating deep networks,
S. Venkataramani, A. Ranjan, S. Banerjee, D. Das, S. Avancha, A. Jagannathan, A. Durg, D. Nagaraj, B. Kaul, P . Dubey, and A. Raghunathan, “Scaledeep: A scalable compute architecture for learning and evaluating deep networks,” in Proc. Int. Symp. Computer Architecture, June 2017, pp. 13–26
2017
-
[3]
Image classi- fication at supercomputer scale,
C. Ying, S. Kumar, D. Chen, T. Wang, and Y. Cheng, “Image classi- fication at supercomputer scale,” arXiv preprint arXiv:1811.06992, Dec. 2018
arXiv 2018
-
[4]
Ultra-performance Pascal GPU and NVLink interconnect,
D. Foley and J. Danskin, “Ultra-performance Pascal GPU and NVLink interconnect,” IEEE Micro, vol. 37, no. 2, pp. 7–17, Mar. 2017
2017
-
[5]
Volta: Performance and programmability,
J. Choquette, O. Giroux, and D. Foley, “Volta: Performance and programmability,”IEEE Micro, vol. 38, no. 2, pp. 42–52, Mar. 2018
2018
-
[6]
A network- centric hardware/algorithm co-design to accelerate distributed training of deep neural networks,
Y. Li, J. Park, M. Alian, Y. Yuan, Z. Qu, P . Pan, R. Wang, A. Schwing, H. Esmaeilzadeh, and N. S. Kim, “A network- centric hardware/algorithm co-design to accelerate distributed training of deep neural networks,” in Proc. IEEE/ACM Int. Symp. Microarchitecture, Oct. 2018, pp. 175–188
2018
Show all 101 references
-
[7]
Escher: A CNN accelerator with flexible buffering to minimize off-chip transfer,
Y. Shen, M. Ferdman, and P . Milder, “Escher: A CNN accelerator with flexible buffering to minimize off-chip transfer,” in Proc. Int. Symp. Field-Programmable Custom Computing Machines , Apr. 2017, pp. 93–100
2017
-
[8]
Scalpel: Customizing DNN pruning to the under- lying hardware parallelism,
J. Yu, A. Lukefahr, D. Palframan, G. Dasika, R. Das, and S. Mahlke, “Scalpel: Customizing DNN pruning to the under- lying hardware parallelism,” in Proc. Int. Symp. Computer Archi- tecture, June 2017, pp. 548–560
2017
-
[9]
Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks,
H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, V . Chandra, and H. Esmaeilzadeh, “Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks,” in Proc. Int. Symp. Computer Architecture, June 2018, pp. 764–775
2018
-
[10]
MAERI: Enabling flexi- ble dataflow mapping over DNN accelerators via reconfigurable interconnects,
H. Kwon, A. Samajdar, and T. Krishna, “MAERI: Enabling flexi- ble dataflow mapping over DNN accelerators via reconfigurable interconnects,” in Proc. Int. Conf. Architectural Support Program- ming Languages Operating Syst., Mar. 2018, pp. 461–475
2018
-
[11]
A reconfigurable fabric for accelerating large-scale datacenter services,
A. Putnam, A. M. Caulfield, E. S. Chung, D. Chiou, K. Con- stantinides, J. Demme, H. Esmaeilzadeh, J. Fowers, G. P . Gopal, J. Gray, M. Haselman, S. Hauck, S. Heil, A. Hormati, J.-Y. Kim, S. Lanka, J. Larus, E. Peterson, S. Pope, A. Smith, J. Thong, P . Y. Xiao, and D. Burger, ...
2014
-
[12]
FPGA based implementation of deep neural networks using on-chip memory only,
J. Park and W. Sung, “FPGA based implementation of deep neural networks using on-chip memory only,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Processing, Mar. 2016, pp. 1011–1015. 12
2016
-
[13]
Overcoming resource underutilization in spatial CNN accelerators,
Y. Shen, M. Ferdman, and P . Milder, “Overcoming resource underutilization in spatial CNN accelerators,” in Proc. Int. Conf. Field Programmable Logic Applications, Aug. 2016, pp. 1–4
2016
-
[14]
Going deeper with embedded FPGA platform for convolutional neural network,
J. Qiu, J. Wang, S. Yao, K. Guo, B. Li, E. Zhou, J. Yu, T. Tang, N. Xu, S. Song, Y. Wang, and H. Yang, “Going deeper with embedded FPGA platform for convolutional neural network,” in Proc. ACM/SIGDA Int. Symp. Field-Programmable Gate Arrays , 2016, pp. 26–35
2016
-
[15]
ShiDianNao: Shifting vision processing closer to the sensor,
Z. Du, R. Fasthuber, T. Chen, P . Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam, “ShiDianNao: Shifting vision processing closer to the sensor,” in Proc. ACM/IEEE Int. Symp. Computer Architecture, June 2015, pp. 92–104
2015
-
[16]
Chain-NN: An energy-efficient 1D chain architecture for accelerating deep convolutional neural networks,
S. Wang, D. Zhou, X. Han, and T. Yoshimura, “Chain-NN: An energy-efficient 1D chain architecture for accelerating deep convolutional neural networks,” in Proc. Design, Automation Test Europe Conf. Exhibition, Mar. 2017, pp. 1032–1037
2017
-
[17]
CirCNN: Accelerating and compressing deep neural networks using block-circulant weight matrices,
C. Ding, S. Liao, Y. Wang, Z. Li, N. Liu, Y. Zhuo, C. Wang, X. Qian, Y. Bai, G. Yuan, X. Ma, Y. Zhang, J. Tang, Q. Qiu, X. Lin, and B. Yuan, “CirCNN: Accelerating and compressing deep neural networks using block-circulant weight matrices,” in Proc. IEEE/ACM Int. Symp. Microarc...
2017
-
[18]
TETRIS: Scalable and efficient neural network acceleration with 3D mem- ory,
M. Gao, J. Pu, X. Yang, M. Horowitz, and C. Kozyrakis, “TETRIS: Scalable and efficient neural network acceleration with 3D mem- ory,” in Proc. Int. Conf. Architectural Support Programming Lan- guages Operating Syst., 2017, pp. 751–764
2017
-
[19]
In-datacenter performance analysis of a tensor processing unit,
N. P . Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P .-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V . Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D...
2017
-
[20]
Fused-layer CNN accelerators,
M. Alwani, H. Chen, M. Ferdman, and P . Milder, “Fused-layer CNN accelerators,” in Proc. IEEE/ACM Int. Symp. Microarchitec- ture, Oct. 2016, pp. 1–12
2016
-
[21]
Accelerating CNN algorithm with fine-grained dataflow architectures,
T. Xiang, Y. Feng, X. Ye, X. Tan, W. Li, Y. Zhu, M. Wu, H. Zhang, and D. Fan, “Accelerating CNN algorithm with fine-grained dataflow architectures,” in Proc. IEEE Int. Conf. High Performance Computing Communications; IEEE Int. Conf. Smart City; IEEE Int. Conf. Data Science Syst....
2018
-
[22]
DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,” in Proc. Int. Conf. Architectural Support Programming Languages Operating Syst. , Mar. 2014, pp. 269–284
2014
-
[23]
Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks,
W. Lu, G. Yan, J. Li, S. Gong, Y. Han, and X. Li, “Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks,” in Proc. IEEE Int. Symp. High Performance Computer Architecture, Feb. 2017, pp. 553–564
2017
-
[24]
Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,
Y. Chen, J. Emer, and V . Sze, “Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,” in Proc. ACM/IEEE Int. Symp. Computer Architecture , June 2016, pp. 367–379
2016
-
[25]
EIE: Efficient inference engine on compressed deep neural network,
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “EIE: Efficient inference engine on compressed deep neural network,” in Proc. Int. Symp. Computer Architecture , June 2016, pp. 243–254
2016
-
[26]
SparseNN: An energy- efficient neural network accelerator exploiting input and output sparsity,
J. Zhu, J. Jiang, X. Chen, and C. Tsui, “SparseNN: An energy- efficient neural network accelerator exploiting input and output sparsity,” in Proc. Design, Automation Test Europe Conf. Exhibition , Mar. 2018, pp. 241–244
2018
-
[27]
SCNN: An accelerator for compressed-sparse convolutional neural net- works,
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “SCNN: An accelerator for compressed-sparse convolutional neural net- works,” in Proc. ACM/IEEE Int. Symp. Computer Architecture, June 2017, pp. 27–40
2017
-
[28]
UCNN: Exploiting computational reuse in deep neural networks via weight repetition,
K. Hegde, J. Yu, R. Agrawal, M. Yan, M. Pellauer, and C. W. Fletcher, “UCNN: Exploiting computational reuse in deep neural networks via weight repetition,” in Proc. Int. Symp. Computer Architecture, June 2018, pp. 674–687
2018
-
[29]
Cnvlutin: Ineffectual-neuron-free deep neural network computing,
J. Albericio, P . Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” in Proc. ACM/IEEE Int. Symp. Computer Architecture, June 2016, pp. 1–13
2016
-
[30]
Cambricon-X: An accelerator for sparse neural networks,
S. Zhang, Z. Du, L. Zhang, H. Lan, S. Liu, L. Li, Q. Guo, T. Chen, and Y. Chen, “Cambricon-X: An accelerator for sparse neural networks,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Oct. 2016, pp. 1–12
2016
-
[31]
Imagenet classi- fication with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classi- fication with deep convolutional neural networks,” in Proc. Int. Conf. Neural Information Processing Syst., Dec. 2012, pp. 1097–1105
2012
-
[33]
(2014) Hybrid Mem- ory Cube specification 2.1
HMC Consortium. (2014) Hybrid Mem- ory Cube specification 2.1. [Online]. Avail- able: http://hybridmemorycube .org/files/SiteDownloads/ HMC-30G-VSR HMCC Specification Rev2.1 20151105.pdf
2014
-
[34]
(2016) Samsung begins mass producing world’s fastest DRAM - based on newest High Bandwidth Memory (HBM) interface
Samsung Newsroom. (2016) Samsung begins mass producing world’s fastest DRAM - based on newest High Bandwidth Memory (HBM) interface. [Online]. Available: https://www.samsung.com/semiconductor/insights/news- events/samsung-begins-mass-producing-worlds-fastest-dram- based-on-new...
2016
-
[35]
Application- transparent near-memory processing architecture with memory channel network,
M. Alian, S. W. Min, H. Asgharimoghaddam, A. Dhar, D. K. Wang, T. Roewer, A. McPadden, O. O’Halloran, D. Chen, J. Xiong, D. Kim, W. Hwu, and N. S. Kim, “Application- transparent near-memory processing architecture with memory channel network,” in Proc. IEEE/ACM Int. Symp. Micr...
2018
-
[36]
Chameleon: Versatile and practical near-DRAM acceleration architecture for large memory systems,
H. Asghari-Moghaddam, Y. H. Son, J. H. Ahn, and N. S. Kim, “Chameleon: Versatile and practical near-DRAM acceleration architecture for large memory systems,” in Proc. IEEE/ACM Int. Symp. Microarchitecture, Oct. 2016, pp. 1–13
2016
-
[37]
PRIME: A novel processing-in-memory architecture for neural network computation in ReRAM-based main memory,
P . Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “PRIME: A novel processing-in-memory architecture for neural network computation in ReRAM-based main memory,” in Proc. Int. Symp. Computer Architecture, June 2016, pp. 27–39
2016
-
[38]
Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory,
D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopadhyay, “Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory,” SIGARCH Computer Architecture News, vol. 44, no. 3, pp. 380–392, June 2016
2016
-
[39]
ISAAC: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars,
A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P . Strachan, M. Hu, R. S. Williams, and V . Srikumar, “ISAAC: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars,” in Proc. Int. Symp. Computer Architecture, June 2016, pp. 14–26
2016
-
[40]
TIME: A training-in-memory architecture for memristor- based deep neural networks,
M. Cheng, L. Xia, Z. Zhu, Y. Cai, Y. Xie, Y. Wang, and H. Yang, “TIME: A training-in-memory architecture for memristor- based deep neural networks,” in Proc. ACM/EDAC/IEEE Design Automation Conf., June 2017, pp. 1–6
2017
-
[41]
NNPIM: A pro- cessing in-memory architecture for neural network acceleration,
S. Gupta, M. Imani, H. Kaur, and T. S. Rosing, “NNPIM: A pro- cessing in-memory architecture for neural network acceleration,” IEEE Trans. Computers, vol. 68, no. 9, pp. 1325–1337, Sep. 2019
2019
-
[42]
TensorDIMM: A practical near- memory processing architecture for embeddings and tensor op- erations in deep learning,
Y. Kwon, Y. Lee, and M. Rhu, “TensorDIMM: A practical near- memory processing architecture for embeddings and tensor op- erations in deep learning,” in Proc. IEEE/ACM Int. Symp. Microar- chitecture, Oct. 2019, pp. 740–753
2019
-
[43]
Deep learning inference in facebook data centers: Characterization, performance optimizations and hardware implications,
J. Park, M. Naumov, P . Basu, S. Deng, A. Kalaiah, D. Khu- dia, J. Law, P . Malani, A. Malevich, S. Nadathur et al. , “Deep learning inference in facebook data centers: Characterization, performance optimizations and hardware implications,” arXiv preprint arXiv:1811.09886, 2018
2018 arXiv
-
[44]
FloatPIM: In- memory acceleration of deep neural network training with high precision,
M. Imani, S. Gupta, Y. Kim, and T. Rosing, “FloatPIM: In- memory acceleration of deep neural network training with high precision,” in Proc. Int. Symp. Computer Architecture , June 2019, pp. 802–815
2019
-
[46]
Modeling the re- source requirements of convolutional neural networks on mobile devices,
Z. Lu, S. Rallapalli, K. Chan, and T. La Porta, “Modeling the re- source requirements of convolutional neural networks on mobile devices,” in Proc. ACM Int. Conf. Multimedia, 2017, pp. 1663–1671
2017
-
[47]
Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks,
C. Zhang, G. Sun, Z. Fang, P . Zhou, P . Pan, and J. Cong, “Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks,” in Proc. IEEE/ACM Int. Conf. Computer-Aided Design, Nov. 2016, pp. 1–8
2016
-
[48]
Deep compression: Compress- ing deep neural networks with pruning, trained quantization and Huffman coding,
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compress- ing deep neural networks with pruning, trained quantization and Huffman coding,” arXiv preprint arXiv:1510.00149, 2015
2015 arXiv
-
[49]
Learning both weights and connections for efficient neural networks,
S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both weights and connections for efficient neural networks,” in Proc. Int. Conf. Neural Information Processing Syst., 2015, pp. 1135–1143
2015
-
[50]
Automatic performance tuning of sparse matrix kernels,
R. W. Vuduc, “Automatic performance tuning of sparse matrix kernels,” Ph.D. dissertation, University of California, Berkeley, 2003
2003
-
[51]
Stitch-X: An accelerator architecture for exploiting unstructured sparsity in deep neural networks,
C. Lee, Y. Shao, J.-F. Zhang, A. Parashar, J. Emer, S. Keckler, and Z. Zhang, “Stitch-X: An accelerator architecture for exploiting unstructured sparsity in deep neural networks,” in Proc. SysML Conference, 2018
2018
-
[52]
Dynamic warp formation and scheduling for efficient GPU control flow,
W. W. L. Fung, I. Sham, G. Yuan, and T. M. Aamodt, “Dynamic warp formation and scheduling for efficient GPU control flow,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Dec. 2007, pp. 407–420
2007
-
[53]
Thread block compaction for efficient SIMT control flow,
W. W. L. Fung and T. M. Aamodt, “Thread block compaction for efficient SIMT control flow,” in Proc. IEEE Int. Symp. High Performance Computer Architecture, Feb. 2011, pp. 25–36
2011
-
[54]
Improving GPU performance via large warps and two-level warp scheduling,
V . Narasiman, M. Shebanow, C. J. Lee, R. Miftakhutdinov, O. Mutlu, and Y. N. Patt, “Improving GPU performance via large warps and two-level warp scheduling,” in Proc. IEEE/ACM Int. Symp. Microarchitecture, Dec. 2011, pp. 308–317
2011
-
[55]
Convergence and scalarization for data-parallel architectures,
Y. Lee, R. Krashinsky, V . Grover, S. W. Keckler, and K. Asanovi ´c, “Convergence and scalarization for data-parallel architectures,” in Proc. IEEE/ACM Int. Symp. Code Generation Optimization , Feb. 2013, pp. 1–11
2013
-
[56]
Cambricon-S: Addressing irreg- ularity in sparse neural networks through a cooperative soft- ware/hardware approach,
X. Zhou, Z. Du, Q. Guo, S. Liu, C. Liu, C. Wang, X. Zhou, L. Li, T. Chen, and Y. Chen, “Cambricon-S: Addressing irreg- ularity in sparse neural networks through a cooperative soft- ware/hardware approach,” in Proc. IEEE/ACM Int. Symp. Mi- croarchitecture, Oct. 2018, pp. 15–28
2018
-
[57]
Stripes: Bit-serial deep neural network comput- ing,
P . Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, “Stripes: Bit-serial deep neural network comput- ing,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Oct. 2016, pp. 1–12
2016
-
[58]
Bit-pragmatic deep neural network comput- ing,
J. Albericio, A. Delm ´as, P . Judd, S. Sharify, G. O’Leary, R. Genov, and A. Moshovos, “Bit-pragmatic deep neural network comput- ing,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Oct. 2017, pp. 382–394
2017
-
[59]
Scalable distributed DNN training using commodity GPU cloud computing,
N. Strom, “Scalable distributed DNN training using commodity GPU cloud computing,” in Proc. Conf. Int. Speech Communication Association, Sep. 2015
2015
-
[60]
Distributed train- ing large-scale deep architectures,
S.-X. Zou, C.-Y. Chen, J.-L. Wu, C.-N. Chou, C.-C. Tsao, K.-C. Tung, T.-W. Lin, C.-L. Sung, and E. Y. Chang, “Distributed train- ing large-scale deep architectures,” in Proc. Int. Conf. Advanced Data Mining Applications, Oct. 2017, pp. 18–32
2017
-
[61]
Distributed training strategies for a computer vision deep learning algorithm on a distributed GPU cluster,
V . Campos, F. Sastre, M. Yag ¨ues, M. Bellver, X. Gir ´o-i Nieto, and J. Torres, “Distributed training strategies for a computer vision deep learning algorithm on a distributed GPU cluster,” Procedia Computer Science, vol. 108, pp. 315–324, May 2017
2017
-
[62]
Highly scalable deep learning training system with mixed- precision: Training Imagenet in four minutes,
X. Jia, S. Song, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu, T. Chen, G. Hu, S. Shi, and X. Chu, “Highly scalable deep learning training system with mixed- precision: Training Imagenet in four minutes,” arXiv preprint arXiv:1807.11205, July 2018
2018 arXiv
-
[63]
Mixed precision training,
P . Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Gar- cia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in Proc. Int. Conf. Learning Representations, May 2018
2018
-
[64]
The vanishing gradient problem during learning recurrent neural nets and problem solutions,
S. Hochreiter, “The vanishing gradient problem during learning recurrent neural nets and problem solutions,” Int. Journal Uncer- tainty, Fuzziness Knowledge-Based Syst. , vol. 6, no. 2, pp. 107–116, Apr. 1998
1998
-
[65]
Dynamically scaled fixed point arithmetic,
D. Williamson, “Dynamically scaled fixed point arithmetic,” in Proc. IEEE Pacific Rim Conf. Communications, Computers Signal Processing, May 1991, pp. 315–318 vol.1
1991
-
[66]
Training deep neural networks with low precision multiplications,
M. Courbariaux, Y. Bengio, and J.-P . David, “Training deep neural networks with low precision multiplications,” in Proc. Int. Conf. Learning Representations, May 2015
2015
-
[67]
Mixed precision training of convo- lutional neural networks using integer operations,
D. Das, N. Mellempudi, D. Mudigere, D. Kalamkar, S. Avancha, K. Banerjee, S. Sridharan, K. Vaidyanathan, B. Kaul, E. Geor- ganas, A. Heinecke, P . Dubey, J. Corbal, N. Shustrov, R. Dubtsov, E. Fomenko, and V . Pirogov, “Mixed precision training of convo- lutional neural networ...
2018
-
[68]
(2015) High Bandwidth Memory
AMD. (2015) High Bandwidth Memory. [Online]. Available: https://www.amd.com/en/technologies/hbm
2015
-
[69]
Energy-efficient monolithic three- dimensional on-chip memory architectures,
Y. Yu and N. K. Jha, “Energy-efficient monolithic three- dimensional on-chip memory architectures,” IEEE Trans. Nan- otechnology, vol. 17, no. 4, pp. 620–633, July 2018
2018
-
[70]
A monolithic 3D hybrid architecture for energy-efficient computation,
——, “A monolithic 3D hybrid architecture for energy-efficient computation,” IEEE Trans. Multi-Scale Computing Syst. , vol. 4, no. 4, pp. 533–547, Oct. 2018
2018
-
[71]
A 5ns fast write multi-level non-volatile 1 K bits RRAM memory with advance write scheme,
S. Sheu, P . Chiang, W. Lin, H. Lee, P . Chen, Y. Chen, T. Wu, F. T. Chen, K. Su, M. Kao, K. Cheng, and M. Tsai, “A 5ns fast write multi-level non-volatile 1 K bits RRAM memory with advance write scheme,” in Proc. Symp VLSI Circuits, June 2009, pp. 82–83
2009
-
[72]
(2013, May) Technology roadmap of DRAM for three major manufacturers: Samsung, SK-Hynix and Micron
Techinsights. (2013, May) Technology roadmap of DRAM for three major manufacturers: Samsung, SK-Hynix and Micron. [Online]. Available: https://www .techinsights.com/ uploadedFiles/Public Website/Content - Primary/Marketing/ 2013/DRAM Roadmap/Report/TechInsights-DRAM- ROADMAP-0...
2013
-
[73]
(2016) The crossbar RRAM advantage
Crossbar. (2016) The crossbar RRAM advantage. [On- line]. Available: http://www .crossbar-inc.com/technology/ rram-advantages/
2016
-
[74]
3D sequential integration opportunities and technology optimization,
P . Batude, B. Sklenard, C. Fenouillet-Beranger, B. Previtali, C. Tabone, O. Rozeau, O. Billoint, O. Turkyilmaz, H. Sarhan, S. Thuries, G. Cibrario, L. Brunet, F. Deprat, J. Michallet, F. Cler- midy, and M. Vinet, “3D sequential integration opportunities and technology optimiz...
2014
-
[75]
Batch normalization: Accelerating deep network training by reducing internal covariate shift,
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proc. Int. Conf. Machine Learning, July 2015, pp. 448–456
2015
-
[76]
(2018) Design Compiler
Synopsys. (2018) Design Compiler. [Online]. Available: https://www .synopsys.com/support/training/rtl- synthesis/design-compiler-rtl-synthesis .html
2018
-
[77]
Hybrid monolithic 3-D IC floorplanner,
A. Guler and N. K. Jha, “Hybrid monolithic 3-D IC floorplanner,” IEEE Trans. Very Large Scale Integration Syst. , vol. 26, no. 10, pp. 1868–1880, Oct. 2018
2018
-
[78]
Capo: Robust and scalable open-source min- cut floorplacer,
J. A. Roy, D. A. Papa, S. N. Adya, H. H. Chan, A. N. Ng, J. F. Lu, and I. L. Markov, “Capo: Robust and scalable open-source min- cut floorplacer,” in Proc. Int. Symp. Physical design , Apr. 2005, pp. 224–226
2005
-
[79]
FinCACTI: Ar- chitectural analysis and modeling of caches with deeply-scaled FinFET devices,
A. Shafaei, Y. Wang, X. Lin, and M. Pedram, “FinCACTI: Ar- chitectural analysis and modeling of caches with deeply-scaled FinFET devices,” in Proc. IEEE Computer Society Annual Symp. VLSI, July 2014, pp. 290–295
2014
-
[80]
CACTI 6.0: A tool to model large caches,
N. Muralimanohar, R. Balasubramonian, and N. P . Jouppi, “CACTI 6.0: A tool to model large caches,” HP Laboratories, pp. 22–31, 2009
2009
-
[81]
NVSim: A circuit-level performance, energy, and area model for emerging nonvolatile memory,
X. Dong, C. Xu, Y. Xie, and N. P . Jouppi, “NVSim: A circuit-level performance, energy, and area model for emerging nonvolatile memory,” IEEE Trans. Comput.-Aided Design Integr. Circuits Syst. , vol. 31, no. 7, pp. 994–1007, July 2012
2012
-
[82]
NVMain 2.0: A user-friendly memory simulator to model (non-)volatile memory systems,
M. Poremba, T. Zhang, and Y. Xie, “NVMain 2.0: A user-friendly memory simulator to model (non-)volatile memory systems,” IEEE Comput. Archit. Lett., vol. 14, no. 2, pp. 140–143, July 2015
2015
-
[83]
Ten- sorflow: A system for large-scale machine learning,
M. Abadi, P . Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Lev- enberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P . Tucker, V . Vasudevan, P . Warden, M. Wicke, Y. Yu, and X. Zheng, “Ten- sorflow: A system for larg...
2016
-
[84]
Software-defined design space exploration for an efficient AI accelerator architec- ture,
Y. Yu, Y. Li, S. Che, N. K. Jha, and W. Zhang, “Software-defined design space exploration for an efficient AI accelerator architec- ture,” arXiv preprint arXiv:1903.07676, 2019
1903 arXiv
-
[85]
Inception- v4, Inception-Resnet and the impact of residual connections on learning,
C. Szegedy, S. Ioffe, V . Vanhoucke, and A. A. Alemi, “Inception- v4, Inception-Resnet and the impact of residual connections on learning,” in Proc. AAAI Conf. Artificial Intelligence , Feb. 2017
2017
-
[86]
Rethinking the Inception architecture for computer vision,
C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception architecture for computer vision,” in 14 Proc. IEEE Conf. Computer Vision Pattern Recognition , June 2016, pp. 2818–2826
2016
-
[87]
MobileNetV2: Inverted residuals and linear bottlenecks,
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in Proc. IEEE Conf. Computer Vision Pattern Recognition , June 2018, pp. 4510–4520
2018
-
[88]
Learning transfer- able architectures for scalable image recognition,
B. Zoph, V . Vasudevan, J. Shlens, and Q. V . Le, “Learning transfer- able architectures for scalable image recognition,” in Proc. IEEE Conf. Computer Vision Pattern Recognition , June 2018, pp. 8697– 8710
2018
-
[89]
Progressive neural architecture search,
C. Liu, B. Zoph, M. Neumann, J. Shlens, W. Hua, L.-J. Li, L. Fei- Fei, A. Yuille, J. Huang, and K. Murphy, “Progressive neural architecture search,” in Proc. European Conf. Computer Vision, Sep. 2018, pp. 19–34
2018
-
[90]
Identity mappings in deep residual networks,
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in Proc. European Conf. Computer Vision , Oct. 2016, pp. 630–645
2016
-
[91]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014
2014 arXiv
-
[92]
ImageNet Large Scale Visual Recognition Challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” Int. Journal Computer Vision, vol. 115, no. 3, pp. 211–252, 2015
2015
-
[93]
Silberman and S
N. Silberman and S. Guadarrama. (2016) Tensorflow-slim image classification model library. [Online]. Available: https: //github.com/tensorflow/models/tree/master/research/slim
2016
-
[94]
Physical design solutions to tackle FEOL/BEOL degradation in gate-level monolithic 3D ICs,
B. W. Ku, P . Debacker, D. Milojevic, P . Raghavan, D. Verkest, A. Thean, and S. K. Lim, “Physical design solutions to tackle FEOL/BEOL degradation in gate-level monolithic 3D ICs,” in Proc. ACM Int. Symp. Low Power Electron. Design , 2016, pp. 76–81
2016
-
[95]
How to cope with slow transistors in the top-tier of monolithic 3D ICs: Design studies and CAD solutions,
S. K. Samal, D. Nayak, M. lchihashi, S. Banna, and S. K. Lim, “How to cope with slow transistors in the top-tier of monolithic 3D ICs: Design studies and CAD solutions,” in Proc. ACM Int. Symp. Low Power Electron. Design, 2016, pp. 320–325
2016
-
[96]
Power-performance study of block-level monolithic 3D-ICs considering inter-tier performance variations,
S. Panth, K. Samadi, Y. Du, and S. K. Lim, “Power-performance study of block-level monolithic 3D-ICs considering inter-tier performance variations,” in Proc. ACM Annual Design Auto. Conf., 2014
2014
-
[97]
Ultra-high density 3D SRAM cell designs for monolithic 3D integration,
C. Liu and S. K. Lim, “Ultra-high density 3D SRAM cell designs for monolithic 3D integration,” in Proc. IEEE Int. Interconnect Technol. Conf., June 2012, pp. 1–3
2012
-
[98]
Compact 6T SRAM cell with robust read/write stabilizing de- sign in 45nm monolithic 3D IC technology,
O. Thomas, M. Vinet, O. Rozeau, P . Batude, and A. Valentian, “Compact 6T SRAM cell with robust read/write stabilizing de- sign in 45nm monolithic 3D IC technology,” in Proc. IEEE Int. Conf. IC Design Technology, May 2009, pp. 195–198
2009
-
[99]
Intermediate BEOL process influence on power and performance for 3DVLSI,
H. Sarhan, S. Thuries, O. Billoint, F. Deprat, A. A. D. Sousa, P . Batude, C. Fenouillet-Beranger, and F. Clermidy, “Intermediate BEOL process influence on power and performance for 3DVLSI,” in Proc. IEEE Int. 3D Syst. Integration Conf., Aug. 2015, pp. TS1.3.1– TS1.3.5
2015
-
[100]
Supporting compressed-sparse activations and weights on SIMD-like accelerator for sparse convolutional neural networks,
C. Lin and B. Lai, “Supporting compressed-sparse activations and weights on SIMD-like accelerator for sparse convolutional neural networks,” in Proc. Asia South Pacific Design Automation Conf., Jan. 2018, pp. 105–110
2018
-
[101]
Eager Pruning: Algo- rithm and architecture support for fast training of deep neural networks,
J. Zhang, X. Chen, M. Song, and T. Li, “Eager Pruning: Algo- rithm and architecture support for fast training of deep neural networks,” in Proc. Int. Symp. Computer Architecture , June 2019, pp. 292–303. Ye Yu received the B.Eng. degree in Electronic and Computer Engineering f...
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.