Pith. sign in

REVIEW 4 major objections 5 minor 101 references

SPRING: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read SPRING claims that a binary-mask sparsity scheme plus stochastic-rounded fixed-point arithmetic can make CNN training and inference far faster and more energy-efficient, reporting 15.6x training speedup over a GTX 1080 Ti.

desk verdict A plausible new architecture combining sparsity, reduced-precision training, and monolithic 3D memory, but the headline speedups rest on an unjustified 50% uniform sparsity assumption that conflates activation with weight sparsity, making the numbers conditional estimates rather than verified results. read the letter →

arxiv 1909.00557 v2 pith:QWL3WDSM submitted 2019-09-02 cs.AR

classification cs.AR
keywords CNNacceleratorsparsity-awarecomputingbinarymasksparsityencodingstochasticroundingreduced-precisiontrainingmonolithic3DintegrationRRAMmemoryinterfaceandinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

SPRING is a proposed CNN accelerator that targets both training and inference on the same hardware, combining three ideas: binary-mask sparsity encoding to skip zero-valued activations and weights, stochastic-rounding fixed-point arithmetic so reduced-precision training does not lose accuracy, and a monolithic 3D RRAM interface to supply enough memory bandwidth for training. The paper's central claim is that this combination lets a 151 mm² accelerator beat an Nvidia GeForce GTX 1080 Ti by 15.6x in training speed, 4.2x in power reduction, and 66.0x in energy efficiency, with similar inference gains of 15.5x, 4.5x, and 69.1x. If true, it would mean CNNs can be trained and deployed on a single low-power, high-bandwidth chip rather than relying on GPU clusters for training and separate accelerators for inference.

What carries the argument

The load-bearing mechanism is the binary-mask sparsity scheme. A mask bit accompanies each activation or weight element; zero-valued elements are removed for storage, and before each multiply-accumulate the activation mask and weight mask are ANDed to identify common nonzero positions. A sequential scanning filter drops 'dangling' nonzeros, those where a nonzero activation aligns with a zero weight or vice versa, and a zero-collapsing shifter presents dense zero-free vectors to the MAC lanes. A post-compute sparsity module recompresses outputs after the activation function. Alongside this, each MAC lane embeds stochastic rounding, which rounds fixed-point results up or down with probability proportional to the discarded fraction, so training can run at 20-bit fixed-point precision without, per the paper's argument, extra convergence epochs. The third leg is a monolithic 3D RRAM interface with 1 KB row-wide buses and decoupled read/write interconnects to keep the PEs fed during memory-intensive training.

What would settle it

Measure per-layer activation and weight sparsity during training and inference on all seven CNNs, then rerun the paper's cycle-accurate simulator with those measured densities; if the average sparsity is below 50%, the claimed 15.6x and 15.5x speedups do not follow.

Watch

Extended reading notes

Core claim

The paper argues that sparsity and reduced precision are not just inference techniques: they can be carried through the entire training process. SPRING compresses activations and weights by dropping zero elements and keeping a binary mask that records where the zeros were. Before each multiply-accumulate, the activation mask and weight mask are ANDed so that only positions where both operands are nonzero are sent to the MAC lanes, and a post-compute module compresses newly zeroed outputs after the activation function. To make low-precision training work, each MAC lane embeds stochastic rounding, which rounds a fixed-point result up or down with probability proportional to the discarded fraction, allowing 4 integer bits and 16 fractional bits to train without, the paper argues, extra convergence epochs. A monolithic 3D nonvolatile-memory interface with row-wide 1 KB buses and decoupled read/write paths supplies the bandwidth that training demands. On seven ImageNet CNNs, the paper reports geometric-mean speedups of 15.6x for training and 15.5x for inference over a GTX 1080 Ti, with larger gains on lightweight networks such as MobileNet V2 and smaller gains on memory-bound networks such as VGG-19.

Load-bearing premise

Every reported speedup and energy gain assumes the seven CNNs are uniformly 50% sparse; since SPRING's gains come from skipping zero entries, real sparsity below 50% would shrink all the headline ratios roughly in proportion.

Editorial extensions

If this is right

  • Sparsity-aware training removes the usual wall between the training phase and the deployment phase, so a single chip could handle both without sacrificing the speedups that zero-skipping provides.
  • Stochastic rounding makes low-precision fixed-point arithmetic viable for backpropagation, which would let future accelerators use simpler, more energy-efficient MAC units than the FP32 units GPUs rely on.
  • The reported gains are larger on lightweight CNNs, showing that the monolithic 3D memory interface relieves the bandwidth bottleneck most effectively when the working set fits on chip.
  • If the reported batch-level numbers hold, SPRING would reduce training energy by roughly two orders of magnitude relative to a GTX 1080 Ti, making on-device and edge training far more practical.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The headline numbers lean on an assumed uniform 50% sparsity applied to all seven CNNs, but real networks have per-layer and per-phase sparsity that varies; a measured sparsity profile would likely spread the speedups from roughly 5x to over 50x instead of a single 15x average.
  • Because stochastic rounding injects pseudo-random noise into every rounding decision, the convergence guarantee depends on the quality of the random source and the exact integer/fraction bit split; sweeping IL and FL bit widths during training would map the accuracy-versus-precision tradeoff directly.
  • The pipelined sequential mask filter preserves batch throughput but adds per-image latency, so for single-image edge inference the architecture may be less attractive than the batch-level speedup numbers suggest.
  • The paper's architectural ideas are separable: a designer could adopt the binary-mask sparsity scheme without the 3D RRAM stack, or use stochastic rounding alone to improve an existing fixed-point training pipeline, and still capture part of the claimed benefit.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents SPRING, a sparsity-aware reduced-precision monolithic 3D CNN accelerator architecture that targets both training and inference. The design uses binary masks to compress zero-valued activations and weights, a pre-compute sparsity module to skip ineffectual MACs, stochastic rounding for fixed-point training, and a monolithic 3D RRAM memory interface to increase bandwidth. The authors evaluate SPRING on seven ImageNet CNNs with a custom cycle-accurate simulator, and report large geometric-mean improvements relative to a GTX 1080 Ti: 15.6x performance, 4.2x power, and 66.0x energy for training, and 15.5x, 4.5x, and 69.1x for inference. The paper includes RTL synthesis of the accelerator and uses standard tools such as Design Compiler, NVSim, NVMain, and FinCACTI.

Significance. If the reported results are validated, SPRING would be a meaningful contribution to the sparse CNN accelerator literature, especially because it addresses training as well as inference and combines sparsity, reduced precision, and monolithic 3D NVM in a single architecture. The use of a custom cycle-accurate simulator, RTL synthesis, and standard circuit-level memory tools is a strength. However, the central quantitative claims rest on an unverified blanket sparsity assumption and an undocumented GPU baseline, so the specific speedup and energy numbers are not yet established. The architecture itself, with its binary-mask datapath and stochastic-rounding MAC lanes, is plausible and worth further study.

major comments (4)
  1. [Section 4 (Simulation Methodology)] The blanket assumption "The sparsity of the CNNs are assumed to be 50%" is load-bearing and unsupported. The seven TensorFlow-Slim CNNs are dense, unpruned models, so their exact-zero weight sparsity is negligible; the cited reference [32] reports activation sparsity during training for AlexNet, VGG, and Inception, not weight sparsity. Since SPRING's pre-compute sparsity module skips MACs only when an activation or a weight is exactly zero, the effective fraction of skipped MACs for dense weights is roughly the activation sparsity, not the ~75% that would result from independent 50% activation and 50% weight sparsity. The same issue applies to the memory-compression benefit: with dense weights, weight masks are almost all ones and weight data cannot be compressed. Because skipped-MAC count and compressed-memory traffic feed directly into the cycle-accurate simulator, the reported 15.6x/15.5x speedups and 66.0x/69.1x energy improvements are systematically inflated unless the authors actually prune each network to 50% weight sparsity. Please provide per-layer activation and weight sparsity measurements for each benchmark, or train/prune the models so that the sparsity input to the simulator is realized.
  2. [Section 4 and Section 5] The GTX 1080 Ti baseline is not described with enough detail to make the comparison reproducible. The paper states only the GPU's peak TFLOPS, memory bandwidth, and die size; it does not explain how execution time, power, or energy are estimated (e.g., which profiler, which cycle-level or analytical model, whether board power or chip power is used, and whether the GPU model accounts for sparsity). Without this information, the claimed GPU-normalized improvements cannot be audited, and the 66.0x energy-efficiency ratio in particular cannot be interpreted. Please add a complete description of the baseline estimation methodology, including the source of all power and timing numbers.
  3. [Section 6 (Discussions and Limitations) and Section 5] The training results are batch-level simulations, not end-to-end training. The paper acknowledges this in Section 6 and relies on the reference [45] to argue that stochastic rounding with 16 FL bits converges in a similar number of epochs as FP32 training. However, no convergence experiment is run on any of the seven CNNs, so the headline "training speedup" is actually a per-batch throughput improvement with an unverified convergence assumption. Please either run a convergence experiment (even on a reduced dataset or a smaller proxy network) or reword the claims to state explicitly that the results are per-batch, not wall-clock training time to target accuracy.
  4. [Section 3.1 (Sparsity-aware acceleration), Algorithm 1] The cost of the pre-compute sparsity module is not modeled in sufficient detail to guarantee that the MAC lanes are not stalled. The paper states that the sequential scanning and filtering scheme is pipelined and that throughput is unaffected, but it provides no cycle counts, hardware area, or energy numbers for the mask-generation, dangling-data filtering, and zero-collapsing shifter stages. Since the performance advantages of SPRING come precisely from this module, the cycle-accurate simulator should incorporate its latency and throughput (including the variable-length compression behavior) rather than assume it runs in shadow of the MAC lanes. Please provide implementation data and simulator modeling details for the pre-compute sparsity module.
minor comments (5)
  1. [Section 2.2] The sentence "the sparsity levels of CNN weights typically range from 20% to 80% [48], [49]" should clarify that these numbers are for pruned/compressed networks, not dense networks, to avoid misleading readers about natural weight sparsity.
  2. [Figures 15 and 16] The GTX 1080 Ti bars are invisible because the energy-efficiency scale is dominated by SPRING's values; consider a log scale or a table with exact normalized numbers.
  3. [Equation (4)] The probability expressions should be parenthesized: "with probability (floor(x) + epsilon - x)/epsilon" and "with probability (x - floor(x))/epsilon" to avoid ambiguity.
  4. [Table 1] The parameter "tBU RST" should be "tBURST" (write burst time).
  5. [Section 6] The "at most 5%" mask-overhead claim holds relative to uncompressed data but not relative to the compressed data stream; clarify the denominator to avoid confusion.

Circularity Check

0 steps flagged · score 2.0 of 10

No definitional or constructional circularity; headline gains are conditional on an externally-sourced 50% sparsity assumption and author-supplied simulation parameters, but do not reduce by construction to the paper's inputs.

full rationale

The paper's derivation chain is not circular. The central claims are accelerator speedups computed by a cycle-accurate simulator from the architecture in Section 3 and the RTL/NVSim/NVMain/Capo flow of Section 4, compared against a GTX 1080 Ti baseline. The 50% sparsity rate is an explicit input ('The sparsity of the CNNs are assumed to be 50%, as it is shown in [32]...'), not a quantity inferred from SPRING's outputs; the reported 15.6x/66.0x numbers are conditional on that input and would shrink if actual weight/activation sparsity is lower, but they are not algebraically equal to the assumption. This is an assumption-validity and external-benchmark concern, not a circular reduction. The stochastic-rounding convergence assumption is attributed to an external paper [45], and Section 6 openly discloses that training results are batch-level and that monolithic-3D process degradation is not modeled. There are self-citations to prior work by the same group ([69], [70] for the monolithic 3D RRAM interface, [77] for floorplanning, and [84] for the design-space exploration used to set Table 1), and the evaluation infrastructure is the authors' own. However, no load-bearing argument reduces to an unverified self-citation: the binary-mask sparsity skip, reduced-precision MAC with stochastic rounding, and benchmark comparisons are described concretely and are externally comprehensible. No equation (e.g., Eqs. 1-4, mask/AND filtering in Section 3.1) has an output that is identical to an input by construction, and no uniqueness theorem is imported from the authors' prior work. I therefore find no circular steps; score 2 reflects minor non-load-bearing self-citation in the simulation/methodology, not circularity.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new physical entities; its contributions are architectural combinations of existing techniques: binary masks, stochastic rounding, RRAM, and monolithic 3D integration. The main free parameters are the assumed sparsity and the device configuration, which directly shape the headline ratios.

free parameters (3)
  • Uniform 50% sparsity assumption = 50%
    Section 4 sets sparsity to 50% for all seven CNNs based on averages from other networks; this directly determines the skip rate and compression ratio that drive the reported speedups.
  • Fixed-point precision selection = 4 IL bits, 16 FL bits
    Taken from Gupta et al. because stochastic rounding was shown to converge there, but not validated on the ImageNet-scale models evaluated in this paper.
  • Accelerator configuration = 700 MHz; 64 PEs; 72 MAC lanes/PE; 16 multipliers/MAC; 24 MB weight buffer; 12 MB activation buffer; 4 MB mask buffer…
    Selected using the authors' design space exploration framework; no sensitivity analysis is provided, so the published ratios may reflect a tuned best point rather than a representative configuration.
assumptions (3)
  • domain assumption Stochastic rounding with 16 fractional bits trains CNNs to accuracy comparable to FP32 without extra epochs.
    Section 6 states that training results are batch-level and rely on the convergence assumption from [45]; no accuracy results are presented for the seven evaluated CNNs.
  • domain assumption Monolithic 3D integration process-induced device and interconnect degradation is negligible.
    Section 6 acknowledges and declines to model top-tier device degradation, citing works claiming less than 2% impact; if degradation is larger in practice, power and performance ratios would shrink.
  • domain assumption The GTX 1080 Ti baseline is modeled correctly from its peak TFLOPS and memory bandwidth.
    Section 4 lists peak GPU specifications but does not describe a timing, power, or energy model for the baseline; every normalized speedup inherits this unstated model.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SPRING: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference." pith.science (2026). https://pith.science/paper/QWL3WDSM

@misc{pith2026190900557,
  author       = {Pith},
  title        = {Pith review of: SPRING: A Sparsity-Aware Reduced-Precision Monolithic 3D CNN Accelerator Architecture for Training and Inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QWL3WDSM}},
  note         = {Machine review of arXiv:1909.00557}
}
read the original abstract

CNNs outperform traditional machine learning algorithms across a wide range of applications. However, their computational complexity makes it necessary to design efficient hardware accelerators. Most CNN accelerators focus on exploring dataflow styles that exploit computational parallelism. However, potential performance speedup from sparsity has not been adequately addressed. The computation and memory footprint of CNNs can be significantly reduced if sparsity is exploited in network evaluations. To take advantage of sparsity, some accelerator designs explore sparsity encoding and evaluation on CNN accelerators. However, sparsity encoding is just performed on activation or weight and only in inference. It has been shown that activation and weight also have high sparsity levels during training. Hence, sparsity-aware computation should also be considered in training. To further improve performance and energy efficiency, some accelerators evaluate CNNs with limited precision. However, this is limited to the inference since reduced precision sacrifices network accuracy if used in training. In addition, CNN evaluation is usually memory-intensive, especially in training. In this paper, we propose SPRING, a SParsity-aware Reduced-precision Monolithic 3D CNN accelerator for trainING and inference. SPRING supports both CNN training and inference. It uses a binary mask scheme to encode sparsities in activation and weight. It uses the stochastic rounding algorithm to train CNNs with reduced precision without accuracy loss. To alleviate the memory bottleneck in CNN evaluation, especially in training, SPRING uses an efficient monolithic 3D NVM interface to increase memory bandwidth. Compared to GTX 1080 Ti, SPRING achieves 15.6X, 4.2X and 66.0X improvements in performance, power reduction, and energy efficiency, respectively, for CNN training, and 15.5X, 4.5X and 69.1X improvements for inference.

Figures

Figures reproduced from arXiv: 1909.00557 by the authors.

Figure 1
Figure 1. CNN architecture illustration 2 BACKGROUND In this section, we discuss the background material nec￾essary for understanding our proposed sparsity-aware reduced-precision accelerator architecture. We first give a primer on CNNs. We then discuss existing sparsity-aware designs. Then, we discuss various CNN training algorithms that use low numerical precision. Finally, we describe an efficient on-chip memory interface … view at source ↗
Figure 2
Figure 2. The SPRING architecture 3 SPARSITY-AWARE REDUCED-PRECISION AC￾CELERATOR ARCHITECTURE In this section, we present the proposed architecture, SPRING: a sparsity-aware reduced-precision CNN accel￾erator for both training and inference. We first discuss accelerator architecture design and then dive into sparsity￾aware acceleration, reduced-precision processing, and the monolithic 3D NVRAM interface [PITH_FULL_IMAGE:fig… view at source ↗
Figure 4
Figure 4. shows the main components of a PE. The com￾pressed data are buffered by the activation FIFO and weight FIFO. Then, they enter the pre-compute sparsity module along with the binary masks. Multiple multiplier￾CPU DMA controller RRAM system Activation buffer Control block Weight buffer Mask buffer PE … PE [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (9 more)
Figure 5
Figure 5. Figure 5: The binary mask scheme: An example 3.1 Sparsity-aware acceleration Traditional accelerator designs can only process dense data and do not support sparse-encoded computation. They treat zero elements in the same manner as regular data and thus perform operations that ha…
Figure 6
Figure 6. Figure 6: The pre-compute sparsity module AND of the activation and weight masks. The output mask, together with the activation and weight masks, is used by two more XOR gates for filter mask generation [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: The submodules of the pre-compute sparsity module: (a) mask generation, (b) dangling-data filter, and (c) zero-collapsing shifter [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: The MAC lane where IL denotes the number of bits for the integer portion and FL denotes the number of bits for the fraction part. The zero-free activations and weights from the pre-compute sparsity module are subject to multiplications in the MAC lanes, where the produ…
Figure 9
Figure 9. Figure 9: Read/write decoupled interconnects [70] memory bus (1KB wide) is used in each channel, since the interconnects between SPRING and memory controllers, and between memory controllers and RRAM ranks, are implemented using vertical MIVs. This on-chip memory bus not only re…
Figure 10
Figure 10. Figure 10: shows the simulation flow used to evaluate the proposed SPRING accelerator architecture. We implement components of SPRING at the register-transfer level (RTL) with SystemVerilog to estimate delay, power, and area. The RTL design is synthesized by Design Compiler [76]…
Figure 11
Figure 11. Figure 11: Normalized training performance is 471mm2 and the base operating frequency is 1.48 GHz, which can be boosted to 1.58 GHz. GTX 1080 Ti uses an 11 GB GDDR5X memory with 484 GB/s memory bandwidth to provide 10.16 TFLOPS peak single-precision performance. We evaluate SPRI…
Figure 15
Figure 15. Figure 15: Normalized energy efficiency in training [PITH_FULL_IMAGE:figures/full_fig_p010_15.png]
Figure 16
Figure 16. Figure 16: Normalized energy efficiency in inference [PITH_FULL_IMAGE:figures/full_fig_p010_16.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

101 extracted references · 75 canonical work pages

  1. [32]

    Compressing DMA engine: Leveraging activation sparsity for training deep neural networks,

    M. Rhu, M. O’Connor, N. Chatterjee, J. Pool, Y. Kwon, and S. W. Keckler, “Compressing DMA engine: Leveraging activation sparsity for training deep neural networks,” in Proc. IEEE Int. Symp. High Performance Computer Architecture , Feb. 2018, pp. 78– 91

  2. [45]

    Deep learning with limited numerical precision,

    S. Gupta, A. Agrawal, K. Gopalakrishnan, and P . Narayanan, “Deep learning with limited numerical precision,” in Proc. Int. Conf. Machine Learning, July 2015, pp. 1737–1746. 13

  3. [1]

    DaDianNao: A machine-learning supercomputer,

    Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “DaDianNao: A machine-learning supercomputer,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Dec. 2014, pp. 609–622

  4. [2]

    Scaledeep: A scalable compute architecture for learning and evaluating deep networks,

    S. Venkataramani, A. Ranjan, S. Banerjee, D. Das, S. Avancha, A. Jagannathan, A. Durg, D. Nagaraj, B. Kaul, P . Dubey, and A. Raghunathan, “Scaledeep: A scalable compute architecture for learning and evaluating deep networks,” in Proc. Int. Symp. Computer Architecture, June 2017, pp. 13–26

  5. [3]

    Image classi- fication at supercomputer scale,

    C. Ying, S. Kumar, D. Chen, T. Wang, and Y. Cheng, “Image classi- fication at supercomputer scale,” arXiv preprint arXiv:1811.06992, Dec. 2018

  6. [4]

    Ultra-performance Pascal GPU and NVLink interconnect,

    D. Foley and J. Danskin, “Ultra-performance Pascal GPU and NVLink interconnect,” IEEE Micro, vol. 37, no. 2, pp. 7–17, Mar. 2017

  7. [5]

    Volta: Performance and programmability,

    J. Choquette, O. Giroux, and D. Foley, “Volta: Performance and programmability,”IEEE Micro, vol. 38, no. 2, pp. 42–52, Mar. 2018

  8. [6]

    A network- centric hardware/algorithm co-design to accelerate distributed training of deep neural networks,

    Y. Li, J. Park, M. Alian, Y. Yuan, Z. Qu, P . Pan, R. Wang, A. Schwing, H. Esmaeilzadeh, and N. S. Kim, “A network- centric hardware/algorithm co-design to accelerate distributed training of deep neural networks,” in Proc. IEEE/ACM Int. Symp. Microarchitecture, Oct. 2018, pp. 175–188

Show all 101 references
  1. [7]

    Escher: A CNN accelerator with flexible buffering to minimize off-chip transfer,

    Y. Shen, M. Ferdman, and P . Milder, “Escher: A CNN accelerator with flexible buffering to minimize off-chip transfer,” in Proc. Int. Symp. Field-Programmable Custom Computing Machines , Apr. 2017, pp. 93–100

  2. [8]

    Scalpel: Customizing DNN pruning to the under- lying hardware parallelism,

    J. Yu, A. Lukefahr, D. Palframan, G. Dasika, R. Das, and S. Mahlke, “Scalpel: Customizing DNN pruning to the under- lying hardware parallelism,” in Proc. Int. Symp. Computer Archi- tecture, June 2017, pp. 548–560

  3. [9]

    Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks,

    H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, V . Chandra, and H. Esmaeilzadeh, “Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks,” in Proc. Int. Symp. Computer Architecture, June 2018, pp. 764–775

  4. [10]

    MAERI: Enabling flexi- ble dataflow mapping over DNN accelerators via reconfigurable interconnects,

    H. Kwon, A. Samajdar, and T. Krishna, “MAERI: Enabling flexi- ble dataflow mapping over DNN accelerators via reconfigurable interconnects,” in Proc. Int. Conf. Architectural Support Program- ming Languages Operating Syst., Mar. 2018, pp. 461–475

  5. [11]

    A reconfigurable fabric for accelerating large-scale datacenter services,

    A. Putnam, A. M. Caulfield, E. S. Chung, D. Chiou, K. Con- stantinides, J. Demme, H. Esmaeilzadeh, J. Fowers, G. P . Gopal, J. Gray, M. Haselman, S. Hauck, S. Heil, A. Hormati, J.-Y. Kim, S. Lanka, J. Larus, E. Peterson, S. Pope, A. Smith, J. Thong, P . Y. Xiao, and D. Burger, ...

  6. [12]

    FPGA based implementation of deep neural networks using on-chip memory only,

    J. Park and W. Sung, “FPGA based implementation of deep neural networks using on-chip memory only,” in Proc. IEEE Int. Conf. Acoustics, Speech, Signal Processing, Mar. 2016, pp. 1011–1015. 12

  7. [13]

    Overcoming resource underutilization in spatial CNN accelerators,

    Y. Shen, M. Ferdman, and P . Milder, “Overcoming resource underutilization in spatial CNN accelerators,” in Proc. Int. Conf. Field Programmable Logic Applications, Aug. 2016, pp. 1–4

  8. [14]

    Going deeper with embedded FPGA platform for convolutional neural network,

    J. Qiu, J. Wang, S. Yao, K. Guo, B. Li, E. Zhou, J. Yu, T. Tang, N. Xu, S. Song, Y. Wang, and H. Yang, “Going deeper with embedded FPGA platform for convolutional neural network,” in Proc. ACM/SIGDA Int. Symp. Field-Programmable Gate Arrays , 2016, pp. 26–35

  9. [15]

    ShiDianNao: Shifting vision processing closer to the sensor,

    Z. Du, R. Fasthuber, T. Chen, P . Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam, “ShiDianNao: Shifting vision processing closer to the sensor,” in Proc. ACM/IEEE Int. Symp. Computer Architecture, June 2015, pp. 92–104

  10. [16]

    Chain-NN: An energy-efficient 1D chain architecture for accelerating deep convolutional neural networks,

    S. Wang, D. Zhou, X. Han, and T. Yoshimura, “Chain-NN: An energy-efficient 1D chain architecture for accelerating deep convolutional neural networks,” in Proc. Design, Automation Test Europe Conf. Exhibition, Mar. 2017, pp. 1032–1037

  11. [17]

    CirCNN: Accelerating and compressing deep neural networks using block-circulant weight matrices,

    C. Ding, S. Liao, Y. Wang, Z. Li, N. Liu, Y. Zhuo, C. Wang, X. Qian, Y. Bai, G. Yuan, X. Ma, Y. Zhang, J. Tang, Q. Qiu, X. Lin, and B. Yuan, “CirCNN: Accelerating and compressing deep neural networks using block-circulant weight matrices,” in Proc. IEEE/ACM Int. Symp. Microarc...

  12. [18]

    TETRIS: Scalable and efficient neural network acceleration with 3D mem- ory,

    M. Gao, J. Pu, X. Yang, M. Horowitz, and C. Kozyrakis, “TETRIS: Scalable and efficient neural network acceleration with 3D mem- ory,” in Proc. Int. Conf. Architectural Support Programming Lan- guages Operating Syst., 2017, pp. 751–764

  13. [19]

    In-datacenter performance analysis of a tensor processing unit,

    N. P . Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P .-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V . Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D...

  14. [20]

    Fused-layer CNN accelerators,

    M. Alwani, H. Chen, M. Ferdman, and P . Milder, “Fused-layer CNN accelerators,” in Proc. IEEE/ACM Int. Symp. Microarchitec- ture, Oct. 2016, pp. 1–12

  15. [21]

    Accelerating CNN algorithm with fine-grained dataflow architectures,

    T. Xiang, Y. Feng, X. Ye, X. Tan, W. Li, Y. Zhu, M. Wu, H. Zhang, and D. Fan, “Accelerating CNN algorithm with fine-grained dataflow architectures,” in Proc. IEEE Int. Conf. High Performance Computing Communications; IEEE Int. Conf. Smart City; IEEE Int. Conf. Data Science Syst....

  16. [22]

    DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,

    T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,” in Proc. Int. Conf. Architectural Support Programming Languages Operating Syst. , Mar. 2014, pp. 269–284

  17. [23]

    Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks,

    W. Lu, G. Yan, J. Li, S. Gong, Y. Han, and X. Li, “Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks,” in Proc. IEEE Int. Symp. High Performance Computer Architecture, Feb. 2017, pp. 553–564

  18. [24]

    Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,

    Y. Chen, J. Emer, and V . Sze, “Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,” in Proc. ACM/IEEE Int. Symp. Computer Architecture , June 2016, pp. 367–379

  19. [25]

    EIE: Efficient inference engine on compressed deep neural network,

    S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “EIE: Efficient inference engine on compressed deep neural network,” in Proc. Int. Symp. Computer Architecture , June 2016, pp. 243–254

  20. [26]

    SparseNN: An energy- efficient neural network accelerator exploiting input and output sparsity,

    J. Zhu, J. Jiang, X. Chen, and C. Tsui, “SparseNN: An energy- efficient neural network accelerator exploiting input and output sparsity,” in Proc. Design, Automation Test Europe Conf. Exhibition , Mar. 2018, pp. 241–244

  21. [27]

    SCNN: An accelerator for compressed-sparse convolutional neural net- works,

    A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “SCNN: An accelerator for compressed-sparse convolutional neural net- works,” in Proc. ACM/IEEE Int. Symp. Computer Architecture, June 2017, pp. 27–40

  22. [28]

    UCNN: Exploiting computational reuse in deep neural networks via weight repetition,

    K. Hegde, J. Yu, R. Agrawal, M. Yan, M. Pellauer, and C. W. Fletcher, “UCNN: Exploiting computational reuse in deep neural networks via weight repetition,” in Proc. Int. Symp. Computer Architecture, June 2018, pp. 674–687

  23. [29]

    Cnvlutin: Ineffectual-neuron-free deep neural network computing,

    J. Albericio, P . Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” in Proc. ACM/IEEE Int. Symp. Computer Architecture, June 2016, pp. 1–13

  24. [30]

    Cambricon-X: An accelerator for sparse neural networks,

    S. Zhang, Z. Du, L. Zhang, H. Lan, S. Liu, L. Li, Q. Guo, T. Chen, and Y. Chen, “Cambricon-X: An accelerator for sparse neural networks,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Oct. 2016, pp. 1–12

  25. [31]

    Imagenet classi- fication with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classi- fication with deep convolutional neural networks,” in Proc. Int. Conf. Neural Information Processing Syst., Dec. 2012, pp. 1097–1105

  26. [33]

    (2014) Hybrid Mem- ory Cube specification 2.1

    HMC Consortium. (2014) Hybrid Mem- ory Cube specification 2.1. [Online]. Avail- able: http://hybridmemorycube .org/files/SiteDownloads/ HMC-30G-VSR HMCC Specification Rev2.1 20151105.pdf

  27. [34]

    (2016) Samsung begins mass producing world’s fastest DRAM - based on newest High Bandwidth Memory (HBM) interface

    Samsung Newsroom. (2016) Samsung begins mass producing world’s fastest DRAM - based on newest High Bandwidth Memory (HBM) interface. [Online]. Available: https://www.samsung.com/semiconductor/insights/news- events/samsung-begins-mass-producing-worlds-fastest-dram- based-on-new...

  28. [35]

    Application- transparent near-memory processing architecture with memory channel network,

    M. Alian, S. W. Min, H. Asgharimoghaddam, A. Dhar, D. K. Wang, T. Roewer, A. McPadden, O. O’Halloran, D. Chen, J. Xiong, D. Kim, W. Hwu, and N. S. Kim, “Application- transparent near-memory processing architecture with memory channel network,” in Proc. IEEE/ACM Int. Symp. Micr...

  29. [36]

    Chameleon: Versatile and practical near-DRAM acceleration architecture for large memory systems,

    H. Asghari-Moghaddam, Y. H. Son, J. H. Ahn, and N. S. Kim, “Chameleon: Versatile and practical near-DRAM acceleration architecture for large memory systems,” in Proc. IEEE/ACM Int. Symp. Microarchitecture, Oct. 2016, pp. 1–13

  30. [37]

    PRIME: A novel processing-in-memory architecture for neural network computation in ReRAM-based main memory,

    P . Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “PRIME: A novel processing-in-memory architecture for neural network computation in ReRAM-based main memory,” in Proc. Int. Symp. Computer Architecture, June 2016, pp. 27–39

  31. [38]

    Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory,

    D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopadhyay, “Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory,” SIGARCH Computer Architecture News, vol. 44, no. 3, pp. 380–392, June 2016

  32. [39]

    ISAAC: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars,

    A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P . Strachan, M. Hu, R. S. Williams, and V . Srikumar, “ISAAC: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars,” in Proc. Int. Symp. Computer Architecture, June 2016, pp. 14–26

  33. [40]

    TIME: A training-in-memory architecture for memristor- based deep neural networks,

    M. Cheng, L. Xia, Z. Zhu, Y. Cai, Y. Xie, Y. Wang, and H. Yang, “TIME: A training-in-memory architecture for memristor- based deep neural networks,” in Proc. ACM/EDAC/IEEE Design Automation Conf., June 2017, pp. 1–6

  34. [41]

    NNPIM: A pro- cessing in-memory architecture for neural network acceleration,

    S. Gupta, M. Imani, H. Kaur, and T. S. Rosing, “NNPIM: A pro- cessing in-memory architecture for neural network acceleration,” IEEE Trans. Computers, vol. 68, no. 9, pp. 1325–1337, Sep. 2019

  35. [42]

    TensorDIMM: A practical near- memory processing architecture for embeddings and tensor op- erations in deep learning,

    Y. Kwon, Y. Lee, and M. Rhu, “TensorDIMM: A practical near- memory processing architecture for embeddings and tensor op- erations in deep learning,” in Proc. IEEE/ACM Int. Symp. Microar- chitecture, Oct. 2019, pp. 740–753

  36. [43]

    Deep learning inference in facebook data centers: Characterization, performance optimizations and hardware implications,

    J. Park, M. Naumov, P . Basu, S. Deng, A. Kalaiah, D. Khu- dia, J. Law, P . Malani, A. Malevich, S. Nadathur et al. , “Deep learning inference in facebook data centers: Characterization, performance optimizations and hardware implications,” arXiv preprint arXiv:1811.09886, 2018

  37. [44]

    FloatPIM: In- memory acceleration of deep neural network training with high precision,

    M. Imani, S. Gupta, Y. Kim, and T. Rosing, “FloatPIM: In- memory acceleration of deep neural network training with high precision,” in Proc. Int. Symp. Computer Architecture , June 2019, pp. 802–815

  38. [46]

    Modeling the re- source requirements of convolutional neural networks on mobile devices,

    Z. Lu, S. Rallapalli, K. Chan, and T. La Porta, “Modeling the re- source requirements of convolutional neural networks on mobile devices,” in Proc. ACM Int. Conf. Multimedia, 2017, pp. 1663–1671

  39. [47]

    Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks,

    C. Zhang, G. Sun, Z. Fang, P . Zhou, P . Pan, and J. Cong, “Caffeine: Towards uniformed representation and acceleration for deep convolutional neural networks,” in Proc. IEEE/ACM Int. Conf. Computer-Aided Design, Nov. 2016, pp. 1–8

  40. [48]

    Deep compression: Compress- ing deep neural networks with pruning, trained quantization and Huffman coding,

    S. Han, H. Mao, and W. J. Dally, “Deep compression: Compress- ing deep neural networks with pruning, trained quantization and Huffman coding,” arXiv preprint arXiv:1510.00149, 2015

  41. [49]

    Learning both weights and connections for efficient neural networks,

    S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both weights and connections for efficient neural networks,” in Proc. Int. Conf. Neural Information Processing Syst., 2015, pp. 1135–1143

  42. [50]

    Automatic performance tuning of sparse matrix kernels,

    R. W. Vuduc, “Automatic performance tuning of sparse matrix kernels,” Ph.D. dissertation, University of California, Berkeley, 2003

  43. [51]

    Stitch-X: An accelerator architecture for exploiting unstructured sparsity in deep neural networks,

    C. Lee, Y. Shao, J.-F. Zhang, A. Parashar, J. Emer, S. Keckler, and Z. Zhang, “Stitch-X: An accelerator architecture for exploiting unstructured sparsity in deep neural networks,” in Proc. SysML Conference, 2018

  44. [52]

    Dynamic warp formation and scheduling for efficient GPU control flow,

    W. W. L. Fung, I. Sham, G. Yuan, and T. M. Aamodt, “Dynamic warp formation and scheduling for efficient GPU control flow,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Dec. 2007, pp. 407–420

  45. [53]

    Thread block compaction for efficient SIMT control flow,

    W. W. L. Fung and T. M. Aamodt, “Thread block compaction for efficient SIMT control flow,” in Proc. IEEE Int. Symp. High Performance Computer Architecture, Feb. 2011, pp. 25–36

  46. [54]

    Improving GPU performance via large warps and two-level warp scheduling,

    V . Narasiman, M. Shebanow, C. J. Lee, R. Miftakhutdinov, O. Mutlu, and Y. N. Patt, “Improving GPU performance via large warps and two-level warp scheduling,” in Proc. IEEE/ACM Int. Symp. Microarchitecture, Dec. 2011, pp. 308–317

  47. [55]

    Convergence and scalarization for data-parallel architectures,

    Y. Lee, R. Krashinsky, V . Grover, S. W. Keckler, and K. Asanovi ´c, “Convergence and scalarization for data-parallel architectures,” in Proc. IEEE/ACM Int. Symp. Code Generation Optimization , Feb. 2013, pp. 1–11

  48. [56]

    Cambricon-S: Addressing irreg- ularity in sparse neural networks through a cooperative soft- ware/hardware approach,

    X. Zhou, Z. Du, Q. Guo, S. Liu, C. Liu, C. Wang, X. Zhou, L. Li, T. Chen, and Y. Chen, “Cambricon-S: Addressing irreg- ularity in sparse neural networks through a cooperative soft- ware/hardware approach,” in Proc. IEEE/ACM Int. Symp. Mi- croarchitecture, Oct. 2018, pp. 15–28

  49. [57]

    Stripes: Bit-serial deep neural network comput- ing,

    P . Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, “Stripes: Bit-serial deep neural network comput- ing,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Oct. 2016, pp. 1–12

  50. [58]

    Bit-pragmatic deep neural network comput- ing,

    J. Albericio, A. Delm ´as, P . Judd, S. Sharify, G. O’Leary, R. Genov, and A. Moshovos, “Bit-pragmatic deep neural network comput- ing,” in Proc. IEEE/ACM Int. Symp. Microarchitecture , Oct. 2017, pp. 382–394

  51. [59]

    Scalable distributed DNN training using commodity GPU cloud computing,

    N. Strom, “Scalable distributed DNN training using commodity GPU cloud computing,” in Proc. Conf. Int. Speech Communication Association, Sep. 2015

  52. [60]

    Distributed train- ing large-scale deep architectures,

    S.-X. Zou, C.-Y. Chen, J.-L. Wu, C.-N. Chou, C.-C. Tsao, K.-C. Tung, T.-W. Lin, C.-L. Sung, and E. Y. Chang, “Distributed train- ing large-scale deep architectures,” in Proc. Int. Conf. Advanced Data Mining Applications, Oct. 2017, pp. 18–32

  53. [61]

    Distributed training strategies for a computer vision deep learning algorithm on a distributed GPU cluster,

    V . Campos, F. Sastre, M. Yag ¨ues, M. Bellver, X. Gir ´o-i Nieto, and J. Torres, “Distributed training strategies for a computer vision deep learning algorithm on a distributed GPU cluster,” Procedia Computer Science, vol. 108, pp. 315–324, May 2017

  54. [62]

    Highly scalable deep learning training system with mixed- precision: Training Imagenet in four minutes,

    X. Jia, S. Song, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu, T. Chen, G. Hu, S. Shi, and X. Chu, “Highly scalable deep learning training system with mixed- precision: Training Imagenet in four minutes,” arXiv preprint arXiv:1807.11205, July 2018

  55. [63]

    Mixed precision training,

    P . Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Gar- cia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in Proc. Int. Conf. Learning Representations, May 2018

  56. [64]

    The vanishing gradient problem during learning recurrent neural nets and problem solutions,

    S. Hochreiter, “The vanishing gradient problem during learning recurrent neural nets and problem solutions,” Int. Journal Uncer- tainty, Fuzziness Knowledge-Based Syst. , vol. 6, no. 2, pp. 107–116, Apr. 1998

  57. [65]

    Dynamically scaled fixed point arithmetic,

    D. Williamson, “Dynamically scaled fixed point arithmetic,” in Proc. IEEE Pacific Rim Conf. Communications, Computers Signal Processing, May 1991, pp. 315–318 vol.1

  58. [66]

    Training deep neural networks with low precision multiplications,

    M. Courbariaux, Y. Bengio, and J.-P . David, “Training deep neural networks with low precision multiplications,” in Proc. Int. Conf. Learning Representations, May 2015

  59. [67]

    Mixed precision training of convo- lutional neural networks using integer operations,

    D. Das, N. Mellempudi, D. Mudigere, D. Kalamkar, S. Avancha, K. Banerjee, S. Sridharan, K. Vaidyanathan, B. Kaul, E. Geor- ganas, A. Heinecke, P . Dubey, J. Corbal, N. Shustrov, R. Dubtsov, E. Fomenko, and V . Pirogov, “Mixed precision training of convo- lutional neural networ...

  60. [68]

    (2015) High Bandwidth Memory

    AMD. (2015) High Bandwidth Memory. [Online]. Available: https://www.amd.com/en/technologies/hbm

  61. [69]

    Energy-efficient monolithic three- dimensional on-chip memory architectures,

    Y. Yu and N. K. Jha, “Energy-efficient monolithic three- dimensional on-chip memory architectures,” IEEE Trans. Nan- otechnology, vol. 17, no. 4, pp. 620–633, July 2018

  62. [70]

    A monolithic 3D hybrid architecture for energy-efficient computation,

    ——, “A monolithic 3D hybrid architecture for energy-efficient computation,” IEEE Trans. Multi-Scale Computing Syst. , vol. 4, no. 4, pp. 533–547, Oct. 2018

  63. [71]

    A 5ns fast write multi-level non-volatile 1 K bits RRAM memory with advance write scheme,

    S. Sheu, P . Chiang, W. Lin, H. Lee, P . Chen, Y. Chen, T. Wu, F. T. Chen, K. Su, M. Kao, K. Cheng, and M. Tsai, “A 5ns fast write multi-level non-volatile 1 K bits RRAM memory with advance write scheme,” in Proc. Symp VLSI Circuits, June 2009, pp. 82–83

  64. [72]

    (2013, May) Technology roadmap of DRAM for three major manufacturers: Samsung, SK-Hynix and Micron

    Techinsights. (2013, May) Technology roadmap of DRAM for three major manufacturers: Samsung, SK-Hynix and Micron. [Online]. Available: https://www .techinsights.com/ uploadedFiles/Public Website/Content - Primary/Marketing/ 2013/DRAM Roadmap/Report/TechInsights-DRAM- ROADMAP-0...

  65. [73]

    (2016) The crossbar RRAM advantage

    Crossbar. (2016) The crossbar RRAM advantage. [On- line]. Available: http://www .crossbar-inc.com/technology/ rram-advantages/

  66. [74]

    3D sequential integration opportunities and technology optimization,

    P . Batude, B. Sklenard, C. Fenouillet-Beranger, B. Previtali, C. Tabone, O. Rozeau, O. Billoint, O. Turkyilmaz, H. Sarhan, S. Thuries, G. Cibrario, L. Brunet, F. Deprat, J. Michallet, F. Cler- midy, and M. Vinet, “3D sequential integration opportunities and technology optimiz...

  67. [75]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift,

    S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proc. Int. Conf. Machine Learning, July 2015, pp. 448–456

  68. [76]

    (2018) Design Compiler

    Synopsys. (2018) Design Compiler. [Online]. Available: https://www .synopsys.com/support/training/rtl- synthesis/design-compiler-rtl-synthesis .html

  69. [77]

    Hybrid monolithic 3-D IC floorplanner,

    A. Guler and N. K. Jha, “Hybrid monolithic 3-D IC floorplanner,” IEEE Trans. Very Large Scale Integration Syst. , vol. 26, no. 10, pp. 1868–1880, Oct. 2018

  70. [78]

    Capo: Robust and scalable open-source min- cut floorplacer,

    J. A. Roy, D. A. Papa, S. N. Adya, H. H. Chan, A. N. Ng, J. F. Lu, and I. L. Markov, “Capo: Robust and scalable open-source min- cut floorplacer,” in Proc. Int. Symp. Physical design , Apr. 2005, pp. 224–226

  71. [79]

    FinCACTI: Ar- chitectural analysis and modeling of caches with deeply-scaled FinFET devices,

    A. Shafaei, Y. Wang, X. Lin, and M. Pedram, “FinCACTI: Ar- chitectural analysis and modeling of caches with deeply-scaled FinFET devices,” in Proc. IEEE Computer Society Annual Symp. VLSI, July 2014, pp. 290–295

  72. [80]

    CACTI 6.0: A tool to model large caches,

    N. Muralimanohar, R. Balasubramonian, and N. P . Jouppi, “CACTI 6.0: A tool to model large caches,” HP Laboratories, pp. 22–31, 2009

  73. [81]

    NVSim: A circuit-level performance, energy, and area model for emerging nonvolatile memory,

    X. Dong, C. Xu, Y. Xie, and N. P . Jouppi, “NVSim: A circuit-level performance, energy, and area model for emerging nonvolatile memory,” IEEE Trans. Comput.-Aided Design Integr. Circuits Syst. , vol. 31, no. 7, pp. 994–1007, July 2012

  74. [82]

    NVMain 2.0: A user-friendly memory simulator to model (non-)volatile memory systems,

    M. Poremba, T. Zhang, and Y. Xie, “NVMain 2.0: A user-friendly memory simulator to model (non-)volatile memory systems,” IEEE Comput. Archit. Lett., vol. 14, no. 2, pp. 140–143, July 2015

  75. [83]

    Ten- sorflow: A system for large-scale machine learning,

    M. Abadi, P . Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Lev- enberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P . Tucker, V . Vasudevan, P . Warden, M. Wicke, Y. Yu, and X. Zheng, “Ten- sorflow: A system for larg...

  76. [84]

    Software-defined design space exploration for an efficient AI accelerator architec- ture,

    Y. Yu, Y. Li, S. Che, N. K. Jha, and W. Zhang, “Software-defined design space exploration for an efficient AI accelerator architec- ture,” arXiv preprint arXiv:1903.07676, 2019

  77. [85]

    Inception- v4, Inception-Resnet and the impact of residual connections on learning,

    C. Szegedy, S. Ioffe, V . Vanhoucke, and A. A. Alemi, “Inception- v4, Inception-Resnet and the impact of residual connections on learning,” in Proc. AAAI Conf. Artificial Intelligence , Feb. 2017

  78. [86]

    Rethinking the Inception architecture for computer vision,

    C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception architecture for computer vision,” in 14 Proc. IEEE Conf. Computer Vision Pattern Recognition , June 2016, pp. 2818–2826

  79. [87]

    MobileNetV2: Inverted residuals and linear bottlenecks,

    M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in Proc. IEEE Conf. Computer Vision Pattern Recognition , June 2018, pp. 4510–4520

  80. [88]

    Learning transfer- able architectures for scalable image recognition,

    B. Zoph, V . Vasudevan, J. Shlens, and Q. V . Le, “Learning transfer- able architectures for scalable image recognition,” in Proc. IEEE Conf. Computer Vision Pattern Recognition , June 2018, pp. 8697– 8710

  81. [89]

    Progressive neural architecture search,

    C. Liu, B. Zoph, M. Neumann, J. Shlens, W. Hua, L.-J. Li, L. Fei- Fei, A. Yuille, J. Huang, and K. Murphy, “Progressive neural architecture search,” in Proc. European Conf. Computer Vision, Sep. 2018, pp. 19–34

  82. [90]

    Identity mappings in deep residual networks,

    K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in Proc. European Conf. Computer Vision , Oct. 2016, pp. 630–645

  83. [91]

    Very deep convolutional networks for large-scale image recognition,

    K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv preprint arXiv:1409.1556, 2014

  84. [92]

    ImageNet Large Scale Visual Recognition Challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” Int. Journal Computer Vision, vol. 115, no. 3, pp. 211–252, 2015

  85. [93]

    Silberman and S

    N. Silberman and S. Guadarrama. (2016) Tensorflow-slim image classification model library. [Online]. Available: https: //github.com/tensorflow/models/tree/master/research/slim

  86. [94]

    Physical design solutions to tackle FEOL/BEOL degradation in gate-level monolithic 3D ICs,

    B. W. Ku, P . Debacker, D. Milojevic, P . Raghavan, D. Verkest, A. Thean, and S. K. Lim, “Physical design solutions to tackle FEOL/BEOL degradation in gate-level monolithic 3D ICs,” in Proc. ACM Int. Symp. Low Power Electron. Design , 2016, pp. 76–81

  87. [95]

    How to cope with slow transistors in the top-tier of monolithic 3D ICs: Design studies and CAD solutions,

    S. K. Samal, D. Nayak, M. lchihashi, S. Banna, and S. K. Lim, “How to cope with slow transistors in the top-tier of monolithic 3D ICs: Design studies and CAD solutions,” in Proc. ACM Int. Symp. Low Power Electron. Design, 2016, pp. 320–325

  88. [96]

    Power-performance study of block-level monolithic 3D-ICs considering inter-tier performance variations,

    S. Panth, K. Samadi, Y. Du, and S. K. Lim, “Power-performance study of block-level monolithic 3D-ICs considering inter-tier performance variations,” in Proc. ACM Annual Design Auto. Conf., 2014

  89. [97]

    Ultra-high density 3D SRAM cell designs for monolithic 3D integration,

    C. Liu and S. K. Lim, “Ultra-high density 3D SRAM cell designs for monolithic 3D integration,” in Proc. IEEE Int. Interconnect Technol. Conf., June 2012, pp. 1–3

  90. [98]

    Compact 6T SRAM cell with robust read/write stabilizing de- sign in 45nm monolithic 3D IC technology,

    O. Thomas, M. Vinet, O. Rozeau, P . Batude, and A. Valentian, “Compact 6T SRAM cell with robust read/write stabilizing de- sign in 45nm monolithic 3D IC technology,” in Proc. IEEE Int. Conf. IC Design Technology, May 2009, pp. 195–198

  91. [99]

    Intermediate BEOL process influence on power and performance for 3DVLSI,

    H. Sarhan, S. Thuries, O. Billoint, F. Deprat, A. A. D. Sousa, P . Batude, C. Fenouillet-Beranger, and F. Clermidy, “Intermediate BEOL process influence on power and performance for 3DVLSI,” in Proc. IEEE Int. 3D Syst. Integration Conf., Aug. 2015, pp. TS1.3.1– TS1.3.5

  92. [100]

    Supporting compressed-sparse activations and weights on SIMD-like accelerator for sparse convolutional neural networks,

    C. Lin and B. Lai, “Supporting compressed-sparse activations and weights on SIMD-like accelerator for sparse convolutional neural networks,” in Proc. Asia South Pacific Design Automation Conf., Jan. 2018, pp. 105–110

  93. [101]

    Eager Pruning: Algo- rithm and architecture support for fast training of deep neural networks,

    J. Zhang, X. Chen, M. Song, and T. Li, “Eager Pruning: Algo- rithm and architecture support for fast training of deep neural networks,” in Proc. Int. Symp. Computer Architecture , June 2019, pp. 292–303. Ye Yu received the B.Eng. degree in Electronic and Computer Engineering f...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.