REVIEW 4 major objections 5 minor 119 references
MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read This paper claims that decomposing INT8-quantized LLM weights into bit-slice matrices exposes ~70% sparsity and repeated column vectors that a co-designed accelerator can exploit, delivering 9.43x speedup and 31.1x energy efficiency over an
desk verdict Bit-level repetition and prediction ideas are new, but the A100 comparison is overprovisioned and the energy metric is suspect. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Grouped bit-slice matrices with enumeration-index factorization. An INT8 weight tile is decomposed into eight bit-slice matrices; m=4 rows are grouped so each column is a 4-bit vector with at most 16 possible types. The pigeonhole principle guarantees repetition because hidden dimension H (4k-12k) vastly exceeds 2^m. Multiplication G x X becomes E x (I x X): I x X merges activations that share a bit-column type in H(1-bs) additions, and E x X' reconstructs outputs in m x 2^{m-1} additions. A CAM-based match unit finds matching columns in one cycle; BSTC's two-state code (0 versus {1,data}) compresses the sparse high-order slices; BGPP's progressive threshold theta = max(A_hat) - alpha*radius
What would settle it
Compute the average bit-slice sparsity and column-repetition rate for a target model under a specific quantization scheme: BRCR's operation-count formula shows the advantage disappears as bit sparsity approaches 50%, and BSTC's compression ratio drops below 1 when sparsity falls below about 65%. Alternatively, run BGPP on a long-context task where attention is dense (e.g., retrieval across many similar tokens) and measure whether KV-load reduction falls to near zero while generation quality drops.
Extended reading notes
Core claim
On the paper's own terms, MCBP's central discovery is that INT8-quantized LLM weight matrices, when decomposed into bit-slice matrices, exhibit about 70% average bit-sparsity and abundant repeated column vectors, while value-level sparsity is only around 7%. The paper establishes that rewriting a grouped bit-slice matrix W as E x I, where E is an enumeration matrix of unique column vectors and I maps each column to its representative, turns redundant GEMM additions into a merge step (I x X) and a reconstruction step (E x X'), cutting additions by up to 12.1x versus value sparsity and 3.8x versus naive bit-serial computing. It further claims BSTC losslessly compresses high-order bit-slices wi
Load-bearing premise
The whole design rests on the empirical observation that bit-slice matrices of INT8-quantized LLM weights are highly sparse (about 70%) with many repeated column vectors, and that attention can be pruned bit-by-bit at a tuned threshold without losing accuracy; if quantized weights had near-random bit distributions or attention were dense, the reported speedups would shrink or vanish.
Editorial extensions
If this is right
- BRCR and BSTC are lossless: they remove operations and bytes, so the gains transfer directly to any INT8-quantized model whose bit slices show the claimed ~70% sparsity and repetition.
- BGPP is lossy: its pruning strength is controlled by alpha_r, and the paper shows accuracy can be held at 0% or 1% loss relative to INT8 by tuning alpha_r per benchmark.
- Across 26 benchmarks, MCBP's standard/aggressive configurations give average 8.72x/9.43x speedup and 29.2x/31.1x energy efficiency over A100 at the same batch size; on GPU the algorithms alone give 1.03x, so the custom hardware is what realizes the gains.
- MCBP reports 75.8% average memory-access reduction, with BSTC cutting weight traffic in decoding and BGPP cutting KV traffic, addressing the bottleneck that shifts with prompt length.
- The scheme extends to INT4: bit sparsity remains about 4x higher than value sparsity, yielding 51% computation and 41% memory-access reduction under PTQ INT4, and MCBP is reported to beat a reproduced INT4 lookup accelerator by 1.5-1.8x while consuming 33-50% less energy.
Reading between the lines
- The empirical basis (bit sparsity ~70%, abundant repetition) is measured on five LLMs, not guaranteed by theory; quantization schemes that whiten or balance bit distributions could erase BRCR/BSTC's advantage, so the obvious stress test is measuring bit statistics on more diverse models and quantizers.
- The group-size sweet spot m=4 balances repetition against the exponential 2^m reconstruction cost; a natural extension is hierarchical or approximate matching that keeps repetition high for larger m, which the fixed m=4 design would miss.
- BGPP's threshold depends on per-benchmark alpha_r tuning; an accuracy-aware adaptive alpha that self-adjusts during decoding could make the technique parameter-free and extend its KV-traffic savings to longer contexts.
- The same enumeration-index factorization could in principle apply to activations or KV cache entries, not just static weights; MCBP limits BRCR to weights, so exploiting bit repetition on dynamic values is an open direction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MCBP, a bit-slice-level algorithm-hardware co-design for LLM inference targeting GEMM computation, weight access, and KV-cache access. Three mechanisms are introduced: BRCR reduces GEMM work by grouping bit-slice weight vectors and reusing repeated column-vector computations via a CAM-based matcher; BSTC losslessly compresses high-order bit-slice weight matrices using a two-state coding aligned with the BRCR group granularity; BGPP performs progressive bit-level top-k prediction with early termination to reduce KV-cache traffic. The authors implement the design in RTL (TSMC 28nm, with CACTI/Ramulator modeling) and evaluate it on five LLMs and 26 benchmark settings, reporting a 9.43x inference speedup and 31.1x energy-efficiency improvement over an NVIDIA A100, and 35x/5.2x/3.2x energy savings over SpAtten/FACT/SOFA.
Significance. If the headline comparisons are made rigorous, MCBP would be a valuable contribution: the bit-slice repetition and sparsity observations are clearly demonstrated across several LLMs (Figs. 5, 8, 25), the BRCR/BSTC mechanisms are lossless, and the hardware designs (CAM match, lightweight codec, clock-gated BGPP) are concrete and backed by RTL simulation and area/power characterization. The paper also ships a thorough ablation study and a sensitivity analysis across PTQ/QAT/INT4 quantization, which are strengths. However, the contribution is not yet fully supported because the A100 comparison uses an inconsistent processor-scaling calculation, the BGPP threshold is tuned per benchmark on the test set to achieve the reported 0%/1% accuracy loss, and the cross-accelerator GOPS/W comparisons rely on heterogeneous operation definitions. These issues directly affect the claimed speedup and energy numbers, so the evaluation needs substantial revision before the central claims can be accepted.
major comments (4)
- [§5.3, Table 4, Fig. 20] The A100 comparison is arithmetically inconsistent as printed. The text says: 'Given the INT8 compute power of A100 is 624 TOPS, we use 148 MCBP processors (total with 622TOPS@INT8)'. Table 4 lists one MCBP processor as 54,463 GOPS ≈ 54.5 TOPS. 148 such processors would provide ≈8,060 TOPS, not 622; their total area/power would be ≈1,409 mm² and ≈354 W, versus the A100's 826 mm² and 400 W. If 'processor' instead means a PE cluster (20 per chip, Table 3), then 148 clusters correspond to 7.4 chips and provide ≈403 TOPS—again not 622. Neither reading supports the stated 622 TOPS, and the comparison is neither iso-area nor iso-power. Moreover, MCBP's GOPS likely counts bit-level additions/shifts rather than INT8 MACs, so normalizing to A100 INT8 TOPS is not established. The 8.72x/9.43x speedup and 31.1x efficiency claims need a corrected scaling and a common metric definition.
- [§5.2 and §6 (Fig. 24a)] The 'standard (0% loss)' configuration is obtained by per-benchmark tuning of alpha_r on the test set. Section 5.2 states: 'we adjust the value of alpha_r in 0.1 increments to evaluate the accuracy and overhead for each benchmark. This yields two MCBP configurations: standard (0% loss), aggressive (1% loss)'. Section 6 acknowledges that BGPP introduces alpha_r and that accuracy is affected. Thus the reported latency/KV-reduction numbers are not achieved by a single fixed configuration; they depend on thresholds selected to meet an accuracy target on the evaluation benchmarks. Fig. 24(a) even shows MBPP accuracy already degrading at alpha_r=0.6, so a global alpha_r in the claimed 0.5–0.6 range may not give 0% loss. Please report results for a fixed alpha_r (or a validation-set selection) and show per-benchmark accuracy relative to the INT8 baseline without test-set tuning.
- [Table 4, §5.4] The cross-accelerator energy-efficiency comparison is not apples-to-apples. The text states that 'the average energy efficiency is evaluated using the metric from each respective paper'. MCBP's GOPS/W is a bit-level operation count, whereas SpAtten/FACT/SOFA report value-level MAC-based GOPS/W. Different papers also use different workloads, memory systems, and technology normalizations. Comparing these absolute GOPS/W numbers (35x, 5.2x, 3.2x savings) is therefore not a reliable measure of relative efficiency. The authors should normalize all designs to a common operation definition (e.g., INT8 MAC-equivalents) and, ideally, report per-workload energy and latency rather than only aggregate GOPS/W.
- [§3.1 (BRCR complexity)] The printed complexity comparison for the value-based sparsity scheme is inconsistent with the stated conclusion. The text says the value-based scheme consumes H×m×k×vs additions, where vs≈0.07 is the value sparsity ratio (fraction of zeros). For H=4000, m=4, k=8, this gives 8960 additions for a k-bit GEMV, which is far less than the claimed 12.1x reduction relative to BRCR's k(H(1−bs)+m·2^{m−1}) ≈ 9856 additions. The claim requires H×m×k×(1−vs) additions. Please correct the formula and re-derive the 12.1x/3.8x ratios.
minor comments (5)
- [Fig. 5(g), §3.3] The label 'Theoretically optimal' in Fig. 5(g) is not defined. Clarify what bound this represents (e.g., lower bound on KV traffic given bit-level early termination).
- [Table 2] The '0% loss' standard configuration is not reflected exactly in the table: several MCBP(S) entries are 0.1–0.2 absolute points below the corresponding INT8 baseline (e.g., MMLU 34.6 vs 34.7, MBPP 17.1 vs 17.2). State whether 0% loss means rounding to one decimal or a tolerance, and report the actual worst-case delta.
- [Fig. 9 / §3.3] The bit ordering for BGPP is described inconsistently: the text says 'fetch the MSB of all Ks' while the figure caption says 'Load 1st bit (K)'. Since weights are stored in sign-magnitude format (§3.2), clarify whether bit 0 is the sign bit and which bit is fetched in round 1.
- [§4.1] The text says the workflow operates through 'eight key steps' with numbered markers ❶–❽, but the paragraph only enumerates ❶–❺ and ❻–❽ in the BGPP description. Make the correspondence between markers and steps explicit in the text.
- [§7 / abstract] The claim that MCBP is 'the first work that uses bit-level strategies for LLM inference' is too strong given that FuseKNA and Bitwave are bit-level accelerators and are adapted to GEMV in this paper. Soften the novelty claim to 'first to jointly apply bit-level repetition, sparsity, and prediction to both prefill and decoding'.
Circularity Check
BGPP's accuracy-preservation claim is fitted per benchmark via alpha_r, making the '0% loss' label true by construction; headline speedups are reported at these tuned operating points.
-
fitted input called prediction
[Section 5.2 (Algorithm settings); used in Table 2, Fig. 20 and Abstract/Section 5.3 headline numbers.]
"we adjust the value of αr in 0.1 increments to evaluate the accuracy and overhead for each benchmark. This yields two MCBP configurations: standard (0% loss), aggressive (1% loss), representing the minimal and maximal performance optimizations, respectively."
The 0%/1% loss targets are the selection criteria for αr, not outputs of an independent prediction. 'Standard' and 'aggressive' are defined by the loss achieved after tuning αr on the same benchmark, so Table 2's accuracy preservation is an input constraint. Since αr directly sets the BGPP pruning threshold (Eq. 1), the per-benchmark tuned αr determines the KV reduction; the subsequent 8.72×/9.43× speedup and 29.2×/31.1× efficiency gains are measured at these fitted operating points rather than at a single pre-specified threshold, making the accuracy part of the claim circular and the performance claims contingent on in-sample tuning.
full rationale
BRCR and BSTC are self-contained: the bit-slice decomposition, pigeonhole argument, and operation-count formulas (Sec. 3.1, 3.2) derive reductions from measured bit-level sparsity/repetition without circularity. The softmax-gap property is cited to the authors' FACT [72], but it is a standard mathematical fact and not load-bearing. The main circularity is confined to BGPP evaluation: alpha_r is tuned per benchmark to hit 0%/1% loss, so the 'standard (0% loss)' accuracy is true by construction, and the headline speedups are reported at these benchmark-specific operating points. The DSE for group size m on the same workloads is a design choice rather than a prediction, and the A100 comparison contains an internal scaling inconsistency (148 MCBP processors vs Table 4 throughput) that is a validity concern but not circularity. Overall, the central algorithmic derivation is mostly independent, so the score is moderate rather than high.
Assumptions & free parameters
free parameters (5)
- Group size m (BRCR/BSTC) =
4
- BGPP threshold coefficient alpha_r =
0.5-0.6 per benchmark (standard config 0% loss; aggressive 1% loss)
- BGPP radius =
3
- BSTC compressed bit-slice set =
bits 3-7 compressed; bits 1, 2, 8 uncompressed
- BGPP number of filtering rounds =
unreported
assumptions (5)
- standard math Pigeonhole principle: any m-row binary matrix has at most 2^m distinct column vectors, so with H >> 2^m repetitions must occur.
- standard math For large H, all 2^m possible column vectors appear, bounding the reconstruction add cost by m*2^(m-1).
- domain assumption Quantized LLM weights are near-Gaussian, so high-order bit-slices are mostly zero and lower-order slices are roughly uniform, giving high bit sparsity and repetition.
- domain assumption Attention scores are sparse, and a low-bit progressive estimate with early termination can select the true top-k keys without accuracy loss.
- domain assumption The RTL/Verilator + CACTI + Ramulator + TSMC 28nm synthesis flow accurately represents the performance, power, and area of a real MCBP chip.
Cite this review
Pith. "Pith review of MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness." pith.science (2026). https://pith.science/paper/TWMJCLWO
@misc{pith2026250910372,
author = {Pith},
title = {Pith review of: MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness},
year = {2026},
howpublished = {\url{https://pith.science/paper/TWMJCLWO}},
note = {Machine review of arXiv:2509.10372}
}
read the original abstract
Large language models (LLMs) face significant inference latency due to inefficiencies in GEMM operations, weight access, and KV cache access, especially in real-time scenarios. This highlights the need for a versatile compute-memory efficient accelerator. Unfortunately, existing Transformer accelerators struggle to address both aspects simultaneously, as they focus on value-level processing, missing fine-grained opportunities to optimize computation and memory collaboratively. This paper introduces MCBP, a bit-grained compute-memory efficient algorithm-hardware co-design that leverages bit-slice (BS) enabled repetitiveness and sparsity to accelerate LLM inference. MCBP features three key innovations: 1) BS-repetitiveness-enabled computation reduction (BRCR), which eliminates redundant GEMM computations via leveraging redundancy hidden among BS vectors; 2) BS-sparsity-enabled two-state coding (BSTC), which reduces weight access via exploiting significant sparsity in high-order bit-slice weight; 3) Bit-grained progressive prediction (BGPP), which reduces KV cache access by leveraging early-termination-based bit-grained prediction. These techniques, supported by custom accelerator designs, effectively alleviate the burden in GEMM, weight access, and KV cache access. Extensive experiments on 26 benchmarks show that MCBP achieves 9.43x speed up and 31.1x higher energy efficiency than Nvidia A100 GPU. Compared to SOTA Transformer accelerators, MCBP achieves 35x, 5.2x and 3.2x energy saving than Spatten, FACT and SOFA, respectively.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
Miklós Ajtai. 1994. The complexity of the pigeonhole principle.Combinatorica 14 (1994), 417–433
1994
-
[2]
Jorge Albericio, Alberto Delmás, Patrick Judd, Sayeh Sharify, Gerard O’Leary, Roman Genov, and Andreas Moshovos. 2017. Bit-pragmatic deep neural network computing. InProceedings of the 50th annual IEEE/ACM international symposium on microarchitecture. 382–394
2017
-
[3]
Renzo Andri, Lukas Cavigelli, Davide Rossi, and Luca Benini. 2016. YodaNN: An ultra-low power convolutional neural network accelerator based on binary weights. InProceedings of the IEEE Computer Society Annual Symposium on VLSI (ISVLSI). 236–241
2016
-
[4]
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report.arXiv preprint arXiv:2305.10403(2023)
arXiv 2023
-
[5]
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021. Program synthesis with large language models.arXiv preprint arXiv:2108.07732(2021)
arXiv 2021
-
[6]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng X...
arXiv 2023
-
[7]
Zhenyu Bai, Pranav Dangi, Huize Li, and Tulika Mitra. 2024. SWAT: Scalable and efficient window attention-based transformers acceleration on FPGAs. In Proceedings of the 61st ACM/IEEE Design Automation Conference. 1–6
2024
-
[8]
Lukas Cavigelli and Luca Benini. 2016. Origami: A 803-GOp/s/W convolu- tional network accelerator.IEEE Transactions on Circuits and Systems for Video Technology27, 11 (2016), 2461–2475
2016
Show all 119 references
-
[9]
Yu, Qiang Yang, and Xing Xie
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024. A survey on evaluation of large language models.ACM Transactions on Intelligent S...
2024
-
[10]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...
2021 arXiv
-
[11]
Yi Chen, Yongwei Zhao, Yifan Hao, Yuanbo Wen, Yuntao Dai, Xiaqing Li, Yang Liu, Rui Zhang, Mo Zou, Xinkai Song, Xing Hu, Zidong Du, Huaping Chen, Qi Guo, and Tianqi Chen. 2024. Cambricon-C: Efficient 4-Bit Matrix Unit via Primi- tivization. In2024 57th IEEE/ACM International S...
2024
-
[12]
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al
-
[13]
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free Dolly: Introducing the world’s first truly open instruction-tuned LLM.Company Blog of Databricks(2023)
2023
-
[14]
Alberto Delmas, Patrick Judd, Sayeh Sharify, and Andreas Moshovos. 2017. Dynamic stripes: Exploiting the dynamic precision requirements of activation values in neural networks.arXiv preprint arXiv:1706.00504(2017)
2017 arXiv
-
[15]
Alberto Delmas Lascorz, Patrick Judd, Dylan Malone Stuart, Zissis Poulos, Mostafa Mahmoud, Sayeh Sharify, Milos Nikolic, Kevin Siu, and Andreas Moshovos. 2019. Bit-tactical: A software/hardware approach to exploiting value and bit sparsity in neural networks. InProceedings of ...
2019
-
[16]
Chunhua Deng, Yang Sui, Siyu Liao, Xuehai Qian, and Bo Yuan. 2021. GoSPA: An energy-efficient high-performance globally optimized sparse convolutional neural network accelerator. InProceedings of the ACM/IEEE 48th Annual Inter- national Symposium on Computer Architecture (ISCA...
2021
-
[17]
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. GPT3.Int8 (): 8-bit matrix multiplication for Transformers at scale.Advances in Neural Information Processing Systems35 (2022), 30318–30332. 14 MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leve...
2022
-
[18]
Claire Cardie Faisal Ladhak, Esin Durmus and Kathleen McKeown. 2020. Wik- iLingua: A new benchmark dataset for multilingual abstractive summarization. InFindings of EMNLP, 2020
2020
-
[19]
Hongxiang Fan, Thomas Chau, Stylianos I Venieris, Royson Lee, Alexandros Kouris, Wayne Luk, Nicholas D Lane, and Mohamed S Abdelfattah. 2022. Adapt- able butterfly accelerator for attention-based NNs via hardware and algorithm co-design. InProceedings of the 55th IEEE/ACM Inte...
2022
-
[20]
Zichen Fan, Qirui Zhang, Pierre Abillama, Sara Shoouri, Changwoo Lee, David Blaauw, Hun-Seok Kim, and Dennis Sylvester. 2023. Taskfusion: An efficient transfer learning architecture with dual delta sparsity for multi-task natural language processing. InProceedings of the 50th ...
2023
-
[21]
Chao Fang, Shouliang Guo, Wei Wu, Jun Lin, Zhongfeng Wang, Ming Kai Hsu, and Lingzhi Liu. 2022. An efficient hardware accelerator for sparse Transformer neural networks. In2022 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 2670–2674
2022
-
[22]
Chao Fang, Aojun Zhou, and Zhongfeng Wang. 2022. An algorithm-hardware co-optimized framework for accelerating N:M sparse Transformers.IEEE Trans- actions on Very Large Scale Integration (VLSI) Systems30, 11 (2022), 1573–1586
2022
-
[23]
Ashish Gondimalla, Noah Chesnut, Mithuna Thottethodi, and TN Vijaykumar
-
[24]
Sumanth Gudaparthi, Sarabjeet Singh, Surya Narayanan, Rajeev Balasubra- monian, and Visvesh Sathe. 2022. CANDLES: Channel-aware novel dataflow- microarchitecture co-design for low energy sparse neural network acceleration. InProceedings of the IEEE International Symposium on h...
2022
-
[25]
Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W Lee, and Deog- Kyoon Jeong. 2020. A3: Accelerating attention mechanisms in neural networks with approximation. InProceedings of the IEEE Internation...
2020
-
[26]
Tae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim, Hyunji Choi, Sung Jun Jung, and Jae W Lee. 2021. ELSA: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks. InPrceedings of the 48th ACM/IEEE Annual International Symposium on C...
2021
-
[27]
Meng Han, Liang Wang, Limin Xiao, Hao Zhang, Tianhao Cai, Jiale Xu, Yibo Wu, Chenhao Zhang, and Xiangrong Xu. 2024. BitNN: A bit-serial accelerator for k-nearest neighbor search in point clouds. InProceddings of the ACM/IEEE 51st Annual International Symposium on Computer Arch...
2024
-
[28]
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally. 2016. EIE: Efficient inference engine on compressed deep neural network.ACM SIGARCH Computer Architecture News44, 3 (2016), 243– 254
2016
-
[29]
Song Han, Huizi Mao, and William J Dally. 2015. Deep compression: Com- pressing deep neural networks with pruning, trained quantization and huffman coding.arXiv preprint arXiv:1510.00149(2015)
2015 arXiv
-
[30]
Edward Hanson, Shiyu Li, Hai‘Helen’ Li, and Yiran Chen. 2022. Cascading structured pruning: Enabling high data reuse for sparse DNN accelerators. In Proceedings of the 49th Annual International Symposium on Computer Architecture. 522–535
2022
-
[31]
Kartik Hegde, Jiyong Yu, Rohit Agrawal, Mengjia Yan, Michael Pellauer, and Christopher Fletcher. 2018. UCNN: Exploiting computational reuse in deep neu- ral networks via weight repetition. In2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). IEE...
2018
-
[32]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language un- derstanding.arXiv preprint arXiv:2009.03300(2020)
2020 arXiv
-
[33]
Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dong- soo Lee, and Joo-Young Kim. 2022. DFX: A low-latency multi-FPGA appliance for accelerating Transformer-based text generation. InProceedings of the 55th IEEE/ACM International Symposium on Microarchitectur...
2022
-
[34]
Yuxuan Hu, Xiaodong Chen, Cuiping Li, Hong Chen, and Jing Zhang. 2025. QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition.arXiv preprint arXiv:2503.19353(2025)
2025 arXiv
-
[35]
Dongseok Im, Gwangtae Park, Zhiyong Li, Junha Ryu, and Hoi-Jun Yoo. 2023. Sibia: Signed bit-slice architecture for dense dnn acceleration with slice-level sparsity exploitation. In2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 69–80
2023
-
[36]
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, An- drew Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Infer- ence. InProceedings of the IEEE Conference on Computer ...
2018
-
[37]
Patrick Judd, Jorge Albericio, Tayler Hetherington, Tor M Aamodt, and Andreas Moshovos. 2016. Stripes: Bit-serial deep neural network computing. InProceed- ings of the 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 1–12
2016
-
[38]
Dongyun Kam, Myeongji Yun, Sunwoo Yoo, Seungwoo Hong, Zhengya Zhang, and Youngjoo Lee. 2024. Panacea: Novel DNN Accelerator using Accuracy- Preserving Asymmetric Quantization and Energy-Saving Bit-Slice Sparsity. arXiv preprint arXiv:2412.10059(2024)
2024 arXiv
-
[39]
Sanghoon Kang, Donghyeon Han, Juhyoung Lee, Dongseok Im, Sangyeob Kim, Soyeon Kim, Junha Ryu, and Hoi-Jun Yoo. 2021. GANPU: An energy-efficient multi-DNN training processor for GANs with speculative dual-sparsity exploita- tion.IEEE Journal of Solid-State Circuits56, 9 (2021),...
2021
-
[40]
Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. Codeaid: Evaluating a classroom deployment of an LLM-based programming assistant that balances student and educator needs. InProceedings of the CHI Confe...
2024
-
[41]
Yoongu Kim, Weikun Yang, and Onur Mutlu. 2015. Ramulator: A fast and extensible DRAM simulator.IEEE Computer architecture letters15, 1 (2015), 45–49
2015
-
[42]
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176B-parameter open-access multilingual language model.arXiv preprint arXiv:2211.05100(2022)
2022 arXiv
-
[43]
Jinmook Lee, Changhyeon Kim, Sanghoon Kang, Dongjoo Shin, Sangyeob Kim, and Hoi-Jun Yoo. 2018. UNPU: A 50.6 TOPS/W unified deep neural network accelerator with 1b-to-16b fully-variable weight bit-precision. InProceedings of IEEE International Solid-State Circuits Conference-(I...
2018
-
[44]
Brian Leibowitz, Robert Palmer, John Poulton, Yohan Frans, Simon Li, John Wilson, Michael Bucher, Andrew M Fuller, John Eyles, Marko Aleksic, Trey Greer, and Nhat M Nguyen. 2010. A 4.3 GB/s mobile memory interface with power-efficient bandwidth scaling.IEEE Journal of Solid-St...
2010
-
[45]
Jonathan S Lew, Yunpeng Liu, Wenyi Gong, Negar Goli, R David Evans, and Tor M Aamodt. 2022. Anticipating and eliminating redundant computations in accelerated sparse training. InProceedings of the 49th Annual International Symposium on Computer Architecture. 536–551
2022
-
[46]
Bingbing Li, Santosh Pandey, Haowen Fang, Yanjun Lyv, Ji Li, Jieyang Chen, Mimi Xie, Lipeng Wan, Hang Liu, and Caiwen Ding. 2020. FTRANS: Energy- efficient acceleration of Transformers using FPGA. InProceedings of the ACM/IEEE International Symposium on Low Power Electronics a...
2020
-
[47]
Gang Li, Weixiang Xu, Zhuoran Song, Naifeng Jing, Jian Cheng, and Xiaoyao Liang. 2022. Ristretto: An atomized processing architecture for sparsity- condensed stream flow in CNN. InProceedings of the 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). 1434–1450
2022
-
[48]
Guoyu Li, Shengyu Ye, Chunyun Chen, Yang Wang, Fan Yang, Ting Cao, Cheng Liu, Mohamed M Sabry, and Mao Yang. 2025. LUT-DLA: Lookup Ta- ble as Efficient Extreme Low-Bit Deep Learning Accelerator.arXiv preprint arXiv:2501.10658(2025)
2025 arXiv
-
[49]
Shiyu Li, Edward Hanson, Xuehai Qian, Hai" Helen" Li, and Yiran Chen. 2021. ESCALATE: Boosting the efficiency of sparse CNN accelerator with kernel decomposition. InProceedings of the 54th Annual IEEE/ACM International Sym- posium on Microarchitecture. 992–1004
2021
-
[50]
Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, and Mingu Kang. 2022. Accelerating attention through gradient-based learned runtime pruning. InProceedings of the 49th Annual International Symposium on Computer Architecture. 902–915
2022
-
[51]
Bin Lin, Chen Zhang, Tao Peng, Hanyu Zhao, Wencong Xiao, Minmin Sun, Anmin Liu, Zhipeng Zhang, Lanbo Li, Xiafei Qiu, Li Shen, Zhigang Ji, Tao Xie, Yong Li, and Wei Lin. 2024. Infinite-LLM: Efficient LLM service for long context with distattention and distributed kvcache.arXiv ...
2024 arXiv
-
[52]
Fangxin Liu, Ning Yang, Haomin Li, Zongwu Wang, Zhuoran Song, Songwen Pei, and Li Jiang. 2024. SPARK: Scalable and precision-aware acceleration of neural networks via efficient encoding. InProceedings of the IEEE International Symposium on High-Performance Computer Architectur...
2024
-
[53]
Jing Liu, Ruihao Gong, Xiuying Wei, Zhiwei Dong, Jianfei Cai, and Bohan Zhuang. 2023. QLLM: Accurate and efficient low-bitwidth quantization for large language models.arXiv preprint arXiv:2310.08041(2023)
2023 arXiv
-
[54]
Siqin Liu, Prakash Chand Kuve, and Avinash Karanth. 2024. HSCONN: Hardware-Software Co-Optimization of Self-Attention Neural Networks for Large Language Models. InProceedings of the Great Lakes Symposium on VLSI
2024
-
[55]
Shiwei Liu, Peizhe Li, Jinshan Zhang, Yunzhengmao Wang, Haozhe Zhu, Wen- ning Jiang, Shan Tang, Chixiao Chen, Qi Liu, and Ming Liu. 2023. 16.2 A 28nm 53.8 TOPS/W 8b sparse Transformer accelerator with in-memory butterfly zero skipper for unstructured-pruned NN and CIM-based lo...
2023
-
[56]
Zhi-Gang Liu, Paul N Whatmough, Yuhao Zhu, and Matthew Mattina. 2022. S2TA: Exploiting structured sparsity for energy-efficient mobile CNN accelera- tion. InProceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA). 573–586
2022
-
[57]
Yun-Chen Lo and Ren-Shuo Liu. 2023. Bit-serial cache: Exploiting input bit vector repetition to accelerate bit-serial inference. InProceedings of the 60th ACM/IEEE Design Automation Conference (DAC). 1–6
2023
-
[58]
Hang Lu, Liang Chang, Chenglong Li, Zixuan Zhu, Shengjian Lu, Yanhuan Liu, and Mingzhe Zhang. 2021. Distilling bit-level sparsity parallelism for general purpose deep learning acceleration. InMICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture. 963–976
2021
-
[59]
Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021. Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture. InProceedings of the 54th Annual IEEE/ACM International Symposium on Microarchitecture. 977–991
2021
-
[60]
Mostafa Mahmoud, Isak Edo, Ali Hadi Zadeh, Omar Mohamed Awad, Gennady Pekhimenko, Jorge Albericio, and Andreas Moshovos. 2020. TensorDash: Ex- ploiting sparsity to accelerate deep neural network training. InProceedings of the 53rd Annual IEEE/ACM International Symposium on Mic...
2020
-
[61]
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. InProceedings of the International Conference on Learning Representations
2016
-
[62]
Bert Moons, Roel Uytterhoeven, Wim Dehaene, and Marian Verhelst. 2017. 14.5 Envision: A 0.26-to-10TOPS/W subword-parallel dynamic-voltage-accuracy- frequency-scalable convolutional neural network processor in 28nm FDSOI. InProceddings of the IEEE International Solid-State Circ...
2017
-
[63]
Bert Moons and Marian Verhelst. 2016. An energy-efficient precision-scalable ConvNet processor in 40-nm CMOS.IEEE Journal of solid-state Circuits52, 4 (2016), 903–914
2016
-
[64]
Naveen Muralimanohar, Rajeev Balasubramonian, and Norman P Jouppi. 2009. CACTI 6.0: A tool to model large caches.HP laboratories27 (2009), 28
2009
-
[65]
Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an LLM to help with code understanding. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–13
2024
-
[66]
Nvidia. 2023. TensorRT-LLM. https://github.com/NVIDIA/TensorRT-LLM? tab=readme-ov-file
2023
-
[67]
Mike O’Connor, Niladrish Chatterjee, Donghyuk Lee, John Wilson, Aditya Agrawal, Stephen W Keckler, and William J Dally. 2017. Fine-grained DRAM: Energy-efficient DRAM for extreme bandwidth systems. InProceedings of the 50th Annual IEEE/ACM International Symposium on Microarchi...
2017
-
[68]
Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rang- harajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W Keckler, and William J Dally. 2017. SCNN: An accelerator for compressed-sparse convo- lutional neural networks.ACM SIGARCH computer architecture n...
2017
-
[69]
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer
-
[70]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient generative LLM infer- ence using phase splitting. In2024 ACM/IEEE 51st Annual International Sympo- sium on Computer Architecture (ISCA). IEEE...
2024
-
[71]
Eric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella, Sudarshan Srini- vasan, Dipankar Das, Bharat Kaul, and Tushar Krishna. 2020. Sigma: A sparse and irregular GEMM accelerator with flexible interconnects for DNN training. In Proceedings of the IEEE International Symposiu...
2020
-
[72]
Yubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao, Xiaolong Yang, Leibo Liu, Shaojun Wei, Yang Hu, and Shouyi Yin. 2023. Fact: FFN-attention co-optimized Transformer architecture with eager correlation prediction. InProceedings of the 50th Annual International Symposium on Compu...
2023
-
[73]
Yubin Qin, Yang Wang, Jiachen Wang, Zhiwei Lin, Yushu Zhao, Shaojun Wei, Yang Hu, and Shouyi Yin. 2025. 23.8 An 88.36 TOPS/W Bit-Level-Weight- Compressed Large-Language-Model Accelerator with Cluster-Aligned INT-FP- GEMM and Bi-Dimensional Workflow Reformulation. In2025 IEEE I...
2025
-
[74]
Yubin Qin, Yang Wang, Zhiren Zhao, Xiaolong Yang, Yang Zhou, Shaojun Wei, Yang Hu, and Shouyi Yin. 2024. MECLA: Memory-compute-efficient LLM accelerator with scaling sub-matrix partition. InProceedings of the 51st ACM/IEEE Annual International Symposium on Computer Architectur...
2024
-
[75]
Zheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen, Yufei Ding, and Yuan Xie. 2022. DOTA: Detect and omit weak attentions for scalable Transformer acceleration. InProceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Sy...
2022
-
[76]
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiao- qing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al. 2023. Code LLaMa: Open foundation models for code.arXiv preprint arXiv:2308.12950 (2023)
2023 arXiv
-
[77]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial Winograd schema challenge at scale.Commun. ACM64, 9 (2021), 99–106
2021
-
[78]
Sayeh Sharify, Alberto Delmas Lascorz, Mostafa Mahmoud, Milos Nikolic, Kevin Siu, Dylan Malone Stuart, Zissis Poulos, and Andreas Moshovos. 2019. Laconic deep learning inference acceleration. InProceedings of the 46th International Symposium on Computer Architecture. 304–317
2019
-
[79]
Hardik Sharma, Jongse Park, Naveen Suda, Liangzhen Lai, Benson Chau, Joon Kyung Kim, Vikas Chandra, and Hadi Esmaeilzadeh. 2018. Bit Fusion: Bit-level dynamically composable architecture for accelerating deep neural net- work. InProceedings of the ACM/IEEE 45th Annual Internat...
2018
-
[80]
Guan Shen, Jieru Zhao, Quan Chen, Jingwen Leng, Chao Li, and Minyi Guo
-
[81]
Man Shi, Vikram Jain, Antony Joseph, Maurice Meijer, and Marian Verhelst. 2024. BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning Accel- eration. InProceedings of IEEE International Symposium on High-Performance Computer Architecture (HPCA). 732–746
2024
-
[82]
Jong Hoon Shin, Ali Shafiee, Ardavan Pedram, Hamzah Abdel-Aziz, Ling Li, and Joseph Hassoun. 2022. Griffin: Rethinking sparse optimization for deep learning architectures. InProceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA). 861–875
2022
-
[83]
Wilson Snyder. 2004. Verilator and SystemPerl. InNorth American SystemC Users’ Group, Design Automation Conference
2004
-
[84]
Benjamin Spector and Chris Re. 2023. Accelerating LLM inference with staged speculative decoding.arXiv preprint arXiv:2308.04623(2023)
2023 arXiv
-
[85]
Salmonn Talebi, Elizabeth Tong, and Mohammad RK Mofrad. 2023. Beyond the Hype: Assessing the Performance, Trustworthiness, and Clinical Suitability of GPT3. 5.arXiv preprint arXiv:2306.15887(2023)
2023 arXiv
-
[86]
Thierry Tambe, Coleman Hooper, Lillian Pentecost, Tianyu Jia, En-Yu Yang, Marco Donato, Victor Sanh, Paul Whatmough, Alexander M Rush, David Brooks, and Gu-Yeon Wei. 2021. Edgebert: Sentence-level energy optimizations for latency-aware multi-task NLP inference. InProceedings o...
2021
-
[87]
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://crfm.stanford.edu/2023/03/13/alpaca. html
2023
-
[88]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288(2023)
2023 arXiv
-
[89]
Shikhar Tuli and Niraj K Jha. 2023. AccelTran: A sparsity-aware accelerator for dynamic inference with Transformers.IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems42, 11 (2023), 4038–4051
2023
-
[90]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in neural information processing systems30 (2017)
2017
-
[91]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. GLUE: A multi-task benchmark and analysis plat- form for natural language understanding. InProceedings of the International Conference on Learning Representations
2018
-
[92]
Huizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue, Jinxi Li, Yubin Qin, Si- han Guan, Qize Yang, Yang Wang, Chao Li, Yang Hu, and Shouyi Yin. 2024. SOFA: A compute-memory optimized sparsity accelerator via cross-stage coordinated tiling.arXiv preprint arXiv:2407.10416(2024)
2024 arXiv
-
[93]
Huizheng Wang, Weihong Xu, Zaichen Zhang, Xiaohu You, and Chuan Zhang
-
[94]
Hanrui Wang, Zhekai Zhang, and Song Han. 2021. SpAtten: Efficient sparse attention architecture with cascade token and head pruning. InProceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA). 97–110
2021
-
[95]
Huizheng Wang, Zaichen Zhang, Xiaohu You, and Chuan Zhang. 2018. Low- complexity Winograd convolution architecture based on stochastic computing. In2018 IEEE 23rd International Conference on Digital Signal Processing (DSP). IEEE, 1–5
2018
-
[96]
Yizhi Wang, Jun Lin, and Zhongfeng Wang. 2017. An energy-efficient archi- tecture for binary weight convolutional neural networks.IEEE Transactions on Very Large Scale Integration (VLSI) Systems26, 2 (2017), 280–293. 16 MCBP: A Memory-Compute Efficient LLM Inference Accelerato...
2017
-
[97]
Yang Wang, Yubin Qin, Dazheng Deng, Jingchuan Wei, Yang Zhou, Yuanqi Fan, Tianbao Chen, Hao Sun, Leibo Liu, Shaojun Wei, and Shouyi Yin. 2022. An energy-efficient Transformer processor exploiting dynamic weak relevances in global attention.IEEE Journal of Solid-State Circuits5...
2022
-
[98]
Xiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong, Shanghang Zhang, Qi Zhang, Fengwei Yu, and Xianglong Liu. 2022. Outlier suppression: Pushing the limit of low-bit transformer language models.Advances in Neural Information Processing Systems35 (2022), 17402–17414
2022
-
[99]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement De- langue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Ma...
2020
-
[100]
Yannan Nellie Wu, Po-An Tsai, Saurav Muralidharan, Angshuman Parashar, Vivienne Sze, and Joel Emer. 2023. HighLight: Efficient and flexible DNN accel- eration with hierarchical structured sparsity. InProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchi...
2023
-
[101]
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. InInternational Conference on Machine Learning. PMLR, 38087–38099
2023
-
[102]
Dingqing Yang, Amin Ghasemazar, Xiaowei Ren, Maximilian Golub, Guy Lemieux, and Mieszko Lis. 2020. Procrustes: A dataflow and accelerator for sparse deep neural network training. InProceedings of the 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 711–724
2020
-
[103]
Jianxun Yang, Zhao Zhang, Zhuangzhi Liu, Jing Zhou, Leibo Liu, Shaojun Wei, and Shouyi Yin. 2021. FuseKNA: Fused kernel convolution based accelerator for deep neural networks. InProceddings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA). 894–907
2021
-
[104]
Tao Yang, Fei Ma, Xiaoling Li, Fangxin Liu, Yilong Zhao, Zhezhi He, and Li Jiang. 2022. DTATrans: Leveraging dynamic token-based quantization with accuracy compensation mechanism for efficient Transformer architecture.IEEE Transactions on Computer-Aided Design of Integrated Ci...
2022
-
[105]
Yifan Yang, Joel S Emer, and Daniel Sanchez. 2023. ISOSceles: Accelerating sparse CNNs through inter-layer pipelining. In2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 598–610
2023
-
[106]
Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, and Mingu Kang
-
[107]
Eunji Yoo, Gunho Park, Jung Gyu Min, Se Jung Kwon, Baeseong Park, Dong- soo Lee, and Youngjoo Lee. 2023. TF-MVP: Novel sparsity-aware transformer accelerator with mixed-length vector pruning. In2023 60th ACM/IEEE Design Automation Conference (DAC). IEEE, 1–6
2023
-
[108]
Haoran You, Zhanyi Sun, Huihong Shi, Zhongzhi Yu, Yang Zhao, Yongan Zhang, Chaojian Li, Baopu Li, and Yingyan Lin. 2023. ViTCoD: Vision Transformer acceleration via dedicated algorithm and accelerator co-design. InProceedings of the IEEE International Symposium on High-Perform...
2023
-
[109]
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019. Q8BERT: Quantized 8bit BERT. In2019 Fifth Workshop on Energy Efficient Machine Learn- ing and Cognitive Computing-NeurIPS Edition (EMC2-NIPS). IEEE, 36–39
2019
-
[110]
Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1–12
2016
-
[111]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...
2022 arXiv
-
[112]
InProceedings of the 55th IEEE/ACM International Sym- posium on Microarchitecture (MICRO)
Sparse attention acceleration with synergistic in-memory pruning and on-chip recomputation. InProceedings of the 55th IEEE/ACM International Sym- posium on Microarchitecture (MICRO). 744–762
-
[113]
Zhe Zhou, Junlin Liu, Zhenyu Gu, and Guangyu Sun. 2022. Energon: Toward efficient acceleration of Transformers using dynamic sparse attention.IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems42, 1 (2022), 136–149. 17
2022
-
[118]
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024. Atom: Low- bit quantization for efficient and accurate llm serving.Proceedings of Machine Learning and Systems6 (2024), 196–209
2024
-
[2017]
Automatic differentiation in PyTorch. (2017)
2017
-
[2019]
InProceedings of the 52nd Annual IEEE/ACM International Symposium on Mi- croarchitecture
SparTen: A sparse tensor accelerator for convolutional neural networks. InProceedings of the 52nd Annual IEEE/ACM International Symposium on Mi- croarchitecture. 151–165
-
[2021]
An efficient stochastic convolution architecture based on fast FIR algo- rithm.IEEE Transactions on Circuits and Systems II: Express Briefs69, 3 (2021), 984–988
2021
-
[2022]
InProceedings of the 59th ACM/IEEE Design Automation Conference
SALO: An efficient spatial accelerator enabling hybrid sparse attention- mechanisms for long sequences. InProceedings of the 59th ACM/IEEE Design Automation Conference. 571–576
-
[2023]
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality.See https://vicuna. lmsys. org (accessed 14 April 2023)2, 3 (2023), 6
2023
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.