REVIEW 4 major objections 6 minor 89 references
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Anda shows that replacing FP16 activations with a variable-length, group-shared-exponent format gives 2.4x faster LLM inference, 4.0x better area efficiency, and 3.1x better energy efficiency at near-unchanged perplexity.
desk verdict Anda is a serious hardware-algorithm co-design paper with a genuinely new variable-length BFP format and matching bit-serial architecture; the main soft spot is that accuracy rests on perplexity alone, but that is a fixable gap, not a fatal one. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Anda data type, a block-floating-point (BFP) format with a sign bit, a group-shared exponent (group size 64), and a variable-length mantissa selectable per tensor from 1 to 16 bits. The argument is carried by three pieces working together: (1) the adaptive precision combination search, which reuses the calibration data of post-training weight quantization to choose mantissa widths for the four FP-INT GeMM activation tensors; (2) a bit-plane data layout that stores mantissa bits of the same significance across 64 values together, keeping memory access regular despite variable lengths; and (3) an Anda-enhanced bit-serial processing unit and a runtime bit-plane compressor that convert FP16 outputs to the compressed format on the fly. The shared exponent removes exponent alignment and normalization inside a group, turning FP-INT dot products into integer operations with a single FP32 accumulation across groups.
What would settle it
Measure the exact precision combinations found by the search (e.g., [7,7,6,5] for OPT-125M) on a downstream benchmark such as question answering or reasoning; if task accuracy drops by more than the tolerance while perplexity stays within it, the claimed accuracy-efficiency balance does not transfer. Alternatively, build or simulate the Anda accelerator with the stated HBM2 memory model and verify whether the 2.4x speedup and 3.1x energy efficiency hold at the system level, since the paper's numbers come from a cycle-accurate simulator.
Extended reading notes
Core claim
The paper's central claim is that FP16 activations in W4A16 LLMs can be replaced by variable-length, group-shared-exponent activations with only a small, controlled perplexity change, and that this replacement is what unlocks large efficiency gains in FP-INT GeMM operations. The underlying empirical finding is a sensitivity pattern: different LLMs and different modules within a model tolerate different amounts of mantissa truncation, with Aqkv consistently the most sensitive and the feed-forward down-projection Ad often the least. On this basis the paper defines the Anda format and an iterative module-wise search over the 4-tuple [Mqkv, Mo, Mu, Md] that maximizes BOPs reduction subject to an accuracy-loss tolerance. The claim is not that all activations can be aggressively quantized, but that precision can be assigned per module, and that the resulting format, together with bit-plane memory layout and bit-serial processing, delivers the reported system-level gains.
Load-bearing premise
The load-bearing premise is that perplexity on WikiText2, PTB, and C4 is a faithful enough measure of model quality that keeping it within 0.1% or 1% of the baseline guarantees the same tolerance for real downstream tasks.
Editorial extensions
If this is right
- FP-INT GeMMs, which make up over 90% of operations in sub-4K-token weight-only LLM inference, can be executed as integer dot products with a shared exponent, cutting out per-element exponent alignment.
- A user can choose the accuracy-efficiency operating point after training: relaxing the tolerance from 0.1% to 5% raises the reported speedup from 1.73x to 2.74x and energy efficiency from 2.95x to 3.22x for LLaMA-13B.
- Because the search reuses the calibration data already used for weight-only quantization and runs in at most 32 iterations, Anda slots into existing post-training deployment pipelines without retraining.
- Compared with FIGNA's fixed 14-bit mantissa conversion, Anda cuts bit operations by 1.46x to 2.69x at similar perplexity loss, because different tensors get different mantissa widths.
- Anda's bit-plane storage and on-the-fly compressor reduce SRAM and DRAM access energy by roughly 2.2x and 2.0x over FIGNA, so memory, not just arithmetic, shares in the gain.
Reading between the lines
- If the module-sensitivity pattern generalizes beyond the nine tested models, the 4-tuple search could be replaced by a learned or heuristic prior (Aqkv high precision, Ad low), making deployment even faster.
- Anda's gains depend on the system actually exploiting variable precision; on a workload where all tensors need full mantissas, the bit-serial design is less efficient than bit-parallel FIGNA at fixed width, as the paper itself notes.
- Combining Anda activation compression with KV-cache quantization, which the paper mentions as future work, could extend the same variable-length idea to long-context inference where FP-INT GeMMs are no longer the sole bottleneck.
- The same format and search could be transferred to other transformer workloads such as vision transformers or encoder-only models, but the sensitivity ranking of the four tensor types would need re-measuring.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Anda, a variable-length grouped block-floating-point activation format for weight-only quantized LLMs, together with a post-training adaptive precision search that selects per-module mantissa widths under a user-specified perplexity-loss tolerance, and a hardware architecture (bit-plane memory layout, bit-serial processing units, runtime bit-plane compressor) to exploit the format. Evaluations on nine OPT/LLaMA/LLaMA-2 models and three datasets report perplexity close to the Omniquant weight-only baseline while cutting bit operations, and RTL/cycle-accurate hardware evaluation claims average 2.4x speedup, 4.0x area efficiency, and 3.1x energy efficiency over a GPU-like FP-FP accelerator.
Significance. If the accuracy-efficiency balance holds, this is a significant contribution to efficient LLM inference: it targets the FP-activation bottleneck that weight-only quantization leaves behind, offers a practical training-free search over four module-wise precisions, and backs the proposal with a coherent hardware design and broad model/dataset coverage. The paper honestly discloses occasional tolerance breaches and provides sensitivity analyses across group sizes, models, and modules. The main weakness is that the accuracy side of the central trade-off is validated only through perplexity, with no downstream-task accuracy, and the reported tolerance semantics are not fully pinned down.
major comments (4)
- [Sec. V-A/V-B, Table II] The central accuracy claim is supported only by perplexity on WikiText2, PTB, and C4. The precision combinations in Fig. 14 are selected to keep PPL within tolerance, and the hardware efficiency numbers inherit those mantissa widths. If these PPL-preserving precisions degrade downstream task accuracy (e.g., reasoning or knowledge benchmarks) by more than the user tolerance, the headline accuracy-efficiency trade-off is overstated. Please report at least a few downstream-task benchmarks (e.g., HellaSwag, ARC, BoolQ, MMLU) for a representative subset of models and precisions, or explicitly restrict the claim and the search objective to language-modeling perplexity.
- [Table II and Sec. V-B] The user-set tolerance is not a guarantee on the reported validation metric. Under the 1% constraint, Table II shows LLaMA2-7B at 1.07% loss on WikiText2, LLaMA-13B at 1.16% loss on WikiText2, and OPT-6.7B at 1.01% loss on C4. The text attributes this to calibration/validation drift, but Algorithm 1 enforces the tolerance on the calibration data. If the stated tolerance is meant to hold during actual inference, the search should validate on held-out data or apply a margin; otherwise the abstract's 'within user-set tolerances' should be qualified as calibration-only.
- [Table II] The sign convention and definition of the red 'accuracy drop' percentages are internally inconsistent. For example, on WikiText2 OPT-1.3B, Omniquant gives 14.88 PPL and Ours(1%) gives 14.99 PPL (a 0.74% increase), but the table reports -0.74%; the text in Sec. V-B says '0.74% accuracy loss' for the same case. If the intended formula is (PPL_ref - PPL_method)/PPL_ref, the signs should be positive for these values; if the intended formula is (PPL_ref/PPL_method - 1), the label should be 'relative accuracy' rather than 'accuracy drop'. The current presentation prevents readers from verifying tolerance compliance.
- [Algorithm 1, line 4 and Sec. V-A] The baseline used as fpacc in Algorithm 1 is ambiguous. If L is the weight-only quantized model, then the tolerance is relative to Omniquant and this should be stated explicitly; if fpacc is evaluated on the original FP16 model, then the red percentages in Table II, which are relative to Omniquant, are not on the same baseline as the search constraint. One clarifying sentence is needed to ensure that the reported 0.1%/1% values correspond to the same reference used in the search.
minor comments (6)
- [Sec. III-C and Fig. 9] The text says the search 'efficiently finds the global optimum within 10 iterations,' but the greedy relaxation strategy is acknowledged later to possibly miss the global optimum; please rephrase to 'near-optimal solution' or provide exhaustive-search evidence for the specific case.
- [Sec. III-D] The claim that the search 'operates approximately twice as fast as Omniquant and ten times faster than GPTQ' is not accompanied by measured search-time data; please add timing measurements or remove the quantitative comparison.
- [Sec. V-A] The statement that all hardware baselines are configured with 'equivalent peak throughput' needs a concrete architectural specification (PE counts, array dimensions, dataflow assumptions) so that the system-level comparisons in Fig. 16 are reproducible.
- [Abstract and Sec. V-D] The abstract reports a single 2.4x average speedup, while Sec. V-D reports 2.14x at 0.1% loss and 2.49x at 1% loss; please clarify whether 2.4x is the geometric mean over the two tolerances or the average over all settings.
- [Algorithm 1] There is a typo in line 1: 'P riorityQueue' should read 'PriorityQueue'.
- [References] Reference [6] spells the vendor name as 'Candence'; it should be 'Cadence'.
Circularity Check
No significant circularity: efficiency results rest on external RTL/simulator baselines and are not defined by the paper's own outputs.
full rationale
The paper's central efficiency claims (2.4x speedup, 4.0x area efficiency, 3.1x energy efficiency over the GPU-like FP-FP baseline) are produced by a cycle-accurate simulator and RTL synthesis at 16nm, with hardware baselines (FP-FP, FP-INT, iFPU, FIGNA) configured to the same clock frequency, peak throughput, and on-chip memory resources; these are external, independently specified comparison points, not quantities constructed from the paper's own fitted outputs. The adaptive precision search (Algorithm 1) minimizes BOPs subject to a perplexity constraint on calibration data, and Table II then evaluates perplexity on held-out validation datasets (WikiText2, PTB, C4). Using the same metric for calibration and validation is a common and legitimate evaluation design, and it is not a reduction by construction: the validation perplexity is not forced to satisfy the tolerance, and indeed the paper explicitly acknowledges calibration/validation drift, noting that 'the occasional slight exceedance of the validation accuracy loss over the constraint is normal.' The few Table II entries that exceed the stated 1% tolerance (e.g., LLaMA2-7B at 1.07% and LLaMA-13B at 1.16% on WikiText2) are an accuracy-validation limitation, not evidence of circular reasoning. The self-citations present in the paper (e.g., [68] for the fair-comparison setup, [57] for GPU kernels, [71] for the BOPs metric) are contextual and do not carry the load of the central accuracy-efficiency claim; no uniqueness theorem or design-forcing ansatz is imported from the authors' prior work. The absence of downstream-task accuracy (e.g., MMLU, HellaSwag) is a scope limitation of the evaluation, not a circularity. Overall, the derivation chain is self-contained and no key result reduces by definition to its own inputs.
Assumptions & free parameters
free parameters (2)
- BFP group size (GS) =
64
- Per-module mantissa tuple [M_qkv, M_o, M_u, M_d] =
e.g., WikiText2 1%: OPT-1.3B [6,5,5,4]; values per model/dataset in Fig. 14
assumptions (4)
- domain assumption Perplexity on WikiText2/PTB/C4 is a sufficient proxy for model accuracy, and 1% relative PPL increase is an acceptable loss.
- domain assumption A single precision tuple applied uniformly across all transformer layers preserves accuracy within the tolerance.
- domain assumption The BOPs metric (one FP16-INT4 MAC = 64 BOPs) is a reliable proxy for hardware cost for search prioritization.
- domain assumption The hardware simulator and RTL synthesis correctly capture energy, area, and timing of the Anda architecture and baselines.
Cite this review
Pith. "Pith review of Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format." pith.science (2026). https://pith.science/paper/6TAG52MY
@misc{pith2026241115982,
author = {Pith},
title = {Pith review of: Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format},
year = {2026},
howpublished = {\url{https://pith.science/paper/6TAG52MY}},
note = {Machine review of arXiv:2411.15982}
}
read the original abstract
The widely-used, weight-only quantized large language models (LLMs), which leverage low-bit integer (INT) weights and retain floating-point (FP) activations, reduce storage requirements while maintaining accuracy. However, this shifts the energy and latency bottlenecks towards the FP activations that are associated with costly memory accesses and computations. Existing LLM accelerators focus primarily on computation optimizations, overlooking the potential of jointly optimizing FP computations and data movement, particularly for the dominant FP-INT GeMM operations in LLM inference. To address these challenges, we investigate the sensitivity of activation precision across various LLM modules and its impact on overall model accuracy. Based on our findings, we first propose the Anda data type: an adaptive data format with group-shared exponent bits and dynamic mantissa bit allocation. Secondly, we develop an iterative post-training adaptive precision search algorithm that optimizes the bit-width for different LLM modules to balance model accuracy, energy efficiency, and inference speed. Lastly, a suite of hardware optimization techniques is proposed to maximally exploit the benefits of the Anda format. These include a bit-plane-based data organization scheme, Anda-enhanced processing units with bit-serial computation, and a runtime bit-plane Anda compressor to simultaneously optimize storage, computation, and memory footprints. Our evaluations on FPINT GeMM operations show that Anda achieves a 2.4x speedup, 4.0x area efficiency, and 3.1x energy efficiency improvement on average for popular LLMs including OPT, LLaMA, and LLaMA-2 series over the GPU-like FP-FP baseline. Anda demonstrates strong adaptability across various application scenarios, accuracy requirements, and system performance, enabling efficient LLM inference across a wide range of deployment scenarios.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Resq: Residual quantization for video perception,
D. Abati, H. Ben Yahia, M. Nagel, and A. Habibian, “Resq: Residual quantization for video perception,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 17 119– 17 129
2023
-
[2]
Bit-pragmatic deep neural network computing,
J. Albericio, A. Delm ´as, P. Judd, S. Sharify, G. O’Leary, R. Genov, and A. Moshovos, “Bit-pragmatic deep neural network computing,” in Proceedings of the 50th annual IEEE/ACM international symposium on microarchitecture (MICRO), 2017, pp. 382–394
2017
-
[3]
Explaining neural scaling laws,
Y . Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma, “Explaining neural scaling laws,” Proceedings of the National Academy of Sciences (PNAS), vol. 121, no. 27, p. e2311878121, 2024
2024
-
[4]
Longbench: A bilingual, multitask benchmark for long context understanding,
Y . Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, Y . Dong, J. Tang, and J. Li, “Longbench: A bilingual, multitask benchmark for long context understanding,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL) , 2024, pp. 3119–3137
2024
-
[5]
Demystifying chatgpt: An in-depth survey of openai’s robust large language models,
P. Bhattacharya, V . K. Prasad, A. Verma, D. Gupta, A. Sapsomboon, W. Viriyasitavat, and G. Dhiman, “Demystifying chatgpt: An in-depth survey of openai’s robust large language models,” Archives of Compu- tational Methods in Engineering , pp. 1–44, 2024
2024
-
[6]
Genus synthesis solution,
Candence, “Genus synthesis solution,” https://www.cadence.com/en US/home/tools/digital-design-and-signoff/synthesis/genus-synthesis- solution.html, 2024, online; accessed 2024-07-16
2024
-
[7]
General purpose deep learning accelerator based on bit interleaving,
L. Chang, H. Lu, C. Li, X. Zhao, Z. Hu, J. Zhou, and X. Li, “General purpose deep learning accelerator based on bit interleaving,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD), 2023
2023
-
[8]
Quip: 2-bit quanti- zation of large language models with guarantees,
J. Chee, Y . Cai, V . Kuleshov, and C. M. De Sa, “Quip: 2-bit quanti- zation of large language models with guarantees,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 36, 2024
2024
Show all 89 references
-
[9]
Efficientqat: Efficient quantization-aware training for large language models,
M. Chen, W. Shao, P. Xu, J. Wang, P. Gao, K. Zhang, Y . Qiao, and P. Luo, “Efficientqat: Efficient quantization-aware training for large language models,” arXiv preprint arXiv:2407.11062 , 2024
2024 arXiv
-
[10]
Nacl: A general and effective kv cache eviction framework for llm at inference time,
Y . Chen, G. Wang, J. Shang, S. Cui, Z. Zhang, T. Liu, S. Wang, Y . Sun, D. Yu, and H. Wu, “Nacl: A general and effective kv cache eviction framework for llm at inference time,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume ...
2024
-
[11]
Palm: Scaling language modeling with pathways,
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y . Tay, N. Shazeer, V . Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. ...
2023
-
[12]
Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,
S. Dai, R. Venkatesan, M. Ren, B. Zimmer, W. Dally, and B. Khailany, “Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,” Proceedings of Machine Learning and Sys- tems (MLSys), vol. 3, pp. 873–884, 2021
2021
-
[13]
Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,
B. Darvish Rouhani, D. Lo, R. Zhao, M. Liu, J. Fowers, K. Ovtcharov, A. Vinogradsky, S. Massengill, L. Yang, R. Bittner, A. Forin, H. Zhu, T. Na, P. Patel, S. Che, L. Chand Koppaka, X. Song, S. Som, K. Das, S. T, S. Reinhardt, S. Lanka, E. Chung, and D. Burger, “Pushing the li...
2020
-
[14]
With shared microexponents, a little shifting goes a long way,
B. Darvish Rouhani, R. Zhao, V . Elango, R. Shafipour, M. Hall, M. Mesmakhosroshahi, A. More, L. Melnick, M. Golub, G. Varatkar, L. Shao, G. Kolhe, D. Melts, J. Klar, R. L’Heureux, M. Perry, D. Burger, E. Chung, Z. S. Deng, S. Naghshineh, J. Park, and M. Naumov, “With shared m...
2023
-
[15]
A timing-driven approach to synthesize fast barrel shifters,
S. Das and S. P. Khatri, “A timing-driven approach to synthesize fast barrel shifters,” IEEE Transactions on Circuits and Systems II: Express Briefs (TCAS-II), vol. 55, no. 1, pp. 31–35, 2008
2008
-
[16]
Llm.int8(): 8- bit matrix multiplication for transformers at scale,
T. Dettmers, M. Lewis, Y . Belkada, and L. Zettlemoyer, “Llm.int8(): 8- bit matrix multiplication for transformers at scale,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 35, pp. 30 318–30 332, 2022
2022
-
[17]
The case for 4-bit precision: k- bit inference scaling laws,
T. Dettmers and L. Zettlemoyer, “The case for 4-bit precision: k- bit inference scaling laws,” in International Conference on Machine Learning (ICML). PMLR, 2023, pp. 7750–7774
2023
-
[18]
Hawq: Hessian aware quantization of neural networks with mixed-precision,
Z. Dong, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Hawq: Hessian aware quantization of neural networks with mixed-precision,” in Proceedings of the IEEE/CVF international conference on computer vision (ICCV), 2019, pp. 293–302
2019
-
[19]
Training dnns with hybrid block floating point,
M. Drumond, T. Lin, M. Jaggi, and B. Falsafi, “Training dnns with hybrid block floating point,” Advances in Neural Information Processing Systems (NeurIPS), vol. 31, 2018
2018
-
[20]
Skvq: Sliding-window key and value cache quantization for large language models,
H. Duanmu, Z. Yuan, X. Li, J. Duan, X. Zhang, and D. Lin, “Skvq: Sliding-window key and value cache quantization for large language models,” in First Conference on Language Modeling (COLM) , 2024
2024
-
[21]
Extreme compression of large language models via additive quantization,
V . Egiazarian, A. Panferov, D. Kuznedelev, E. Frantar, A. Babenko, and D. Alistarh, “Extreme compression of large language models via additive quantization,” inInternational Conference on Machine Learning (ICML). PMLR, 2024
2024
-
[22]
Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,
H. Fan, H.-C. Ng, S. Liu, Z. Que, X. Niu, and W. Luk, “Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,” in 28th International Conference on Field Programmable Logic and Applications (FPL) . IEEE, 2018, pp. 287– 2877
2018
-
[23]
Static block floating-point quantization for convolutional neural networks on fpga,
H. Fan, G. Wang, M. Ferianc, X. Niu, and W. Luk, “Static block floating-point quantization for convolutional neural networks on fpga,” in International Conference on Field-Programmable Technology (ICFPT) . IEEE, 2019, pp. 28–35
2019
-
[24]
Optq: Accurate quantization for generative pre-trained transformers,
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “Optq: Accurate quantization for generative pre-trained transformers,” in The Eleventh International Conference on Learning Representations (ICLR) , 2023
2023
-
[25]
Llmc: Benchmarking large language model quantization with a versatile compression toolkit,
R. Gong, Y . Yong, S. Gu, Y . Huang, Y . Zhang, X. Liu, and D. Tao, “Llmc: Benchmarking large language model quantization with a versatile compression toolkit,” arXiv preprint arXiv:2405.06001 , 2024
2024 arXiv
-
[26]
Boost: block minifloat-based on-device cnn training accelerator with transfer learning,
C. Guo, B. Lou, X. Liu, D. Boland, P. H. Leong, and C. Zhuo, “Boost: block minifloat-based on-device cnn training accelerator with transfer learning,” in IEEE/ACM International Conference on Computer Aided Design (ICCAD). IEEE, 2023, pp. 1–9
2023
-
[27]
Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,
C. Guo, J. Tang, W. Hu, J. Leng, C. Zhang, F. Yang, Y . Liu, M. Guo, and Y . Zhu, “Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,” in Proceedings of the 50th Annual International Symposium on Computer Architecture (ISCA) , 20...
2023
-
[28]
Ese: Efficient speech recognition engine with sparse lstm on fpga,
S. Han, J. Kang, H. Mao, Y . Hu, X. Li, Y . Li, D. Xie, H. Luo, S. Yao, Y . Wang, H. Yang, and W. B. J. Dally, “Ese: Efficient speech recognition engine with sparse lstm on fpga,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (F...
2017
-
[29]
Kvquant: Towards 10 million context length llm inference with kv cache quantization,
C. Hooper, S. Kim, H. Mohammadzadeh, M. W. Mahoney, Y . S. Shao, K. Keutzer, and A. Gholami, “Kvquant: Towards 10 million context length llm inference with kv cache quantization,” arXiv preprint arXiv:2401.18079, 2024
2024 arXiv
-
[30]
A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,
L. Huang, C. Fang, Q. Li, J. Lin, and Z. Wang, “A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,” in 29th Asia and South Pacific Design Automation Conference (ASP-DAC). IEEE, 2024, pp. 927–932
2024
-
[31]
Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,
Q. Huang, P.-A. Tsai, J. S. Emer, and A. Parashar, “Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA) , 2024
2024
-
[32]
Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,
J. Jang, Y . Kim, J. Lee, and J.-J. Kim, “Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,” in IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2024, pp. 760–773
2024
-
[33]
Perplexity—a measure of the difficulty of speech recognition tasks,
F. Jelinek, R. L. Mercer, L. R. Bahl, and J. K. Baker, “Perplexity—a measure of the difficulty of speech recognition tasks,” The Journal of the Acoustical Society of America (JASA) , vol. 62, no. S1, pp. S63–S63, 1977
1977
-
[34]
Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,
Y . Jeon, C. Lee, E. Cho, and Y . Ro, “Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 12 329–12 338
2022
-
[35]
Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,
Y . Jeon, B. Park, S. J. Kwon, B. Kim, J. Yun, and D. Lee, “Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,” in International Conference for High Performance Comput- ing, Networking, Storage and Analysis (SC) . IEEE, 2020, pp. 1–14
2020
-
[36]
Ten lessons from three generations shaped google’s tpuv4i: Industrial product,
N. P. Jouppi, D. Hyun Yoon, M. Ashcraft, M. Gottscho, T. B. Jablin, G. Kurian, J. Laudon, S. Li, P. Ma, X. Ma, T. Norrie, N. Patil, S. Prasad, C. Young, Z. Zhou, and D. Patterson, “Ten lessons from three generations shaped google’s tpuv4i: Industrial product,” in ACM/IEEE 48th...
2021
-
[37]
Stripes: Bit-serial deep neural network computing,
P. Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, “Stripes: Bit-serial deep neural network computing,” in 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2016, pp. 1–12
2016
-
[38]
A survey of gpt-3 family large language models including chatgpt and gpt-4,
K. S. Kalyan, “A survey of gpt-3 family large language models including chatgpt and gpt-4,” Natural Language Processing Journal , p. 100048, 2023
2023
-
[39]
A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,
B. Keller, R. Venkatesan, S. Dai, S. G. Tell, B. Zimmer, C. Sakr, W. J. Dally, C. T. Gray, and B. Khailany, “A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,” IEEE Journal of Solid-State Circuits (JSSC) , vol. 58, no. 4, pp. ...
2023
-
[40]
Compressed context mem- ory for online language model interaction,
J.-H. Kim, J. Yeom, S. Yun, and H. O. Song, “Compressed context mem- ory for online language model interaction,” in The Twelfth International Conference on Learning Representations (ICLR) , 2024
2024
-
[41]
Dacapo: Accelerating continuous learning in autonomous systems for video analytics,
Y . Kim, C. Oh, J. Hwang, W. Kim, S. Oh, Y . Lee, H. Sharma, A. Yaz- danbakhsh, and J. Park, “Dacapo: Accelerating continuous learning in autonomous systems for video analytics,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA) , 2024
2024
-
[42]
Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,
Y . Kim, J. Jang, J. Lee, J. Park, J. Kim, B. Kim, B. park, S. J. Kwon, D. Lee, and J.-J. Kim, “Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,” in The Eleventh International Conference on Learning Representations (ICLR) , 2023
2023
-
[43]
One-shot model for mixed-precision quantization,
I. Koryakovskiy, A. Yakovleva, V . Buchnev, T. Isaev, and G. Odinokikh, “One-shot model for mixed-precision quantization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7939–7949
2023
-
[44]
Flexpoint: An adaptive numerical format for efficient training of deep neural networks,
U. K ¨oster, T. J. Webb, X. Wang, M. Nassar, A. K. Bansal, W. H. Constable, O. H. Elibol, S. Gray, S. Hall, L. Hornof, A. Khosrowshahi, C. Kloss, R. J. Pai, and N. Rao, “Flexpoint: An adaptive numerical format for efficient training of deep neural networks,” Advances in Neural...
2017
-
[45]
Tender: Accelerating large language models via tensor decomposition and runtime requantization,
J. Lee, W. Lee, and J. Sim, “Tender: Accelerating large language models via tensor decomposition and runtime requantization,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA), 2024
2024
-
[46]
Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,
A. Li, H. Mo, W. Zhu, Q. Li, S. Yin, S. Wei, and L. Liu, “Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) , vol. 41, no. 11, pp. 4747– 4...
2022
-
[47]
Norm tweaking: High-performance low-bit quantization of large language models,
L. Li, Q. Li, B. Zhang, and X. Chu, “Norm tweaking: High-performance low-bit quantization of large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 38, no. 17, 2024, pp. 18 536–18 544
2024
-
[48]
Geo: Generation and execution optimized stochastic computing accelerator for neural networks,
T. Li, W. Romaszkan, S. Pamarti, and P. Gupta, “Geo: Generation and execution optimized stochastic computing accelerator for neural networks,” in 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 2021, pp. 689–694
2021
-
[49]
Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,
Z. Li, A. Lu, Y . Xie, Z. Kong, M. Sun, H. Tang, Z. J. Xue, P. Dong, C. Ding, Y . Wang, X. Lin, and Z. Fang, “Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,” in Pro- ceedings of the 38th ACM International Conference on Supercomput...
2024
-
[50]
High-performance fpga-based cnn accelerator with block-floating-point arithmetic,
X. Lian, Z. Liu, Z. Song, J. Dai, W. Zhou, and X. Ji, “High-performance fpga-based cnn accelerator with block-floating-point arithmetic,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI) , vol. 27, no. 8, pp. 1874–1885, 2019
2019
-
[51]
Awq: Activation-aware weight quan- tization for llm compression and acceleration,
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quan- tization for llm compression and acceleration,” in The Seventh Annual Conference on Machine Learning and Systems (MLSys) , 2024
2024
-
[52]
Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,
Y . Lin, H. Tang, S. Yang, Z. Zhang, G. Xiao, C. Gan, and S. Han, “Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,” arXiv preprint arXiv:2405.04532 , 2024
2024 arXiv
-
[53]
Llm-qat: Data-free quantization aware training for large language models,
Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y . Mehdad, Y . Shi, R. Kr- ishnamoorthi, and V . Chandra, “Llm-qat: Data-free quantization aware training for large language models,” arXiv preprint arXiv:2305.17888 , 2023
2023 arXiv
-
[54]
Kivi: A tuning-free asymmetric 2bit quantization for kv cache,
Z. Liu, J. Yuan, H. Jin, S. Zhong, Z. Xu, V . Braverman, B. Chen, and X. Hu, “Kivi: A tuning-free asymmetric 2bit quantization for kv cache,” in Forty-first International Conference on Machine Learning (ICML) , 2024
2024
-
[55]
Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,
H. Lu, L. Chang, C. Li, Z. Zhu, S. Lu, Y . Liu, and M. Zhang, “Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,” in 54th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2021, pp. 963–976
2021
-
[56]
Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,
S. Luohe, H. Zhang, Y . Yao, Z. Li et al. , “Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,” in First Conference on Language Modeling (COLM) , 2024
2024
-
[57]
Efficient arbitrary precision acceleration for large language models on gpu tensor cores,
S. Ma, C. Fang, H. Shao, and Z. Wang, “Efficient arbitrary precision acceleration for large language models on gpu tensor cores,” arXiv preprint arXiv:2409.17870, 2024
2024 arXiv
-
[58]
Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,
S. Mach, F. Schuiki, F. Zaruba, and L. Benini, “Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI) , vol. 29, no. 4, pp. 774–787, 2020
2020
-
[59]
The penn treebank: Anno- tating predicate argument structure,
M. Marcus, G. Kim, M. A. Marcinkiewicz, R. MacIntyre, A. Bies, M. Ferguson, K. Katz, and B. Schasberger, “The penn treebank: Anno- tating predicate argument structure,” in Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994, 1994
1994
-
[60]
Pointer sentinel mix- ture models,
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mix- ture models,” in International Conference on Learning Representations (ICLR), 2017
2017
-
[61]
Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,
S.-H. Noh, J. Koo, S. Lee, J. Park, and J. Kung, “Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,” IEEE Transactions on Computers (TC) , vol. 72, no. 9, pp. 2522–2535, 2023
2023
-
[62]
Cutlass,
NVIDIA, “Cutlass,” https://github.com/NVIDIA/cutlass, 2024, online; accessed 2024-07-03
2024
-
[63]
Gpt- 4 technical report,
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V . Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L...
2023 arXiv
-
[64]
LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,
G. Park, B. park, M. Kim, S. Lee, J. Kim, B. Kwon, S. J. Kwon, B. Kim, Y . Lee, and D. Lee, “LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,” in The Twelfth International Conference on Learning Repres...
2024
-
[65]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research (JMLR), vol. 21, no. 140, pp. 1–67, 2020
2020
-
[66]
Omniquant: Omnidirectionally calibrated quantization for large language models,
W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, K. Zhang, P. Gao, Y . Qiao, and P. Luo, “Omniquant: Omnidirectionally calibrated quantization for large language models,” in The Twelfth International Conference on Learning Representations (ICLR) , 2024
2024
-
[67]
Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,
H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, J. K. Kim, V . Chandra, and H. Esmaeilzadeh, “Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,” in Proceedings of the ACM/IEEE 45th Annual International Symposium on Computer Architect...
2018
-
[68]
Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,
M. Shi, V . Jain, A. Joseph, M. Meijer, and M. Verhelst, “Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,” in IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2024, pp. 732–746
2024
-
[69]
Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,
W. Sun, A. Li, T. Geng, S. Stuijk, and H. Corporaal, “Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,” IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 34, no. 1, pp. 246–261, 2022
2022
-
[70]
Gemma: Open models based on gemini research and technology,
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Riviere, M. Kale, J. C. Love, P. D. Tafti, L. Hussenot, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. H’eliou, A. Tacchetti, A. Bulanova, A. Paterson, B. Tsai, B. Sh...
2024 arXiv
-
[71]
Bebert: Efficient and robust binary ensemble bert,
J. Tian, C. Fang, H. Wang, and Z. Wang, “Bebert: Efficient and robust binary ensemble bert,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
-
[72]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971 , 2023
2023 arXiv
-
[73]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Ho...
2023 arXiv
-
[74]
Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,
A. Tseng, J. Chee, Q. Sun, V . Kuleshov, and C. De Sa, “Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,” arXiv preprint arXiv:2402.04396 , 2024
2024 arXiv
-
[75]
Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,
G. Wang, S. Cai, W. Li, D. Lyu, and G. He, “Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,” IEEE Transactions on Circuits and Systems I: Regular Papers (TCAS-I), 2024
2024
-
[76]
Haq: Hardware-aware automated quantization with mixed precision,
K. Wang, Z. Liu, Y . Lin, J. Lin, and S. Han, “Haq: Hardware-aware automated quantization with mixed precision,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), 2019, pp. 8612–8620
2019
-
[77]
Outlier suppression: Pushing the limit of low-bit transformer language models,
X. Wei, Y . Zhang, X. Zhang, R. Gong, S. Zhang, Q. Zhang, F. Yu, and X. Liu, “Outlier suppression: Pushing the limit of low-bit transformer language models,” Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 17 402–17 414, 2022
2022
-
[78]
Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,
H. Xia, Z. Zheng, X. Wu, S. Chen, Z. Yao, S. Youn, A. Bakhtiari, M. Wyatt, D. Zhuang, Z. Zhou, O. Ruwase, Y . He, and S. L. Song, “Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,” in 2024 USENIX Annual Te...
2024
-
[79]
Smoothquant: Accurate and efficient post-training quantization for large language models,
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning (ICML). PMLR, 2023, pp. 38 087–38 099
2023
-
[80]
Efficient streaming language models with attention sinks,
G. Xiao, Y . Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in The Twelfth International Conference on Learning Representations (ICLR) , 2024
2024
-
[81]
Onebit: Towards extremely low-bit large language models,
Y . Xu, X. Han, Z. Yang, S. Wang, Q. Zhu, Z. Liu, W. Liu, and W. Che, “Onebit: Towards extremely low-bit large language models,” arXiv preprint arXiv:2402.11295 , 2024
2024 arXiv
-
[82]
Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,
J. Yuan, H. Liu, S. Zhong, Y .-N. Chuang, S. Li, G. Wang, D. Le, H. Jin, V . Chaudhary, Z. Xu, Z. Liu, and X. Hu, “Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,” in The 2024 Conference on Empirical Methods ...
2024
-
[83]
Llm inference unveiled: Survey and roofline model insights,
Z. Yuan, Y . Shang, Y . Zhou, Z. Dong, Z. Zhou, C. Xue, B. Wu, Z. Li, Q. Gu, Y . J. Lee, Y . Yan, B. Chen, G. Sun, and K. Keutzer, “Llm inference unveiled: Survey and roofline model insights,” arXiv preprint arXiv:2402.16363, 2024
2024 arXiv
-
[84]
Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,
A. H. Zadeh, M. Mahmoud, A. Abdelhadi, and A. Moshovos, “Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,” in Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA) , 2022, pp. 888–901
2022
-
[85]
Fast: Dnn training under variable precision block floating point with stochastic rounding,
S. Q. Zhang, B. McDanel, and H. Kung, “Fast: Dnn training under variable precision block floating point with stochastic rounding,” inIEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2022, pp. 846–860
2022
-
[86]
Opt: Open pre-trained transformer language models,
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V . Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer, “Opt: Open pre-trained transformer language models,” arXiv preprint ...
2022 arXiv
-
[87]
Cam: Cache merging for memory-efficient llms inference,
Y . Zhang, Y . Du, G. Luo, Y . Zhong, Z. Zhang, S. Liu, and R. Ji, “Cam: Cache merging for memory-efficient llms inference,” in Forty- first International Conference on Machine Learning (ICML) , 2024
2024
-
[88]
H2o: Heavy-hitter oracle for efficient generative inference of large language models,
Z. Zhang, Y . Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y . Tian, C. R´e, C. Barrett et al., “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 36, 2024
2024
-
[89]
Atom: Low-bit quantization for efficient and accurate llm serving,
Y . Zhao, C.-Y . Lin, K. Zhu, Z. Ye, L. Chen, S. Zheng, L. Ceze, A. Krishnamurthy, T. Chen, and B. Kasikci, “Atom: Low-bit quantization for efficient and accurate llm serving,” Proceedings of Machine Learning and Systems (MLSys) , vol. 6, pp. 196–209, 2024
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.