Pith. sign in

REVIEW 4 major objections 6 minor 89 references

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Anda shows that replacing FP16 activations with a variable-length, group-shared-exponent format gives 2.4x faster LLM inference, 4.0x better area efficiency, and 3.1x better energy efficiency at near-unchanged perplexity.

desk verdict Anda is a serious hardware-algorithm co-design paper with a genuinely new variable-length BFP format and matching bit-serial architecture; the main soft spot is that accuracy rests on perplexity alone, but that is a fixable gap, not a fatal one. read the letter →

arxiv 2411.15982 v1 pith:6TAG52MY submitted 2024-11-24 cs.AR cs.AIcs.LG

classification cs.ARcs.AIcs.LG
keywords LLMinferenceactivationquantizationblockfloatingpointFP-INTGeMMbit-serialprocessingpost-traininghardwareacceleratorperplexity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to show that the floating-point activations in weight-only quantized large language models are the real efficiency bottleneck, and that they can be compressed and accelerated without retraining. It proposes Anda, a block-floating-point data format in which a group of activations shares one exponent and each activation tensor chooses its own mantissa width, ranging from 1 to 16 bits. A training-free search over just four tensor types (Aqkv, Ao, Au, Ad) finds precision combinations that keep perplexity within a user-set tolerance (0.1% or 1%) while minimizing bit operations. The paper reports that on OPT, LLaMA, and LLaMA-2 models this yields a 2.4x average speedup, 4.0x area efficiency, and 3.1x energy efficiency over a GPU-like FP-FP baseline, with the same models and tolerances. The significance, if correct, is that weight-only quantized LLMs can be deployed faster and cheaper on custom hardware without the costly retraining that earlier block-floating-point methods required.

What carries the argument

The central object is the Anda data type, a block-floating-point (BFP) format with a sign bit, a group-shared exponent (group size 64), and a variable-length mantissa selectable per tensor from 1 to 16 bits. The argument is carried by three pieces working together: (1) the adaptive precision combination search, which reuses the calibration data of post-training weight quantization to choose mantissa widths for the four FP-INT GeMM activation tensors; (2) a bit-plane data layout that stores mantissa bits of the same significance across 64 values together, keeping memory access regular despite variable lengths; and (3) an Anda-enhanced bit-serial processing unit and a runtime bit-plane compressor that convert FP16 outputs to the compressed format on the fly. The shared exponent removes exponent alignment and normalization inside a group, turning FP-INT dot products into integer operations with a single FP32 accumulation across groups.

What would settle it

Measure the exact precision combinations found by the search (e.g., [7,7,6,5] for OPT-125M) on a downstream benchmark such as question answering or reasoning; if task accuracy drops by more than the tolerance while perplexity stays within it, the claimed accuracy-efficiency balance does not transfer. Alternatively, build or simulate the Anda accelerator with the stated HBM2 memory model and verify whether the 2.4x speedup and 3.1x energy efficiency hold at the system level, since the paper's numbers come from a cycle-accurate simulator.

Watch

Extended reading notes

Core claim

The paper's central claim is that FP16 activations in W4A16 LLMs can be replaced by variable-length, group-shared-exponent activations with only a small, controlled perplexity change, and that this replacement is what unlocks large efficiency gains in FP-INT GeMM operations. The underlying empirical finding is a sensitivity pattern: different LLMs and different modules within a model tolerate different amounts of mantissa truncation, with Aqkv consistently the most sensitive and the feed-forward down-projection Ad often the least. On this basis the paper defines the Anda format and an iterative module-wise search over the 4-tuple [Mqkv, Mo, Mu, Md] that maximizes BOPs reduction subject to an accuracy-loss tolerance. The claim is not that all activations can be aggressively quantized, but that precision can be assigned per module, and that the resulting format, together with bit-plane memory layout and bit-serial processing, delivers the reported system-level gains.

Load-bearing premise

The load-bearing premise is that perplexity on WikiText2, PTB, and C4 is a faithful enough measure of model quality that keeping it within 0.1% or 1% of the baseline guarantees the same tolerance for real downstream tasks.

Editorial extensions

If this is right

  • FP-INT GeMMs, which make up over 90% of operations in sub-4K-token weight-only LLM inference, can be executed as integer dot products with a shared exponent, cutting out per-element exponent alignment.
  • A user can choose the accuracy-efficiency operating point after training: relaxing the tolerance from 0.1% to 5% raises the reported speedup from 1.73x to 2.74x and energy efficiency from 2.95x to 3.22x for LLaMA-13B.
  • Because the search reuses the calibration data already used for weight-only quantization and runs in at most 32 iterations, Anda slots into existing post-training deployment pipelines without retraining.
  • Compared with FIGNA's fixed 14-bit mantissa conversion, Anda cuts bit operations by 1.46x to 2.69x at similar perplexity loss, because different tensors get different mantissa widths.
  • Anda's bit-plane storage and on-the-fly compressor reduce SRAM and DRAM access energy by roughly 2.2x and 2.0x over FIGNA, so memory, not just arithmetic, shares in the gain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the module-sensitivity pattern generalizes beyond the nine tested models, the 4-tuple search could be replaced by a learned or heuristic prior (Aqkv high precision, Ad low), making deployment even faster.
  • Anda's gains depend on the system actually exploiting variable precision; on a workload where all tensors need full mantissas, the bit-serial design is less efficient than bit-parallel FIGNA at fixed width, as the paper itself notes.
  • Combining Anda activation compression with KV-cache quantization, which the paper mentions as future work, could extend the same variable-length idea to long-context inference where FP-INT GeMMs are no longer the sole bottleneck.
  • The same format and search could be transferred to other transformer workloads such as vision transformers or encoder-only models, but the sensitivity ranking of the four tensor types would need re-measuring.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes Anda, a variable-length grouped block-floating-point activation format for weight-only quantized LLMs, together with a post-training adaptive precision search that selects per-module mantissa widths under a user-specified perplexity-loss tolerance, and a hardware architecture (bit-plane memory layout, bit-serial processing units, runtime bit-plane compressor) to exploit the format. Evaluations on nine OPT/LLaMA/LLaMA-2 models and three datasets report perplexity close to the Omniquant weight-only baseline while cutting bit operations, and RTL/cycle-accurate hardware evaluation claims average 2.4x speedup, 4.0x area efficiency, and 3.1x energy efficiency over a GPU-like FP-FP accelerator.

Significance. If the accuracy-efficiency balance holds, this is a significant contribution to efficient LLM inference: it targets the FP-activation bottleneck that weight-only quantization leaves behind, offers a practical training-free search over four module-wise precisions, and backs the proposal with a coherent hardware design and broad model/dataset coverage. The paper honestly discloses occasional tolerance breaches and provides sensitivity analyses across group sizes, models, and modules. The main weakness is that the accuracy side of the central trade-off is validated only through perplexity, with no downstream-task accuracy, and the reported tolerance semantics are not fully pinned down.

major comments (4)
  1. [Sec. V-A/V-B, Table II] The central accuracy claim is supported only by perplexity on WikiText2, PTB, and C4. The precision combinations in Fig. 14 are selected to keep PPL within tolerance, and the hardware efficiency numbers inherit those mantissa widths. If these PPL-preserving precisions degrade downstream task accuracy (e.g., reasoning or knowledge benchmarks) by more than the user tolerance, the headline accuracy-efficiency trade-off is overstated. Please report at least a few downstream-task benchmarks (e.g., HellaSwag, ARC, BoolQ, MMLU) for a representative subset of models and precisions, or explicitly restrict the claim and the search objective to language-modeling perplexity.
  2. [Table II and Sec. V-B] The user-set tolerance is not a guarantee on the reported validation metric. Under the 1% constraint, Table II shows LLaMA2-7B at 1.07% loss on WikiText2, LLaMA-13B at 1.16% loss on WikiText2, and OPT-6.7B at 1.01% loss on C4. The text attributes this to calibration/validation drift, but Algorithm 1 enforces the tolerance on the calibration data. If the stated tolerance is meant to hold during actual inference, the search should validate on held-out data or apply a margin; otherwise the abstract's 'within user-set tolerances' should be qualified as calibration-only.
  3. [Table II] The sign convention and definition of the red 'accuracy drop' percentages are internally inconsistent. For example, on WikiText2 OPT-1.3B, Omniquant gives 14.88 PPL and Ours(1%) gives 14.99 PPL (a 0.74% increase), but the table reports -0.74%; the text in Sec. V-B says '0.74% accuracy loss' for the same case. If the intended formula is (PPL_ref - PPL_method)/PPL_ref, the signs should be positive for these values; if the intended formula is (PPL_ref/PPL_method - 1), the label should be 'relative accuracy' rather than 'accuracy drop'. The current presentation prevents readers from verifying tolerance compliance.
  4. [Algorithm 1, line 4 and Sec. V-A] The baseline used as fpacc in Algorithm 1 is ambiguous. If L is the weight-only quantized model, then the tolerance is relative to Omniquant and this should be stated explicitly; if fpacc is evaluated on the original FP16 model, then the red percentages in Table II, which are relative to Omniquant, are not on the same baseline as the search constraint. One clarifying sentence is needed to ensure that the reported 0.1%/1% values correspond to the same reference used in the search.
minor comments (6)
  1. [Sec. III-C and Fig. 9] The text says the search 'efficiently finds the global optimum within 10 iterations,' but the greedy relaxation strategy is acknowledged later to possibly miss the global optimum; please rephrase to 'near-optimal solution' or provide exhaustive-search evidence for the specific case.
  2. [Sec. III-D] The claim that the search 'operates approximately twice as fast as Omniquant and ten times faster than GPTQ' is not accompanied by measured search-time data; please add timing measurements or remove the quantitative comparison.
  3. [Sec. V-A] The statement that all hardware baselines are configured with 'equivalent peak throughput' needs a concrete architectural specification (PE counts, array dimensions, dataflow assumptions) so that the system-level comparisons in Fig. 16 are reproducible.
  4. [Abstract and Sec. V-D] The abstract reports a single 2.4x average speedup, while Sec. V-D reports 2.14x at 0.1% loss and 2.49x at 1% loss; please clarify whether 2.4x is the geometric mean over the two tolerances or the average over all settings.
  5. [Algorithm 1] There is a typo in line 1: 'P riorityQueue' should read 'PriorityQueue'.
  6. [References] Reference [6] spells the vendor name as 'Candence'; it should be 'Cadence'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: efficiency results rest on external RTL/simulator baselines and are not defined by the paper's own outputs.

full rationale

The paper's central efficiency claims (2.4x speedup, 4.0x area efficiency, 3.1x energy efficiency over the GPU-like FP-FP baseline) are produced by a cycle-accurate simulator and RTL synthesis at 16nm, with hardware baselines (FP-FP, FP-INT, iFPU, FIGNA) configured to the same clock frequency, peak throughput, and on-chip memory resources; these are external, independently specified comparison points, not quantities constructed from the paper's own fitted outputs. The adaptive precision search (Algorithm 1) minimizes BOPs subject to a perplexity constraint on calibration data, and Table II then evaluates perplexity on held-out validation datasets (WikiText2, PTB, C4). Using the same metric for calibration and validation is a common and legitimate evaluation design, and it is not a reduction by construction: the validation perplexity is not forced to satisfy the tolerance, and indeed the paper explicitly acknowledges calibration/validation drift, noting that 'the occasional slight exceedance of the validation accuracy loss over the constraint is normal.' The few Table II entries that exceed the stated 1% tolerance (e.g., LLaMA2-7B at 1.07% and LLaMA-13B at 1.16% on WikiText2) are an accuracy-validation limitation, not evidence of circular reasoning. The self-citations present in the paper (e.g., [68] for the fair-comparison setup, [57] for GPU kernels, [71] for the BOPs metric) are contextual and do not carry the load of the central accuracy-efficiency claim; no uniqueness theorem or design-forcing ansatz is imported from the authors' prior work. The absence of downstream-task accuracy (e.g., MMLU, HellaSwag) is a scope limitation of the evaluation, not a circularity. Overall, the derivation chain is self-contained and no key result reduces by definition to its own inputs.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central results depend on two fitted choices (group size and searched precision tuples) and several domain assumptions about the evaluation methodology. No speculative physical entities are introduced; the Anda data format is a concrete designed artifact implemented in RTL and evaluated on real model outputs.

free parameters (2)
  • BFP group size (GS) = 64
    Chosen manually from Fig. 5 as a balance between accuracy and computational efficiency; not optimized per model.
  • Per-module mantissa tuple [M_qkv, M_o, M_u, M_d] = e.g., WikiText2 1%: OPT-1.3B [6,5,5,4]; values per model/dataset in Fig. 14
    Produced by Algorithm 1 to keep relative PPL within tolerance; these exact values determine the reported speedup and energy numbers.
assumptions (4)
  • domain assumption Perplexity on WikiText2/PTB/C4 is a sufficient proxy for model accuracy, and 1% relative PPL increase is an acceptable loss.
    All accuracy constraints and comparisons in Sec. V-B and Table II are expressed as relative perplexity; no task-level benchmarks are run.
  • domain assumption A single precision tuple applied uniformly across all transformer layers preserves accuracy within the tolerance.
    Algorithm 1 searches only over four tensor types and applies the chosen widths to every layer (Sec. III-C). The paper does not measure layer-level sensitivity.
  • domain assumption The BOPs metric (one FP16-INT4 MAC = 64 BOPs) is a reliable proxy for hardware cost for search prioritization.
    Used in Sec. III-C/III-D to rank precision combinations; it correlates with bit width but not with memory access or control overhead.
  • domain assumption The hardware simulator and RTL synthesis correctly capture energy, area, and timing of the Anda architecture and baselines.
    All system-level numbers depend on the in-house simulator described in Sec. V-A; no independent silicon measurements are provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format." pith.science (2026). https://pith.science/paper/6TAG52MY

@misc{pith2026241115982,
  author       = {Pith},
  title        = {Pith review of: Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6TAG52MY}},
  note         = {Machine review of arXiv:2411.15982}
}
read the original abstract

The widely-used, weight-only quantized large language models (LLMs), which leverage low-bit integer (INT) weights and retain floating-point (FP) activations, reduce storage requirements while maintaining accuracy. However, this shifts the energy and latency bottlenecks towards the FP activations that are associated with costly memory accesses and computations. Existing LLM accelerators focus primarily on computation optimizations, overlooking the potential of jointly optimizing FP computations and data movement, particularly for the dominant FP-INT GeMM operations in LLM inference. To address these challenges, we investigate the sensitivity of activation precision across various LLM modules and its impact on overall model accuracy. Based on our findings, we first propose the Anda data type: an adaptive data format with group-shared exponent bits and dynamic mantissa bit allocation. Secondly, we develop an iterative post-training adaptive precision search algorithm that optimizes the bit-width for different LLM modules to balance model accuracy, energy efficiency, and inference speed. Lastly, a suite of hardware optimization techniques is proposed to maximally exploit the benefits of the Anda format. These include a bit-plane-based data organization scheme, Anda-enhanced processing units with bit-serial computation, and a runtime bit-plane Anda compressor to simultaneously optimize storage, computation, and memory footprints. Our evaluations on FPINT GeMM operations show that Anda achieves a 2.4x speedup, 4.0x area efficiency, and 3.1x energy efficiency improvement on average for popular LLMs including OPT, LLaMA, and LLaMA-2 series over the GPU-like FP-FP baseline. Anda demonstrates strong adaptability across various application scenarios, accuracy requirements, and system performance, enabling efficient LLM inference across a wide range of deployment scenarios.

Figures

Figures reproduced from arXiv: 2411.15982 by the authors.

Figure 1
Figure 1. Overview of the drop-in replacement for FP activations using the variable-length grouped Anda data type via a one-shot offline calibration process. This enables online variable-precision LLM inference, significantly improving speed and energy efficiency through the adaptive precision combi￾nation search algorithm and the Anda-aware architecture. OPT-1.3B OPT-2.7B OPT-6.7B LLaMA-7B LLaMA2-7BOPT-13B LLaMA-13B LLaMA2-1… view at source ↗
Figure 2
Figure 2. Proportion of FP-INT GeMM operations in weight-only quantized LLMs across varying model sizes and context lengths for text generation tasks. FP-INT GeMMs dominate (>90%) in prevalent sub-4K token applications and remain significant for 10K+ sequences. To address these challenges, quantization techniques [12], [16], [24], [51], [52], [66], [78] have been widely adopted in LLMs to reduce memory footprint and lower dep… view at source ↗
Figure 3
Figure 3. Illustration of the architecture for a weight-only quantized LLM model. [PITH_FULL_IMAGE:figures/full_fig_p002_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: The process of converting a set of FP16 numbers into different BFP [PITH_FULL_IMAGE:figures/full_fig_p003_4.png]
Figure 6
Figure 6. Figure 6: The relative accuracy to preserved mantissa bits across various LLMs. [PITH_FULL_IMAGE:figures/full_fig_p004_6.png]
Figure 7
Figure 7. Figure 7: The relative accuracy of OPT-6.7B, LLaMA-7B, and LLaMA2-7B [PITH_FULL_IMAGE:figures/full_fig_p004_7.png]
Figure 8
Figure 8. Figure 8: Comparison of (a) the current computation scheme on GPU, (b) and that enhanced with dedicated FP-INT processing unit, (c) FIGNA scheme, and (d) [PITH_FULL_IMAGE:figures/full_fig_p005_8.png]
Figure 9
Figure 9. Figure 9: Search process of the proposed adaptive precision combination search [PITH_FULL_IMAGE:figures/full_fig_p006_9.png]
Figure 11
Figure 11. Figure 11: The architecture of Anda-enhanced bit-serial processing unit, which [PITH_FULL_IMAGE:figures/full_fig_p007_11.png]
Figure 12
Figure 12. Figure 12: The architecture of the on-the-fly bit plane compressor and the man [PITH_FULL_IMAGE:figures/full_fig_p008_12.png]
Figure 14
Figure 14. Figure 14: Identified best precision combinations of various LLMs on different [PITH_FULL_IMAGE:figures/full_fig_p010_14.png]
Figure 16
Figure 16. Figure 16: Speedup, area efficiency, and energy efficiency comparison across accelerators on WikiText2. All data are aligned to the GPU-like FP-FP baseline. pared to the corresponding FIGNA variants, Anda achieves 1.48× and 1.25× higher acceleration, benefiting from efficient ut…
Figure 17
Figure 17. Figure 17: Energy breakdown of Anda in contrast with the baseline accelerators. Energy consumption during the LLaMA-13B inference is evaluated. TABLE III AREA AND POWER CHARACTERISTICS OF ANDA Component Setup Area [mm2 ] Power [mW] MXU 16×16 APUs 0.41 (18.89%) 54.34 (66.94%) BPC…
Figure 18
Figure 18. Figure 18: Speedup and energy efficiency improvement of Anda over FP-FP [PITH_FULL_IMAGE:figures/full_fig_p012_18.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 66 canonical work pages

  1. [1]

    Resq: Residual quantization for video perception,

    D. Abati, H. Ben Yahia, M. Nagel, and A. Habibian, “Resq: Residual quantization for video perception,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023, pp. 17 119– 17 129

  2. [2]

    Bit-pragmatic deep neural network computing,

    J. Albericio, A. Delm ´as, P. Judd, S. Sharify, G. O’Leary, R. Genov, and A. Moshovos, “Bit-pragmatic deep neural network computing,” in Proceedings of the 50th annual IEEE/ACM international symposium on microarchitecture (MICRO), 2017, pp. 382–394

  3. [3]

    Explaining neural scaling laws,

    Y . Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma, “Explaining neural scaling laws,” Proceedings of the National Academy of Sciences (PNAS), vol. 121, no. 27, p. e2311878121, 2024

  4. [4]

    Longbench: A bilingual, multitask benchmark for long context understanding,

    Y . Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, Y . Dong, J. Tang, and J. Li, “Longbench: A bilingual, multitask benchmark for long context understanding,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL) , 2024, pp. 3119–3137

  5. [5]

    Demystifying chatgpt: An in-depth survey of openai’s robust large language models,

    P. Bhattacharya, V . K. Prasad, A. Verma, D. Gupta, A. Sapsomboon, W. Viriyasitavat, and G. Dhiman, “Demystifying chatgpt: An in-depth survey of openai’s robust large language models,” Archives of Compu- tational Methods in Engineering , pp. 1–44, 2024

  6. [6]

    Genus synthesis solution,

    Candence, “Genus synthesis solution,” https://www.cadence.com/en US/home/tools/digital-design-and-signoff/synthesis/genus-synthesis- solution.html, 2024, online; accessed 2024-07-16

  7. [7]

    General purpose deep learning accelerator based on bit interleaving,

    L. Chang, H. Lu, C. Li, X. Zhao, Z. Hu, J. Zhou, and X. Li, “General purpose deep learning accelerator based on bit interleaving,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD), 2023

  8. [8]

    Quip: 2-bit quanti- zation of large language models with guarantees,

    J. Chee, Y . Cai, V . Kuleshov, and C. M. De Sa, “Quip: 2-bit quanti- zation of large language models with guarantees,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 36, 2024

Show all 89 references
  1. [9]

    Efficientqat: Efficient quantization-aware training for large language models,

    M. Chen, W. Shao, P. Xu, J. Wang, P. Gao, K. Zhang, Y . Qiao, and P. Luo, “Efficientqat: Efficient quantization-aware training for large language models,” arXiv preprint arXiv:2407.11062 , 2024

  2. [10]

    Nacl: A general and effective kv cache eviction framework for llm at inference time,

    Y . Chen, G. Wang, J. Shang, S. Cui, Z. Zhang, T. Liu, S. Wang, Y . Sun, D. Yu, and H. Wu, “Nacl: A general and effective kv cache eviction framework for llm at inference time,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume ...

  3. [11]

    Palm: Scaling language modeling with pathways,

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y . Tay, N. Shazeer, V . Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. ...

  4. [12]

    Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,

    S. Dai, R. Venkatesan, M. Ren, B. Zimmer, W. Dally, and B. Khailany, “Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference,” Proceedings of Machine Learning and Sys- tems (MLSys), vol. 3, pp. 873–884, 2021

  5. [13]

    Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point,

    B. Darvish Rouhani, D. Lo, R. Zhao, M. Liu, J. Fowers, K. Ovtcharov, A. Vinogradsky, S. Massengill, L. Yang, R. Bittner, A. Forin, H. Zhu, T. Na, P. Patel, S. Che, L. Chand Koppaka, X. Song, S. Som, K. Das, S. T, S. Reinhardt, S. Lanka, E. Chung, and D. Burger, “Pushing the li...

  6. [14]

    With shared microexponents, a little shifting goes a long way,

    B. Darvish Rouhani, R. Zhao, V . Elango, R. Shafipour, M. Hall, M. Mesmakhosroshahi, A. More, L. Melnick, M. Golub, G. Varatkar, L. Shao, G. Kolhe, D. Melts, J. Klar, R. L’Heureux, M. Perry, D. Burger, E. Chung, Z. S. Deng, S. Naghshineh, J. Park, and M. Naumov, “With shared m...

  7. [15]

    A timing-driven approach to synthesize fast barrel shifters,

    S. Das and S. P. Khatri, “A timing-driven approach to synthesize fast barrel shifters,” IEEE Transactions on Circuits and Systems II: Express Briefs (TCAS-II), vol. 55, no. 1, pp. 31–35, 2008

  8. [16]

    Llm.int8(): 8- bit matrix multiplication for transformers at scale,

    T. Dettmers, M. Lewis, Y . Belkada, and L. Zettlemoyer, “Llm.int8(): 8- bit matrix multiplication for transformers at scale,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 35, pp. 30 318–30 332, 2022

  9. [17]

    The case for 4-bit precision: k- bit inference scaling laws,

    T. Dettmers and L. Zettlemoyer, “The case for 4-bit precision: k- bit inference scaling laws,” in International Conference on Machine Learning (ICML). PMLR, 2023, pp. 7750–7774

  10. [18]

    Hawq: Hessian aware quantization of neural networks with mixed-precision,

    Z. Dong, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Hawq: Hessian aware quantization of neural networks with mixed-precision,” in Proceedings of the IEEE/CVF international conference on computer vision (ICCV), 2019, pp. 293–302

  11. [19]

    Training dnns with hybrid block floating point,

    M. Drumond, T. Lin, M. Jaggi, and B. Falsafi, “Training dnns with hybrid block floating point,” Advances in Neural Information Processing Systems (NeurIPS), vol. 31, 2018

  12. [20]

    Skvq: Sliding-window key and value cache quantization for large language models,

    H. Duanmu, Z. Yuan, X. Li, J. Duan, X. Zhang, and D. Lin, “Skvq: Sliding-window key and value cache quantization for large language models,” in First Conference on Language Modeling (COLM) , 2024

  13. [21]

    Extreme compression of large language models via additive quantization,

    V . Egiazarian, A. Panferov, D. Kuznedelev, E. Frantar, A. Babenko, and D. Alistarh, “Extreme compression of large language models via additive quantization,” inInternational Conference on Machine Learning (ICML). PMLR, 2024

  14. [22]

    Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,

    H. Fan, H.-C. Ng, S. Liu, Z. Que, X. Niu, and W. Luk, “Reconfig- urable acceleration of 3d-cnns for human action recognition with block floating-point representation,” in 28th International Conference on Field Programmable Logic and Applications (FPL) . IEEE, 2018, pp. 287– 2877

  15. [23]

    Static block floating-point quantization for convolutional neural networks on fpga,

    H. Fan, G. Wang, M. Ferianc, X. Niu, and W. Luk, “Static block floating-point quantization for convolutional neural networks on fpga,” in International Conference on Field-Programmable Technology (ICFPT) . IEEE, 2019, pp. 28–35

  16. [24]

    Optq: Accurate quantization for generative pre-trained transformers,

    E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “Optq: Accurate quantization for generative pre-trained transformers,” in The Eleventh International Conference on Learning Representations (ICLR) , 2023

  17. [25]

    Llmc: Benchmarking large language model quantization with a versatile compression toolkit,

    R. Gong, Y . Yong, S. Gu, Y . Huang, Y . Zhang, X. Liu, and D. Tao, “Llmc: Benchmarking large language model quantization with a versatile compression toolkit,” arXiv preprint arXiv:2405.06001 , 2024

  18. [26]

    Boost: block minifloat-based on-device cnn training accelerator with transfer learning,

    C. Guo, B. Lou, X. Liu, D. Boland, P. H. Leong, and C. Zhuo, “Boost: block minifloat-based on-device cnn training accelerator with transfer learning,” in IEEE/ACM International Conference on Computer Aided Design (ICCAD). IEEE, 2023, pp. 1–9

  19. [27]

    Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,

    C. Guo, J. Tang, W. Hu, J. Leng, C. Zhang, F. Yang, Y . Liu, M. Guo, and Y . Zhu, “Olive: Accelerating large language models via hardware- friendly outlier-victim pair quantization,” in Proceedings of the 50th Annual International Symposium on Computer Architecture (ISCA) , 20...

  20. [28]

    Ese: Efficient speech recognition engine with sparse lstm on fpga,

    S. Han, J. Kang, H. Mao, Y . Hu, X. Li, Y . Li, D. Xie, H. Luo, S. Yao, Y . Wang, H. Yang, and W. B. J. Dally, “Ese: Efficient speech recognition engine with sparse lstm on fpga,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (F...

  21. [29]

    Kvquant: Towards 10 million context length llm inference with kv cache quantization,

    C. Hooper, S. Kim, H. Mohammadzadeh, M. W. Mahoney, Y . S. Shao, K. Keutzer, and A. Gholami, “Kvquant: Towards 10 million context length llm inference with kv cache quantization,” arXiv preprint arXiv:2401.18079, 2024

  22. [30]

    A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,

    L. Huang, C. Fang, Q. Li, J. Lin, and Z. Wang, “A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,” in 29th Asia and South Pacific Design Automation Conference (ASP-DAC). IEEE, 2024, pp. 927–932

  23. [31]

    Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,

    Q. Huang, P.-A. Tsai, J. S. Emer, and A. Parashar, “Mind the gap: Attainable data movement and operational intensity bounds for tensor algorithms,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA) , 2024

  24. [32]

    Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,

    J. Jang, Y . Kim, J. Lee, and J.-J. Kim, “Figna: Integer unit-based accel- erator design for fp-int gemm preserving numerical accuracy,” in IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2024, pp. 760–773

  25. [33]

    Perplexity—a measure of the difficulty of speech recognition tasks,

    F. Jelinek, R. L. Mercer, L. R. Bahl, and J. K. Baker, “Perplexity—a measure of the difficulty of speech recognition tasks,” The Journal of the Acoustical Society of America (JASA) , vol. 62, no. S1, pp. S63–S63, 1977

  26. [34]

    Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,

    Y . Jeon, C. Lee, E. Cho, and Y . Ro, “Mr. biq: Post-training non- uniform quantization based on minimizing the reconstruction error,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 12 329–12 338

  27. [35]

    Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,

    Y . Jeon, B. Park, S. J. Kwon, B. Kim, J. Yun, and D. Lee, “Biqgemm: matrix multiplication with lookup table for binary-coding-based quan- tized dnns,” in International Conference for High Performance Comput- ing, Networking, Storage and Analysis (SC) . IEEE, 2020, pp. 1–14

  28. [36]

    Ten lessons from three generations shaped google’s tpuv4i: Industrial product,

    N. P. Jouppi, D. Hyun Yoon, M. Ashcraft, M. Gottscho, T. B. Jablin, G. Kurian, J. Laudon, S. Li, P. Ma, X. Ma, T. Norrie, N. Patil, S. Prasad, C. Young, Z. Zhou, and D. Patterson, “Ten lessons from three generations shaped google’s tpuv4i: Industrial product,” in ACM/IEEE 48th...

  29. [37]

    Stripes: Bit-serial deep neural network computing,

    P. Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, “Stripes: Bit-serial deep neural network computing,” in 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2016, pp. 1–12

  30. [38]

    A survey of gpt-3 family large language models including chatgpt and gpt-4,

    K. S. Kalyan, “A survey of gpt-3 family large language models including chatgpt and gpt-4,” Natural Language Processing Journal , p. 100048, 2023

  31. [39]

    A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,

    B. Keller, R. Venkatesan, S. Dai, S. G. Tell, B. Zimmer, C. Sakr, W. J. Dally, C. T. Gray, and B. Khailany, “A 95.6-tops/w deep learning inference accelerator with per-vector scaled 4-bit quantization in 5 nm,” IEEE Journal of Solid-State Circuits (JSSC) , vol. 58, no. 4, pp. ...

  32. [40]

    Compressed context mem- ory for online language model interaction,

    J.-H. Kim, J. Yeom, S. Yun, and H. O. Song, “Compressed context mem- ory for online language model interaction,” in The Twelfth International Conference on Learning Representations (ICLR) , 2024

  33. [41]

    Dacapo: Accelerating continuous learning in autonomous systems for video analytics,

    Y . Kim, C. Oh, J. Hwang, W. Kim, S. Oh, Y . Lee, H. Sharma, A. Yaz- danbakhsh, and J. Park, “Dacapo: Accelerating continuous learning in autonomous systems for video analytics,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA) , 2024

  34. [42]

    Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,

    Y . Kim, J. Jang, J. Lee, J. Park, J. Kim, B. Kim, B. park, S. J. Kwon, D. Lee, and J.-J. Kim, “Winning both the accuracy of floating point activation and the simplicity of integer arithmetic,” in The Eleventh International Conference on Learning Representations (ICLR) , 2023

  35. [43]

    One-shot model for mixed-precision quantization,

    I. Koryakovskiy, A. Yakovleva, V . Buchnev, T. Isaev, and G. Odinokikh, “One-shot model for mixed-precision quantization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7939–7949

  36. [44]

    Flexpoint: An adaptive numerical format for efficient training of deep neural networks,

    U. K ¨oster, T. J. Webb, X. Wang, M. Nassar, A. K. Bansal, W. H. Constable, O. H. Elibol, S. Gray, S. Hall, L. Hornof, A. Khosrowshahi, C. Kloss, R. J. Pai, and N. Rao, “Flexpoint: An adaptive numerical format for efficient training of deep neural networks,” Advances in Neural...

  37. [45]

    Tender: Accelerating large language models via tensor decomposition and runtime requantization,

    J. Lee, W. Lee, and J. Sim, “Tender: Accelerating large language models via tensor decomposition and runtime requantization,” in Proceedings of the 51st Annual International Symposium on Computer Architecture (ISCA), 2024

  38. [46]

    Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,

    A. Li, H. Mo, W. Zhu, Q. Li, S. Yin, S. Wei, and L. Liu, “Bitcluster: Fine-grained weight quantization for load-balanced bit-serial neural network accelerators,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) , vol. 41, no. 11, pp. 4747– 4...

  39. [47]

    Norm tweaking: High-performance low-bit quantization of large language models,

    L. Li, Q. Li, B. Zhang, and X. Chu, “Norm tweaking: High-performance low-bit quantization of large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), vol. 38, no. 17, 2024, pp. 18 536–18 544

  40. [48]

    Geo: Generation and execution optimized stochastic computing accelerator for neural networks,

    T. Li, W. Romaszkan, S. Pamarti, and P. Gupta, “Geo: Generation and execution optimized stochastic computing accelerator for neural networks,” in 2021 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 2021, pp. 689–694

  41. [49]

    Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,

    Z. Li, A. Lu, Y . Xie, Z. Kong, M. Sun, H. Tang, Z. J. Xue, P. Dong, C. Ding, Y . Wang, X. Lin, and Z. Fang, “Quasar-vit: Hardware-oriented quantization-aware architecture search for vision transformers,” in Pro- ceedings of the 38th ACM International Conference on Supercomput...

  42. [50]

    High-performance fpga-based cnn accelerator with block-floating-point arithmetic,

    X. Lian, Z. Liu, Z. Song, J. Dai, W. Zhou, and X. Ji, “High-performance fpga-based cnn accelerator with block-floating-point arithmetic,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI) , vol. 27, no. 8, pp. 1874–1885, 2019

  43. [51]

    Awq: Activation-aware weight quan- tization for llm compression and acceleration,

    J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han, “Awq: Activation-aware weight quan- tization for llm compression and acceleration,” in The Seventh Annual Conference on Machine Learning and Systems (MLSys) , 2024

  44. [52]

    Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,

    Y . Lin, H. Tang, S. Yang, Z. Zhang, G. Xiao, C. Gan, and S. Han, “Qserve: W4a8kv4 quantization and system co-design for efficient llm serving,” arXiv preprint arXiv:2405.04532 , 2024

  45. [53]

    Llm-qat: Data-free quantization aware training for large language models,

    Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y . Mehdad, Y . Shi, R. Kr- ishnamoorthi, and V . Chandra, “Llm-qat: Data-free quantization aware training for large language models,” arXiv preprint arXiv:2305.17888 , 2023

  46. [54]

    Kivi: A tuning-free asymmetric 2bit quantization for kv cache,

    Z. Liu, J. Yuan, H. Jin, S. Zhong, Z. Xu, V . Braverman, B. Chen, and X. Hu, “Kivi: A tuning-free asymmetric 2bit quantization for kv cache,” in Forty-first International Conference on Machine Learning (ICML) , 2024

  47. [55]

    Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,

    H. Lu, L. Chang, C. Li, Z. Zhu, S. Lu, Y . Liu, and M. Zhang, “Dis- tilling bit-level sparsity parallelism for general purpose deep learning acceleration,” in 54th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2021, pp. 963–976

  48. [56]

    Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,

    S. Luohe, H. Zhang, Y . Yao, Z. Li et al. , “Keep the cost down: A review on methods to optimize llm’s kv-cache consumption,” in First Conference on Language Modeling (COLM) , 2024

  49. [57]

    Efficient arbitrary precision acceleration for large language models on gpu tensor cores,

    S. Ma, C. Fang, H. Shao, and Z. Wang, “Efficient arbitrary precision acceleration for large language models on gpu tensor cores,” arXiv preprint arXiv:2409.17870, 2024

  50. [58]

    Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,

    S. Mach, F. Schuiki, F. Zaruba, and L. Benini, “Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing,” IEEE Transactions on Very Large Scale Integration (VLSI) Systems (TVLSI) , vol. 29, no. 4, pp. 774–787, 2020

  51. [59]

    The penn treebank: Anno- tating predicate argument structure,

    M. Marcus, G. Kim, M. A. Marcinkiewicz, R. MacIntyre, A. Bies, M. Ferguson, K. Katz, and B. Schasberger, “The penn treebank: Anno- tating predicate argument structure,” in Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994, 1994

  52. [60]

    Pointer sentinel mix- ture models,

    S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mix- ture models,” in International Conference on Learning Representations (ICLR), 2017

  53. [61]

    Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,

    S.-H. Noh, J. Koo, S. Lee, J. Park, and J. Kung, “Flexblock: A flexible dnn training accelerator with multi-mode block floating point support,” IEEE Transactions on Computers (TC) , vol. 72, no. 9, pp. 2522–2535, 2023

  54. [62]

    Cutlass,

    NVIDIA, “Cutlass,” https://github.com/NVIDIA/cutlass, 2024, online; accessed 2024-07-03

  55. [63]

    Gpt- 4 technical report,

    OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V . Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L...

  56. [64]

    LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,

    G. Park, B. park, M. Kim, S. Lee, J. Kim, B. Kwon, S. J. Kwon, B. Kim, Y . Lee, and D. Lee, “LUT-GEMM: Quantized matrix multipli- cation based on LUTs for efficient inference in large-scale generative language models,” in The Twelfth International Conference on Learning Repres...

  57. [65]

    Exploring the limits of transfer learning with a unified text-to-text transformer,

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research (JMLR), vol. 21, no. 140, pp. 1–67, 2020

  58. [66]

    Omniquant: Omnidirectionally calibrated quantization for large language models,

    W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, K. Zhang, P. Gao, Y . Qiao, and P. Luo, “Omniquant: Omnidirectionally calibrated quantization for large language models,” in The Twelfth International Conference on Learning Representations (ICLR) , 2024

  59. [67]

    Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,

    H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, J. K. Kim, V . Chandra, and H. Esmaeilzadeh, “Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,” in Proceedings of the ACM/IEEE 45th Annual International Symposium on Computer Architect...

  60. [68]

    Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,

    M. Shi, V . Jain, A. Joseph, M. Meijer, and M. Verhelst, “Bitwave: Exploiting column-based bit-level sparsity for deep learning accelera- tion,” in IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2024, pp. 732–746

  61. [69]

    Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,

    W. Sun, A. Li, T. Geng, S. Stuijk, and H. Corporaal, “Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors,” IEEE Transactions on Parallel and Distributed Systems (TPDS), vol. 34, no. 1, pp. 246–261, 2022

  62. [70]

    Gemma: Open models based on gemini research and technology,

    G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Riviere, M. Kale, J. C. Love, P. D. Tafti, L. Hussenot, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. H’eliou, A. Tacchetti, A. Bulanova, A. Paterson, B. Tsai, B. Sh...

  63. [71]

    Bebert: Efficient and robust binary ensemble bert,

    J. Tian, C. Fang, H. Wang, and Z. Wang, “Bebert: Efficient and robust binary ensemble bert,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5

  64. [72]

    Llama: Open and efficient foundation language models,

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971 , 2023

  65. [73]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V . Goswami, N. Goyal, A. Hartshorn, S. Ho...

  66. [74]

    Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,

    A. Tseng, J. Chee, Q. Sun, V . Kuleshov, and C. De Sa, “Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks,” arXiv preprint arXiv:2402.04396 , 2024

  67. [75]

    Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,

    G. Wang, S. Cai, W. Li, D. Lyu, and G. He, “Bsvit: A bit-serial vision transformer accelerator exploiting dynamic patch and weight bit-group quantization,” IEEE Transactions on Circuits and Systems I: Regular Papers (TCAS-I), 2024

  68. [76]

    Haq: Hardware-aware automated quantization with mixed precision,

    K. Wang, Z. Liu, Y . Lin, J. Lin, and S. Han, “Haq: Hardware-aware automated quantization with mixed precision,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), 2019, pp. 8612–8620

  69. [77]

    Outlier suppression: Pushing the limit of low-bit transformer language models,

    X. Wei, Y . Zhang, X. Zhang, R. Gong, S. Zhang, Q. Zhang, F. Yu, and X. Liu, “Outlier suppression: Pushing the limit of low-bit transformer language models,” Advances in Neural Information Processing Systems (NeurIPS), vol. 35, pp. 17 402–17 414, 2022

  70. [78]

    Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,

    H. Xia, Z. Zheng, X. Wu, S. Chen, Z. Yao, S. Youn, A. Bakhtiari, M. Wyatt, D. Zhuang, Z. Zhou, O. Ruwase, Y . He, and S. L. Song, “Quant-llm: Accelerating the serving of large language models via fp6- centric algorithm-system co-design on modern gpus,” in 2024 USENIX Annual Te...

  71. [79]

    Smoothquant: Accurate and efficient post-training quantization for large language models,

    G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han, “Smoothquant: Accurate and efficient post-training quantization for large language models,” in International Conference on Machine Learning (ICML). PMLR, 2023, pp. 38 087–38 099

  72. [80]

    Efficient streaming language models with attention sinks,

    G. Xiao, Y . Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in The Twelfth International Conference on Learning Representations (ICLR) , 2024

  73. [81]

    Onebit: Towards extremely low-bit large language models,

    Y . Xu, X. Han, Z. Yang, S. Wang, Q. Zhu, Z. Liu, W. Liu, and W. Che, “Onebit: Towards extremely low-bit large language models,” arXiv preprint arXiv:2402.11295 , 2024

  74. [82]

    Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,

    J. Yuan, H. Liu, S. Zhong, Y .-N. Chuang, S. Li, G. Wang, D. Le, H. Jin, V . Chaudhary, Z. Xu, Z. Liu, and X. Hu, “Kv cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches,” in The 2024 Conference on Empirical Methods ...

  75. [83]

    Llm inference unveiled: Survey and roofline model insights,

    Z. Yuan, Y . Shang, Y . Zhou, Z. Dong, Z. Zhou, C. Xue, B. Wu, Z. Li, Q. Gu, Y . J. Lee, Y . Yan, B. Chen, G. Sun, and K. Keutzer, “Llm inference unveiled: Survey and roofline model insights,” arXiv preprint arXiv:2402.16363, 2024

  76. [84]

    Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,

    A. H. Zadeh, M. Mahmoud, A. Abdelhadi, and A. Moshovos, “Mokey: Enabling narrow fixed-point inference for out-of-the-box floating-point transformer models,” in Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA) , 2022, pp. 888–901

  77. [85]

    Fast: Dnn training under variable precision block floating point with stochastic rounding,

    S. Q. Zhang, B. McDanel, and H. Kung, “Fast: Dnn training under variable precision block floating point with stochastic rounding,” inIEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2022, pp. 846–860

  78. [86]

    Opt: Open pre-trained transformer language models,

    S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V . Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer, “Opt: Open pre-trained transformer language models,” arXiv preprint ...

  79. [87]

    Cam: Cache merging for memory-efficient llms inference,

    Y . Zhang, Y . Du, G. Luo, Y . Zhong, Z. Zhang, S. Liu, and R. Ji, “Cam: Cache merging for memory-efficient llms inference,” in Forty- first International Conference on Machine Learning (ICML) , 2024

  80. [88]

    H2o: Heavy-hitter oracle for efficient generative inference of large language models,

    Z. Zhang, Y . Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y . Tian, C. R´e, C. Barrett et al., “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 36, 2024

  81. [89]

    Atom: Low-bit quantization for efficient and accurate llm serving,

    Y . Zhao, C.-Y . Lin, K. Zhu, Z. Ye, L. Chen, S. Zheng, L. Ceze, A. Krishnamurthy, T. Chen, and B. Kasikci, “Atom: Low-bit quantization for efficient and accurate llm serving,” Proceedings of Machine Learning and Systems (MLSys) , vol. 6, pp. 196–209, 2024

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.