REVIEW 3 major objections 5 minor 76 references
Moving weight dequantization into the HBM base die removes the CUDA-core bottleneck for large-batch LLM inference.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 01:09 UTC pith:NFAY3UHD
load-bearing objection Solid, low-intrusion HBM-base-die dequant design for large-batch weight-only LLM inference; numbers are simulator-bound but the architecture and engineering case are real. the 3 major comments →
StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Integrating compact DeQuantization Blocks into the HBM base die and performing inline dequantization on standard, sideband-tagged memory loads eliminates GPU-side CUDA-core dequantization and the extra HBM write-back/reload of dequantized weights that appear at large batch sizes, delivering up to 7.08× mixed-precision GEMM speedup, 90 percent lower energy, and up to 2.2× higher end-to-end LLM decode throughput at negligible area and power cost.
What carries the argument
The DeQuantization Block (DQB): a per-pseudo-channel unit on the HBM base-die read path that, guided by a few-bit sideband tag, converts quantized weights (plus co-located scale/zero-point metadata) into the compute format and returns them through the ordinary load-response path while preserving conventional load semantics.
Load-bearing premise
The simulator that projects these gains from full-precision GPU traces and pre-silicon RTL/thermal models must be accurate enough to represent a real custom-HBM system with sideband tags and pseudo-channel-aware layouts.
What would settle it
Build or cycle-accurately simulate a GPU-plus-custom-HBM stack that implements the sideband tags and DQBs, run the same W4A16 and W8A16 LLaMA/Qwen/Mistral large-batch decode workloads, and check whether the measured latency and energy reductions match the claimed 7× / 90 percent / 2.2× figures.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes StreamDQ, a near-memory architecture that places compact DeQuantization Blocks (DQBs) on the HBM base-die read path of each pseudo-channel. A few-bit sideband tag on ordinary GPU memory loads selects conversion mode (INT4/INT8/FP8 → FP16/BF16) or bypass while preserving address and load semantics. Combined with a one-time pseudo-channel-aware G/S/Z layout, this eliminates CUDA-core dequantization, on-chip traffic, and the HBM write-back/reload of dequantized weights that appears in large-batch split kernels. RTL synthesis (12 nm, pre-CTS) reports 0.127 mm² / 0.355 W per DQB; FloTHERM thermal maps on an HBM3-proxy stack stay inside the base-die budget. Using an in-house StreamDQ-Sim (Accel-Sim + AccelWattch calibrated to A100 Nsight/NVML), the authors claim up to 7.08× mpGEMM speedup and 90.23 % energy reduction versus fused GPTQ/AWQ/TorchAO kernels, and up to 54.68 % lower end-to-end latency / 2.20× decode throughput on three 7–8 B LLMs.
Significance. Weight-only quantization is already the dominant practical path for cloud LLM serving; the paper correctly identifies CUDA-core dequantization and the fused-versus-split kernel trade-off as first-order bottlenecks under large-batch, compute-bound regimes. Relocating dequantization into the HBM base die with only a sideband tag and PC-local metadata is a clean, low-intrusion NMP design that sidesteps the classic VA-to-PA and cross-channel placement problems. The multi-format datapath (shared FP32 ALUs, wire-mapped FP conversion, shared INT LUT with zero-padding) and the explicit area/power/thermal feasibility study are concrete engineering contributions. If the quantitative gains hold under real custom-HBM silicon and modest GPU tag support, StreamDQ would be a high-impact, deployable enhancement for next-generation AI memory stacks. Credit is due for the careful fused/split kernel selection, the public-baseline comparisons, and the RTL + FloTHERM evidence rather than pure analytical claims.
major comments (3)
- [§5.2 Simulation Methodology; §6.3–6.4] §5.2 and §6.3–6.4: All headline speedups (7.08×), energy reductions (90.23 %), and end-to-end numbers (54.68 % latency, 2.20× throughput) are produced by StreamDQ-Sim. Traces are taken from fpGEMM, rewritten as mpGEMM, run through a modified Accel-Sim, then scaled by the simulated-to-measured A100 fpGEMM ratio. Reported MAPE remains 7–28 % even after A100-specific AccelWattch extensions; the paper never demonstrates that the same scale factor remains valid once sideband-tag parsing, S/Z request generation, PC-aware layout, and DQB pipeline latency are present. End-to-end results further multiply the simulated kernel speedup by the Nsight-Systems mpGEMM fraction, so any systematic bias is amplified. A sensitivity study (or cycle-accurate DQB model validated against a micro-benchmark) is required before the quantitative claims can be treated as reliable.
- [§6.1 Area and Power Overhead] §6.1 and Table 5: Area (0.127 mm²) and power (0.355 W) are pre-CTS 12 nm RTL numbers; the stack-level overhead (3.36 % area, 11.36 W) is extrapolated from public HBM3 parameters used as a proxy for custom HBM4-class base dies. No post-CTS, place-and-route, or silicon correlation is provided, nor is the interaction with the real D2D PHY / MC / NoC floorplan quantified. Because the central feasibility argument rests on these numbers remaining “modest,” the manuscript should either supply tighter physical-design results or clearly qualify the claims as pre-silicon estimates.
- [§3.2 Sideband Tagging; §3.3 Pseudo-Channel-Aware Layout] §3.2–3.3 and §4: The design assumes that a few-bit sideband tag can be carried on every weight load (spare metadata bits or a minimal request-path extension) and that a privileged runtime can install a region-lookup table and perform the offline PC-aware G/S/Z reorganization. While the paper correctly notes that these changes are modest compared with full NMP, they are still non-zero GPU-side and software-stack modifications. The evaluation never measures the tag-generation latency, table-miss/reprogramming cost, or any extra memory-controller arbitration introduced by the S/Z request generator. Without that overhead characterization, the claim of “preserving conventional load semantics with minimal GPU-side changes” remains incompletely substantiated.
minor comments (5)
- [Fig. 1] Fig. 1 caption and body: “len = x” is used for both input and output sequence lengths; a short clarification that prefill and decode are both set to the same length would avoid ambiguity.
- [Tables 2–3] Table 2 and Table 3: The S/Z replication and metadata-request fractions are useful, but the assumed Z-bit-width matching the weight precision and S always 16-bit should be stated once in the table captions for self-containment.
- [§6.4] §6.4 Kernel Selection: The heuristic thresholds (GPTQ switch at batch 64, AWQ remaining fused up to 128) are reasonable, yet a one-sentence note on how sensitive the end-to-end ranking is to a ±1 bin shift would strengthen the fairness claim.
- [Front matter / References] References and ACM template: Several arXiv preprints and GitHub links appear; ensure the final camera-ready version follows the venue’s citation style and that the placeholder “Conference acronym ’XX” / 2018 dates are updated.
- [§3.6] Fig. 10 and Fig. 11: The wire-mapping diagrams are dense; a short textual walk-through of one concrete bit pattern (e.g., FP8 E4M3 → FP16) would improve readability for non-specialists.
Circularity Check
No circularity: architecture proposal with external-baseline simulation; claims are measured/simulated outcomes, not self-derived identities.
full rationale
StreamDQ is a systems/architecture paper whose central claims (mpGEMM speedups up to 7.08×, energy reductions up to 90.23 %, E2E latency/throughput gains) are obtained by comparing a proposed hardware design against independent software baselines (GPTQ, AWQ-v1/v2, TorchAO, PyTorch fpGEMM) on real A100 measurements plus a calibrated Accel-Sim model. Design parameters (3-bit tags, 328 MHz, buffer sizes, PC-aware layout) are engineering choices, not quantities fitted to the target metrics and then re-presented as predictions. The StreamDQ-Sim calibration (fpGEMM traces scaled by measured-to-simulated ratios, MAPE reported) is ordinary simulator validation against external silicon data; it does not make the reported speedups tautological by construction. There are no self-definitional equations, no uniqueness theorems imported from the authors’ prior work, no ansatz smuggled via self-citation, and no renaming of a known empirical pattern. The paper is therefore free of the circularity patterns enumerated in the analyzer; any residual modeling uncertainty belongs to correctness risk, not circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- DQB operating clock frequency =
328 MHz
- S/Z request table and buffer sizes =
8 KB table, 4 KB buffer
- Accel-Sim/AccelWattch to A100 calibration scale =
PCC ~0.83–0.87; MAPE ~7–29% by regime
- Input toggle rate for DQB power =
20%
axioms (5)
- domain assumption Weight-only per-group quantization with co-located S/Z metadata is the target deployment model for large-batch LLM inference.
- domain assumption HBM base dies in custom/HBM4-class stacks can host modest logic under practical area, power, and thermal budgets without breaking DRAM retention margins.
- ad hoc to paper A few-bit sideband tag can be carried on GPU memory read requests (spare metadata or minimal path extension) without changing effective addresses or load semantics.
- domain assumption Pseudo-channel interleaving maps can be made weight-group and S/Z co-local via offline layout transformation with only small S/Z replication overhead.
- standard math Standard dequantization arithmetic Dequant(x)=(x−z)·s and the supported format conversions preserve the numerical behavior expected by existing tensor-core GEMM paths.
invented entities (3)
-
DeQuantization Block (DQB)
no independent evidence
-
StreamDQ sideband tag + region-lookup control path
no independent evidence
-
Pseudo-channel-aware G/S/Z memory layout
no independent evidence
read the original abstract
As large language models (LLMs) scale, their memory and computation demands have grown substantially, making weight-only quantization a widely adopted technique for reducing model size with minimal accuracy loss. However, on current GPUs, CUDA-core-based dequantization introduces substantial instruction overhead, on-chip traffic, and pipeline stalls, making it a major bottleneck for high-throughput, cloud-scale LLM serving. To address these limitations, we propose StreamDQ, a lightweight architectural enhancement that enables on-the-fly dequantization in the memory subsystem for high-throughput, large-batch LLM inference. StreamDQ integrates compact DeQuantization Blocks (DQBs) into the base die of high-bandwidth memory (HBM) and performs inline dequantization on standard memory loads. A lightweight sideband tag on each memory read request selects the dequantization mode while preserving conventional load semantics. By relocating dequantization to the memory side, StreamDQ eliminates GPU-side CUDA-core-based dequantization, thereby reducing on-chip traffic on the GPU and avoiding extra HBM write-back and reload of dequantized weights at large batch sizes. Our evaluation shows that StreamDQ achieves up to 7.08$\times$ speedup and 90.23\% lower energy for mixed-precision GEMM, with only 0.127\,mm$^2$ area and 0.355\,W power overhead per DQB in a 12\,nm CMOS process. For end-to-end LLM inference, StreamDQ reduces latency by up to 54.68\% and improves decode throughput by up to 2.20$\times$.
Figures
Reference graph
Works this paper leans on
-
[1]
AWQ GitHub repository.https://github.com/mit-han-lab/llm- awq/
2024. AWQ GitHub repository.https://github.com/mit-han-lab/llm- awq/
2024
-
[2]
NVIDIA H100 NVL GPU.https://www.nvidia.com/content/dam/ en-zz/Solutions/Data-Center/h100/PB-11773-001_v01.pdf
2024. NVIDIA H100 NVL GPU.https://www.nvidia.com/content/dam/ en-zz/Solutions/Data-Center/h100/PB-11773-001_v01.pdf
2024
-
[3]
vLLM GitHub repository.https://github.com/vllm-project/vllm/
2024. vLLM GitHub repository.https://github.com/vllm-project/vllm/
2024
-
[4]
FloTHERM.https://plm.sw.siemens.com/
2025. FloTHERM.https://plm.sw.siemens.com/
2025
-
[5]
NVIDIA Management Library (NVML).https://developer.nvidia
2025. NVIDIA Management Library (NVML).https://developer.nvidia. com/management-library-nvml/
2025
-
[6]
NVIDIA NSight Compute.https://developer.nvidia.com/nsight- compute/
2025. NVIDIA NSight Compute.https://developer.nvidia.com/nsight- compute/
2025
-
[7]
NVIDIA NSight Systems.https://developer.nvidia.com/nsight- systems/
2025. NVIDIA NSight Systems.https://developer.nvidia.com/nsight- systems/
2025
-
[8]
Synopsys Design Compiler.https://www.synopsys.com/
2025. Synopsys Design Compiler.https://www.synopsys.com/
2025
-
[9]
TensorRT-Weight-only-quantization.https://developer.nvidia
2025. TensorRT-Weight-only-quantization.https://developer.nvidia. com/blog/nvidia-tensorrt-10-0-upgrades-usability-performance- and-ai-model-support/
2025
-
[10]
Mohammad Alian, Seung Won Min, Hadi Asgharimoghaddam, Ashutosh Dhar, Dong Kai Wang, Thomas Roewer, Adam McPadden, Oliver O’Halloran, Deming Chen, Jinjun Xiong, et al. 2018. Application- transparent near-memory processing architecture with memory chan- nel network. In2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 802–814
2018
-
[11]
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma. 2024. Explaining neural scaling laws.Proceedings of the National Academy of Sciences121, 27 (2024), e2311878121
2024
-
[12]
Kamalika Chatterjee, Yan Li, Hochan Chang, Mohsen Damadam, Pouya Asrar, Jaechoon Kim, Glen Jeong, and WooPoung Kim. 2024. Thermal and mechanical simulations of 3D packages with custom high band- width memory (HBM). In2024 IEEE 74th Electronic Components and Technology Conference (ECTC). IEEE, 1054–1059
2024
-
[13]
Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher M De Sa. 2023. Quip: 2-bit quantization of large language models with guarantees.Advances in Neural Information Processing Systems36 (2023), 4396–4429
2023
-
[14]
Yuzong Chen, Ahmed F AbouElhamayed, Xilai Dai, Yang Wang, Marta Andronic, George A Constantinides, and Mohamed S Abdelfattah. 2025. Bitmod: Bit-serial mixture-of-datatype llm acceleration. In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 1082–1097
2025
-
[15]
Benjamin Y Cho, Jeageun Jung, and Mattan Erez. 2021. Accelerat- ing bandwidth-bound deep learning inference with main-memory accelerators. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1–14. Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Trovato et al
2021
-
[16]
MH Cho, YI Kim, DS Woo, SW Kim, MS Shim, YJ Park, WS Lee, and BI Ryu. 2006. Analysis of thermal variation of DRAM retention time. In2006 IEEE International Reliability Physics Symposium Proceedings. IEEE, 433–436
2006
-
[17]
Steve Dai, Rangha Venkatesan, Mark Ren, Brian Zimmer, William Dally, and Brucek Khailany. 2021. Vs-quant: Per-vector scaled quantization for accurate low-precision neural network inference.Proceedings of Machine Learning and Systems3 (2021), 873–884
2021
-
[18]
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer
-
[19]
int8 (): 8-bit matrix multiplication for transformers at scale.Advances in neural information processing systems35 (2022), 30318–30332
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale.Advances in neural information processing systems35 (2022), 30318–30332
2022
-
[20]
Chao Fang, Man Shi, Robin Geens, Arne Symons, Zhongfeng Wang, and Marian Verhelst. 2025. Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format. In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 1467–1481
2025
-
[21]
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers.arXiv preprint arXiv:2210.17323(2022)
Pith/arXiv arXiv 2022
-
[22]
Tom Glint, Manu Awasthi, and Joycee Mekie. 2024. Hardware-Software Co-Design of a Collaborative DNN Accelerator for 3D Stacked Mem- ories with Multi-Channel Data. In2024 29th Asia and South Pacific Design Automation Conference (ASP-DAC). IEEE, 454–459
2024
-
[23]
Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. 2023. Olive: Accel- erating large language models via hardware-friendly outlier-victim pair quantization. InProceedings of the 50th Annual International Sym- posium on Computer Architecture. 1–15
2023
-
[24]
Cong Guo, Chen Zhang, Jingwen Leng, Zihan Liu, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. 2022. Ant: Exploiting adaptive numer- ical data type for low-bit deep neural network quantization. In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1414–1433
2022
-
[25]
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko
-
[26]
InProceedings of the IEEE conference on computer vision and pattern recognition
Quantization and training of neural networks for efficient integer- arithmetic-only inference. InProceedings of the IEEE conference on computer vision and pattern recognition. 2704–2713
-
[27]
Jaeyoung Jang, Jun Heo, Yejin Lee, Jaeyeon Won, Seonghak Kim, Sung Jun Jung, Hakbeom Jang, Tae Jun Ham, and Jae W Lee. 2019. Charon: Specialized near-memory processing architecture for clearing dead objects in memory. InProceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture. 726–739
2019
-
[28]
Jaeyong Jang, Yulhwa Kim, Juheun Lee, and Jae-Joon Kim. 2024. Figna: Integer unit-based accelerator design for fp-int gemm preserving numerical accuracy. In2024 IEEE International Symposium on High- Performance Computer Architecture (HPCA). IEEE, 760–773
2024
-
[29]
Yongkweon Jeon, Chungman Lee, Eulrang Cho, and Yeonju Ro. 2022. Mr. biq: Post-training non-uniform quantization based on minimizing the reconstruction error. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 12329–12338
2022
-
[30]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825(2023)
Pith/arXiv arXiv 2023
-
[31]
Hongshin Jun, Jinhee Cho, Kangseol Lee, Ho-Young Son, Kwiwook Kim, Hanho Jin, and Keith Kim. 2017. Hbm (high bandwidth memory) dram technology and architecture. In2017 IEEE International Memory Workshop (IMW). IEEE, 1–4
2017
-
[32]
Vijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan, Amogh Manjunath, Timothy G Rogers, Tor M Aamodt, and Nikos Hardavel- las. 2021. AccelWattch: A power modeling framework for modern GPUs. InMICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture. 738–753
2021
-
[33]
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Ben- jamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361(2020)
Pith/arXiv arXiv 2020
-
[34]
Liu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks, Vikas Chandra, Utku Diril, Amin Firoozshahian, Kim Hazelwood, Bill Jia, Hsien-Hsin S Lee, et al . 2020. Recnmp: Accelerating personalized recommendation with near-memory processing. In2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 790–803
2020
-
[35]
Liu Ke, Xuan Zhang, Jinin So, Jong-Geon Lee, Shin-Haeng Kang, Sukhan Lee, Songyi Han, YeonGon Cho, Jin Hyun Kim, Yongsuk Kwon, et al. 2021. Near-memory processing in action: Accelerating personal- ized recommendation with axdimm.IEEE Micro42, 1 (2021), 116–127
2021
-
[36]
Mahmoud Khairy, Zhesheng Shen, Tor M Aamodt, and Timothy G Rogers. 2020. Accel-sim: An extensible simulation framework for validated gpu modeling. In2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 473–486
2020
-
[37]
Gwangsun Kim, Niladrish Chatterjee, Mike O’Connor, and Kevin Hsieh
-
[38]
InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Toward standardized near-data processing with unrestricted data placement for GPUs. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1–12
-
[39]
Jongmin Kim, Sungmin Yun, Hyesung Ji, Wonseok Choi, Sangpyo Kim, and Jung Ho Ahn. 2025. Anaheim: Architecture and Algorithms for Processing Fully Homomorphic Encryption in Memory. In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 1158–1173
2025
-
[40]
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. 2023. Squeezellm: Dense-and-sparse quantization.arXiv preprint arXiv:2306.07629(2023)
Pith/arXiv arXiv 2023
-
[41]
Taesu Kim, Jongho Lee, Daehyun Ahn, Sarang Kim, Jiwoong Choi, Minkyu Kim, and Hyungjun Kim. 2024. QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference.arXiv preprint arXiv:2402.10076(2024)
Pith/arXiv arXiv 2024
-
[42]
Yulhwa Kim, Jaeyong Jang, Jehun Lee, Jihoon Park, Jeonghoon Kim, Byeongwook Kim, Se Jung Kwon, Dongsoo Lee, et al. 2023. Winning both the accuracy of floating point activation and the simplicity of integer arithmetic. InThe Eleventh International Conference on Learning Representations
2023
-
[43]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica
-
[44]
InProceedings of the 29th symposium on operating systems principles
Efficient memory management for large language model serving with pagedattention. InProceedings of the 29th symposium on operating systems principles. 611–626
-
[45]
Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park. 2024. Owq: Outlier-aware weight quantization for efficient fine- tuning and inference of large language models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 13355–13364
2024
-
[46]
Yiwei Li, Boyu Tian, Yi Ren, and Mingyu Gao. 2024. Stream-Based Data Placement for Near-Data Processing with Extended Memory. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1648–1662
2024
-
[47]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei- Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems6 (2024), 87–100
2024
-
[48]
Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, and Song Han. 2024. Qserve: W4a8kv4 quanti- zation and system co-design for efficient llm serving.arXiv preprint arXiv:2405.04532(2024). StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration Conference acronym ’XX, June 03–05, 2018, Wo...
Pith/arXiv arXiv 2024
-
[49]
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. DeepSeek-V3 Technical Report.CoRR(2024)
2024
-
[50]
Zhiwen Mo, Lei Wang, Jianyu Wei, Zhichen Zeng, Shijie Cao, Lingxiao Ma, Naifeng Jing, Ting Cao, Jilong Xue, Fan Yang, et al . 2025. LUT Tensor Core: A Software-Hardware Co-Design for LUT-Based Low- Bit LLM Inference. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 514–528
2025
-
[51]
Hiroyuki Ootomo and Akira Naruse. 2023. Custom 8-bit floating point value format for reducing shared memory bank conflict in approximate nearest neighbor search.arXiv preprint arXiv:2301.06672(2023)
Pith/arXiv arXiv 2023
-
[52]
Andrew Or, Apurva Jain, Daniel Vega-Myhre, Jesse Cai, Charles David Hernandez, Zhenrui Zheng, Driss Guessous, Vasiliy Kuznetsov, Chris- tian Puhrsch, Mark Saroufim, et al . 2025. TorchAO: PyTorch- Native Training-to-Serving Model Optimization.arXiv preprint arXiv:2507.16099(2025)
Pith/arXiv arXiv 2025
-
[53]
Gunho Park, Hyeokjun Kwon, Jiwoo Kim, Jeongin Bae, Baeseong Park, Dongsoo Lee, and Youngjoo Lee. 2025. FIGLUT: An Energy- Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables. In2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 1098–1111
2025
-
[54]
Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee. 2022. Lut-gemm: Quantized matrix multiplication based on luts for efficient inference in large-scale generative language models.arXiv preprint arXiv:2206.09557(2022)
Pith/arXiv arXiv 2022
-
[55]
Myeong-Jae Park, Jinhyung Lee, Kyungjun Cho, Jihwan Park, Junil Moon, Sung-Hak Lee, Tae-Kyun Kim, Sanghoon Oh, Seokwoo Choi, Yongsuk Choi, et al. 2022. A 192-Gb 12-high 896-GB/s HBM3 DRAM with a TSV auto-calibration scheme and machine-learning-based lay- out optimization.IEEE Journal of Solid-State Circuits58, 1 (2022), 256–269
2022
-
[56]
Javier Picorel, Djordje Jevdjic, and Babak Falsafi. 2017. Near-memory address translation. In2017 26th International Conference on Parallel Architectures and Compilation Techniques (PACT). Ieee, 303–317
2017
-
[57]
Lance Saldanha and Roman Lysecky. 2009. Float-to-fixed and fixed-to- float hardware converters for rapid hardware/software partitioning of floating point software applications to static and dynamic fixed point coprocessors.Design automation for embedded systems13, 3 (2009), 139–157
2009
-
[58]
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. 2023. Omniquant: Omnidirectionally calibrated quantization for large lan- guage models.arXiv preprint arXiv:2308.13137(2023)
Pith/arXiv arXiv 2023
-
[59]
Gian Singh and Sarma Vrudhula. 2024. A DRAM-based near-memory architecture for accelerated and energy-efficient execution of trans- formers. InProceedings of the Great Lakes Symposium on VLSI 2024. 57–62
2024
-
[60]
Jaihyuk Song. 2025. AI Revolution Driven by Memory Technology Innovation. In2025 IEEE International Solid-State Circuits Conference (ISSCC), Vol. 68. IEEE, 26–36
2025
-
[61]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie- Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971(2023)
Pith/arXiv arXiv 2023
-
[62]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. At- tention is all you need.Advances in neural information processing systems30 (2017)
2017
-
[63]
Oreste Villa, Mark Stephenson, David Nellans, and Stephen W Keck- ler. 2019. Nvbit: A dynamic binary instrumentation framework for nvidia gpus. InProceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture. 372–383
2019
-
[64]
Jianyu Wei, Shijie Cao, Ting Cao, Lingxiao Ma, Lei Wang, Yanyong Zhang, and Mao Yang. 2025. T-mac: Cpu renaissance via table lookup for low-bit llm deployment on edge. InProceedings of the Twentieth European Conference on Computer Systems. 278–292
2025
-
[65]
Xiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong, Shang- hang Zhang, Qi Zhang, Fengwei Yu, and Xianglong Liu. 2022. Outlier suppression: Pushing the limit of low-bit transformer language models. Advances in Neural Information Processing Systems35 (2022), 17402– 17414
2022
-
[66]
Christian Weis, Matthias Jung, Peter Ehses, Cristiano Santos, Pascal Vivet, Sven Goossens, Martijn Koedam, and Norbert Wehn. 2015. Re- tention time measurements and modelling of bit error rates of WIDE I/O DRAM in MPSoCs. In2015 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 495–500
2015
-
[67]
Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, et al . 2024. Quant-LLM: accelerating the serving of large language models via FP6-centric algorithm-system co-design on modern GPUs. InProceedings of the 2024 USENIX Conference on Usenix Annual Technical Conference. 699–713
2024
-
[68]
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. InInternational conference on machine learning. PMLR, 38087–38099
2023
-
[69]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025)
Pith/arXiv arXiv 2025
-
[70]
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022. Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.Advances in neural information processing systems35 (2022), 27168–27183
2022
-
[71]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068(2022)
Pith/arXiv arXiv 2022
-
[72]
Yu Zhang, Mingzi Wang, Lancheng Zou, Wulong Liu, Hui-Ling Zhen, Mingxuan Yuan, and Bei Yu. 2024. Mixpe: Quantization and hardware co-design for efficient llm inference.arXiv preprint arXiv:2411.16158 (2024)
Pith/arXiv arXiv 2024
-
[73]
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024. Atom: Low-bit quantization for efficient and accurate llm serving.Proceedings of Machine Learning and Systems6 (2024), 196–209
2024
-
[74]
Zhe Zhou, Cong Li, Fan Yang, and Guangyu Sun. 2023. Dimm-link: Enabling efficient inter-dimm communication for near-memory pro- cessing. In2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 302–316
2023
-
[75]
Kan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Zihao Ye, Keisuke Kamahori, Chien-Yu Lin, et al. 2025. {NanoFlow}: Towards optimal large language model serving through- put. In19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25). 749–765
2025
-
[76]
Jiaxiang Zou, Yonghao Chen, Xingyu Chen, Chenxi Xu, and Xinyu Chen. 2025. AxCore: A Quantization-Aware Approximate GEMM Unit for LLM Inference. InProceedings of the 58th IEEE/ACM International Symposium on Microarchitecture®. 839–853
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.