REVIEW 3 major objections 5 minor 1 cited by
A single continuous offset per layer lets LLMs learn mixed MXFP bit-widths that beat uniform low-precision and greedy heuristics.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 10:56 UTC pith:DHJVQG66
load-bearing objection Solid first gradient-based MXFP mixed-precision PTQ method; Pareto gains look real on the reported 1–2B models, but single-run small-scale evidence is the main soft spot. the 3 major comments →
dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Parameterizing every layer’s MXFP format by a single continuous offset β, annealing it with a temperature-controlled sigmoid, and regularizing the average bit-width yields mixed-precision models that dominate both homogeneous MXFP baselines and KL-divergence layer-selection heuristics on perplexity and zero-shot accuracy for Llama, Qwen3 and SmolLM2.
What carries the argument
The shared offset β that defines the format E(2+β)M(1+β) (or the MXFP6/MXFP4 variant), mapped through a temperature-annealed sigmoid so that continuous values used in the forward pass gradually collapse onto the two hardware-legal endpoints.
Load-bearing premise
That one continuous offset per layer plus a temperature schedule and a simple average-bit-width penalty is enough to capture how quantization errors interact across layers and still land on legal hardware formats.
What would settle it
On the same 1–2 B models and calibration budget, either a KL-ranked mixed-precision assignment or a rounding-plus-STE baseline matching or beating dMX’s Pareto front at intermediate average bit-widths (roughly 5–7).
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes dMX, a gradient-based post-training framework for mixed-precision assignment over OCP microscaling floating-point formats (MXFP8/MXFP6/MXFP4). Each layer is parameterized by a single continuous offset β that defines a format of the form E(2+β)M(1+β) (or the MXFP6/MXFP4 variant), with weight and activation bit-widths tied. Continuous formats are used during calibration; a temperature-annealed sigmoid progressively maps offsets to the discrete hardware endpoints. A target-aware regularizer on average bit-width (simple or size-weighted) steers the budget. Calibration jointly optimizes β and SpinQuant-style rotations on 3200 FineWeb samples (400 steps). On Llama 3.2 1B, Qwen3 1.7B and SmolLM2 1.7B the method reports lower WikiText-2 perplexity and higher average zero-shot accuracy than homogeneous MXFP baselines and a KL-divergence layer-selection heuristic at intermediate average bit-widths, with ablations of regularization form, continuous vs ROUND+STE, and format pairs. Appendix A supplies closed-form STE gradients that match autograd.
Significance. If the empirical Pareto claim holds under modest robustness checks, the work is a useful and timely contribution: it is the first gradient-based bit-width allocation method specialized to MX floating-point formats, supplies a clean continuous parameterization plus annealing that avoids the oscillations of naive rounding+STE, and is compatible with existing PTQ pipelines (Brevitas, SpinQuant rotations). The closed-form gradients in Appendix A and the systematic ablations (target-aware vs simple penalty, continuous vs ROUND+STE, learned vs KL) are concrete strengths. The practical impact is currently limited by the 1–2B model scale and the use of average bit-width as the sole cost proxy, both of which the authors flag as future work.
major comments (3)
- [Tables 1-2, Figures 3-5] Tables 1–2 and Figures 3–5/7–8: the central Pareto-dominance claim (vs homogeneous MXFP and vs KL pre-selection at intermediate bit-widths ~5–7) rests on single calibration runs with no multi-seed error bars or reported variance. Given free parameters (T schedule, Tratio=60%, λ=5, SGD lr=1 for β, 400-step budget) and the discrete nature of the final assignment, modest seed or hyper-parameter variation could move the intermediate-bit-width advantage inside run-to-run noise. At least a small multi-seed study (or sensitivity sweeps) on one model is needed before the claim can be treated as robust.
- [Section 5, Abstract] Sec. 5 and experimental setting: all results are on models ≤1.7B. Cross-layer quantization interactions and the value of end-to-end β optimization may change at larger scale or for MoE architectures. The paper correctly lists scaling as future work, but the abstract and introduction state the Pareto claim without that qualifier; either add a clear scope statement or provide at least one larger-model check so the claim is not over-generalized.
- [Section 2.3] Sec. 2.3: average bit-width (simple or size-weighted) is used as the sole proxy for inference cost. On real MX hardware, latency/energy can depend on format-specific throughput, memory hierarchy, and activation vs weight traffic in ways that a scalar average does not capture. The paper acknowledges this as a coarse proxy; a short discussion of when the proxy is expected to be adequate (or a simple hardware-aware alternative) would strengthen the deployment claim.
minor comments (5)
- [Abstract, Section 2.1] Abstract and Sec. 2.1: repeated word 'format format' and occasional missing spaces around equations; a light copy-edit pass would help.
- [Figure 1] Figure 1 caption and surrounding text: the continuous-grid illustration is useful but the discrete E2M1 reference is only briefly mentioned; a short sentence clarifying what continuous e,m produce would aid readers unfamiliar with the construction.
- [Appendix B] Appendix B: hyper-parameter values (T from 8 to 400, Tratio=60%, λ=5, bit-width SGD lr=1) are listed; stating whether any of these were tuned per model or held fixed would improve reproducibility.
- [Section 4] Related Work (Sec. 4): MicroMix is cited for channel-wise MX assignment; a one-sentence contrast of layer-wise continuous β vs channel-wise heuristics would make the novelty claim sharper.
- [Tables 1-2] Tables 1–2 report only a subset of targets (4.5/5/6/8); the full curves are in the appendix. Cross-referencing the appendix figures more explicitly in the main text would help readers locate the complete Pareto fronts.
Circularity Check
No significant circularity: dMX is a continuous optimization procedure whose quality metrics are independent of the bit-width regularizer and parameterization.
full rationale
The paper formulates per-layer MXFP bit-width assignment as gradient-based minimization of task loss plus a target-aware average-bit-width regularizer (Eqs. 5–9), with continuous offset β and temperature annealing used only as a differentiable search mechanism that is later discretized. Final claims are empirical Pareto comparisons of held-out WikiText-2 perplexity and zero-shot accuracy against homogeneous MXFP baselines and a KL-preselection heuristic (Tables 1–2, Figs. 3–5, 7–8). These metrics are not defined by, nor statistically forced by, the fitted β values or the regularizer; the regularizer merely steers the cost axis. Appendix A derives STE gradients of the continuous FP quantizer from first principles (IEEE-style range and ULP) without circular appeal to the experimental outcomes. Self-citations (Brevitas library, SpinQuant-style rotations) supply implementation infrastructure only and do not underwrite uniqueness or force the Pareto results. No step reduces a claimed prediction or first-principles result to its own inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- Temperature schedule (T_init, T_final, Tratio, exponential ramp)
- Regularization strength λ and target bit-width set
- Bit-width optimizer learning rate (SGD lr=1) and calibration budget (3200 samples / 400 steps)
- Initial β configuration (all layers start at MXFP4)
axioms (5)
- domain assumption Straight-through estimator: derivative of round/floor is treated as 1 while forward values are preserved (Appendix A).
- domain assumption OCP MXFP formats are restricted to the discrete set {E4M3, E2M3, E2M1} (and the paper’s E(2+β)M(1+β) interpolation).
- ad hoc to paper Average bit-width (simple or size-weighted) is a sufficient coarse proxy for inference cost.
- ad hoc to paper Tying weight and activation bit-widths per layer and using a single scalar β is adequate.
- standard math Quantization equations remain well-defined and usefully differentiable for real-valued e, m (or β).
invented entities (2)
-
Continuous MXFP(β) offset parameterization E(2+β)M(1+β)
no independent evidence
-
Temperature-regulated sigmoid annealing F(β̂, T) for format discretization
no independent evidence
read the original abstract
Quantizing large language models (LLMs) to low-precision floating-point representations is central to efficient deployment, yet applying a single bit-width uniformly across all layers is sub-optimal in terms of both performance and accuracy. This work introduces dMX, a differentiable mixed-precision quantization framework for learnable floating-point bit-width assignment. We study its application for the microscaling floating-point (MXFP) family of data types defined by the Open Compute Project (OCP) standard. The per-layer bit-width assignment is formulated as a continuous optimization problem in which each layer's floating-point format format is parameterized by a scalar parameter, folding the multi-variate design space into a single learnable offset. During training this offset takes continuous values, avoiding sudden oscillations between discrete quantization formats. A temperature-based annealing schedule progressively discretizes the learned offsets, ensuring that the final configuration maps to hardware-compatible MXFP formats without abrupt transitions between training and inference behavior. A target-aware regularization term steers the average bit-width toward a user-specified budget, serving as a coarse-grained proxy for inference cost and balancing model quality against deployment efficiency. We performed experiments on different families of LLM, such as Llama, Qwen3, and SmolLM2, evaluating perplexity on WikiText-2 and accuracy on four zero-shot reasoning benchmarks. Across these settings, dMX consistently yields Pareto-dominating models and improves over Kullback-Leibler (KL) divergence-based layer-selection heuristics, efficiently navigating trade-offs between model quality and average bit-width.
Figures
Forward citations
Cited by 1 Pith paper
-
Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models
Learned per-group bit-widths yield a reusable low-bit recipe that makes language models simultaneously larger in parameters and smaller in storage than FP16 baselines, with growing decode speedups.
Reference graph
Works this paper leans on
-
[1]
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: Accurate post-training quantization for generative pre-trained transformers.arXiv preprint arXiv:2210.17323, 2022
Pith/arXiv arXiv 2022
-
[2]
AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. InProceedings of Machine Learning and Systems, 2024
2024
-
[3]
SmoothQuant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. SmoothQuant: Accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning, 2023
2023
-
[4]
FP8 versus INT8 for efficient deep learning inference.arXiv preprint arXiv:2303.17951, 2023
Mart van Baalen, Andrey Kuzmin, Suparna S Nair, Yuwei Ren, Eric Mahurin, Chirag Patel, Sundar Subramanian, Sanghyuk Lee, Markus Nagel, Joseph Soriaga, and Tijmen Blankevoort. FP8 versus INT8 for efficient deep learning inference.arXiv preprint arXiv:2303.17951, 2023
Pith/arXiv arXiv 2023
-
[5]
Microscaling data formats for deep learning.arXiv preprint arXiv:2310.10537, 2023
Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, et al. Microscaling data formats for deep learning.arXiv preprint arXiv:2310.10537, 2023
Pith/arXiv arXiv 2023
-
[6]
OCP microscaling formats (MX) specification
Bita Darvish Rouhani, Nitin Garegrat, Tom Savell, Ankit More, Kyung-Nam Han, Ritchie Zhao, Mathew Hall, Jasmine Klar, Eric Chung, Yuan Yu, Michael Schulte, Ralph Wittig, Ian Bratt, Nigel Stephens, Jelena Milanovic, John Brothers, Pradeep Dubey, Marius Cornea, Alexander Heinecke, Andres Rodriguez, Martin Langhammer, Summer Deng, Maxim Naumov, Paulius Micik...
2023
-
[7]
Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh
Vage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Bridging the gap between promise and performance for microscaling fp4 quantization, 2025
2025
-
[8]
MixQuant: Pushing the limits of block rotations in post-training quantization
Sai Sanjeet, Ian Colbert, Pablo Monteagudo-Lago, Giuseppe Franco, Yaman Umuroglu, and Nicholas J Fraser. MixQuant: Pushing the limits of block rotations in post-training quantization. arXiv preprint arXiv:2601.22347, 2026
Pith/arXiv arXiv 2026
-
[9]
Gradient-free training of quantized neural networks.arXiv preprint arXiv:2410.09734, 2024
Noa Cohen, Omkar Joglekar, Dotan Di Castro, Vladimir Tchuiev, Shir Kozlovsky, and Michal Moshkovitz. Gradient-free training of quantized neural networks.arXiv preprint arXiv:2410.09734, 2024
arXiv 2024
-
[10]
Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer. Mixed precision quantization of ConvNets via differentiable neural architecture search.arXiv preprint arXiv:1812.00090, 2018
Pith/arXiv arXiv 2018
-
[11]
InfoQ: Mixed-precision quantization via global information flow
Mehmet Emre Akbulut, Hazem Hesham Yousef Shalby, Fabrizio Pittorino, and Manuel Roveri. InfoQ: Mixed-precision quantization via global information flow. InProceedings of the AAAI Conference on Artificial Intelligence, 2026
2026
-
[12]
Mix-QSAM: Mixed-precision quantization of the segment anything model
Navin Ranjan and Andreas Savakis. Mix-QSAM: Mixed-precision quantization of the segment anything model. InIEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2025
2025
-
[13]
Mahoney, and Kurt Keutzer
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. HAWQ: Hessian AWare quantization of neural networks with mixed-precision. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 293–302, 2019
2019
-
[14]
Mahoney, and Kurt Keutzer
Zhen Dong, Zhewei Yao, Yaohui Cai, Daiyaan Arfeen, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. HAWQ-V2: Hessian aware trace-weighted quantization of neural networks. InAdvances in Neural Information Processing Systems, volume 33, 2020
2020
-
[15]
FracBits: Mixed precision quantization via fractional bit-widths
Linjie Yang and Qing Jin. FracBits: Mixed precision quantization via fractional bit-widths. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10612–10620, 2021. 11
2021
-
[16]
SDQ: Stochastic differentiable quantization with mixed precision
Xijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu, Xianghong Hu, Jeffry Wicaksana, Eric Xing, and Kwang-Ting Cheng. SDQ: Stochastic differentiable quantization with mixed precision. InProceedings of the 39th International Conference on Machine Learning, pages 9295–9309, 2022
2022
-
[17]
BSQ: Exploring bit-level sparsity for mixed-precision neural network quantization
Huanrui Yang, Lin Duan, Yiran Chen, and Hai Li. BSQ: Exploring bit-level sparsity for mixed-precision neural network quantization. InProceedings of the International Conference on Learning Representations, 2021
2021
-
[18]
Yoshua Bengio, Nicholas L´eonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation.arXiv preprint arXiv:1308.3432, 2013
Pith/arXiv arXiv 2013
-
[19]
Spinquant: Llm quantization with learned rotations.arXiv preprint arXiv:2405.16406, 2024
Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort. Spinquant: Llm quantization with learned rotations.arXiv preprint arXiv:2405.16406, 2024
Pith/arXiv arXiv 2024
-
[20]
FP8 formats for deep learning.arXiv preprint arXiv:2209.05433, 2022
Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellem- pudi, Stuart Oberman, Mohammad Shoeybi, Michael Siu, and Hao Wu. FP8 formats for deep learning.arXiv preprint arXiv:2209.05433, 2022
Pith/arXiv arXiv 2022
-
[21]
The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Pith/arXiv arXiv 2024
-
[22]
Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025
Pith/arXiv arXiv 2025
-
[23]
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Mart´ın Bl´azquez, Guilherme Penedo, Lewis Tunstall, Andr´es Marafioti, Hynek Kydl´ıˇcek, Agust´ın Piqueres Lajar´ın, Vaibhav Srivastav, et al. SmolLM2: When smol goes big – data-centric training of a small language model.arXiv preprint arXiv:2502.02737, 2025
Pith/arXiv arXiv 2025
-
[24]
Guilherme Penedo, Hynek Kydl´ıˇcek, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro V on Werra, and Thomas Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale.arXiv preprint arXiv:2406.17557, 2024
Pith/arXiv arXiv 2024
-
[25]
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. InProceedings of the International Conference on Learning Representations, 2017
2017
-
[26]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018
Pith/arXiv arXiv 2018
-
[27]
HellaSwag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019
2019
-
[28]
WinoGrande: An adversarial Winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: An adversarial Winograd schema challenge at scale. InProceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8732–8740, 2020
2020
-
[29]
LightEval: A lightweight framework for LLM evaluation, 2023
Nathan Habib, Cl ´ementine Fourrier, Hynek Kydl ´ıˇcek, Thomas Wolf, and Lewis Tunstall. LightEval: A lightweight framework for LLM evaluation, 2023
2023
-
[30]
Xilinx/brevitas, 2025
Giuseppe Franco, Alessandro Pappalardo, and Nicholas J Fraser. Xilinx/brevitas, 2025
2025
-
[31]
HAQ: Hardware-aware automated quantization with mixed precision
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. HAQ: Hardware-aware automated quantization with mixed precision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8612–8620, 2019. 12
2019
-
[32]
Mahoney, and Kurt Keutzer
Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael W. Mahoney, and Kurt Keutzer. HAWQ-V3: Dyadic neural network quantization. InProceedings of the 38th International Conference on Machine Learning, pages 11875–11886, 2021
2021
-
[33]
Towards mixed-precision quantization of neural networks via constrained optimization
Weihan Chen, Peisong Wang, and Jian Cheng. Towards mixed-precision quantization of neural networks via constrained optimization. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5350–5359, 2021
2021
-
[34]
APTQ: Attention- aware post-training mixed-precision quantization for large language models
Ziyi Guan, Hantao Huang, Yupeng Su, Hong Huang, Ngai Wong, and Hao Yu. APTQ: Attention- aware post-training mixed-precision quantization for large language models. InProceedings of the 61st IEEE/ACM Design Automation Conference, 2024
2024
-
[35]
Utkarsh Saxena, Sayeh Sharify, Kaushik Roy, and Xin Wang. ResQ: Mixed-precision quan- tization of large language models with low-rank residuals.arXiv preprint arXiv:2412.14363, 2024
Pith/arXiv arXiv 2024
-
[36]
Rethinking differentiable search for mixed-precision neural networks
Zhaowei Cai and Nuno Vasconcelos. Rethinking differentiable search for mixed-precision neural networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2346–2355, 2020
2020
-
[37]
Zhexin Li, Tong Yang, Peisong Wang, and Jian Cheng. Q-ViT: Fully differentiable quantization for vision transformer.arXiv preprint arXiv:2201.07703, 2022
Pith/arXiv arXiv 2022
-
[38]
Jennings, and Arnon Netzer
Hai Victor Habi, Roy H. Jennings, and Arnon Netzer. HMQ: Hardware friendly mixed precision quantization block for CNNs. InComputer Vision – ECCV 2020, pages 448–463. Springer, 2020
2020
-
[39]
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016
Pith/arXiv arXiv 2016
-
[40]
Maddison, Andriy Mnih, and Yee Whye Teh
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. The concrete distribution: A continuous relaxation of discrete random variables. InProceedings of the International Conference on Learning Representations, 2017
2017
-
[41]
Bayesian bits: Unifying quantization and pruning
Mart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad, Ying Wang, Tijmen Blankevoort, and Max Welling. Bayesian bits: Unifying quantization and pruning. InAdvances in Neural Information Processing Systems, volume 33, 2020
2020
-
[42]
Mixed precision DNNs: All you need is a good parametrization
Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso Garcia, Stephen Tiedemann, Thomas Kemp, and Akira Nakamura. Mixed precision DNNs: All you need is a good parametrization. InProceedings of the International Conference on Learning Representations, 2020
2020
-
[43]
Micromix: Efficient mixed-precision quantization with microscaling formats for large language models
Wenyuan Liu, Haoqian Meng, Yilun Luo, Yafei Zhao, Peng Zhang, and Xindian Ma. Micromix: Efficient mixed-precision quantization with microscaling formats for large language models. arXiv preprint arXiv:2508.02343, 2025
arXiv 2025
-
[44]
Mixture compressor for mixture-of-experts LLMs gains more
Wei Huang, Yue Liao, Jianhui Liu, Ruifei He, Haoru Tan, Shiming Zhang, Hongsheng Li, Si Liu, and Xiaojuan Qi. Mixture compressor for mixture-of-experts LLMs gains more. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[45]
Ieee standard for floating-point arithmetic.IEEE Std 754-2019 (Revision of IEEE 754-2008), pages 1–84, 2019
IEEE. Ieee standard for floating-point arithmetic.IEEE Std 754-2019 (Revision of IEEE 754-2008), pages 1–84, 2019
2019
-
[46]
Jain, Albert Gural, Michael Wu, and Chris H
Sambhav R. Jain, Albert Gural, Michael Wu, and Chris H. Dick. Trained quantization thresholds for accurate and efficient fixed-point inference of deep neural networks. InProceedings of the 3rd Machine Learning and Systems (MLSys) Conference, 2020
2020
-
[47]
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K¨opf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, high-perf...
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.