Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

A single continuous offset per layer lets LLMs learn mixed MXFP bit-widths that beat uniform low-precision and greedy heuristics.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 10:56 UTC pith:DHJVQG66

load-bearing objection Solid first gradient-based MXFP mixed-precision PTQ method; Pareto gains look real on the reported 1–2B models, but single-run small-scale evidence is the main soft spot. the 3 major comments →

arxiv 2606.04115 v2 pith:DHJVQG66 submitted 2026-06-02 cs.LG cs.AI

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats

classification cs.LG cs.AI
keywords mixed-precision quantizationMXFPdifferentiable bit-widthtemperature annealingLLM post-training quantizationfloating-point formatsmicroscaling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Uniform low-precision floating-point formats such as MXFP4 hurt LLM quality, yet choosing different formats per layer is a hard combinatorial search. dMX turns that search into ordinary gradient descent: each layer is given one continuous scalar that smoothly interpolates between hardware-legal MXFP formats, a temperature schedule slowly hardens the choice into a discrete format, and a simple average-bit-width penalty steers the model toward a user budget. On 1–2 B models the resulting mixed-precision assignments sit above the uniform MXFP baselines and above KL-based layer ranking on both perplexity and zero-shot accuracy, especially in the intermediate bit-width band where the budget must be spent carefully. The practical payoff is a short post-training calibration that produces Pareto-better quantized models without full retraining or hand-crafted sensitivity metrics.

Core claim

Parameterizing every layer’s MXFP format by a single continuous offset β, annealing it with a temperature-controlled sigmoid, and regularizing the average bit-width yields mixed-precision models that dominate both homogeneous MXFP baselines and KL-divergence layer-selection heuristics on perplexity and zero-shot accuracy for Llama, Qwen3 and SmolLM2.

What carries the argument

The shared offset β that defines the format E(2+β)M(1+β) (or the MXFP6/MXFP4 variant), mapped through a temperature-annealed sigmoid so that continuous values used in the forward pass gradually collapse onto the two hardware-legal endpoints.

Load-bearing premise

That one continuous offset per layer plus a temperature schedule and a simple average-bit-width penalty is enough to capture how quantization errors interact across layers and still land on legal hardware formats.

What would settle it

On the same 1–2 B models and calibration budget, either a KL-ranked mixed-precision assignment or a rounding-plus-STE baseline matching or beating dMX’s Pareto front at intermediate average bit-widths (roughly 5–7).

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes dMX, a gradient-based post-training framework for mixed-precision assignment over OCP microscaling floating-point formats (MXFP8/MXFP6/MXFP4). Each layer is parameterized by a single continuous offset β that defines a format of the form E(2+β)M(1+β) (or the MXFP6/MXFP4 variant), with weight and activation bit-widths tied. Continuous formats are used during calibration; a temperature-annealed sigmoid progressively maps offsets to the discrete hardware endpoints. A target-aware regularizer on average bit-width (simple or size-weighted) steers the budget. Calibration jointly optimizes β and SpinQuant-style rotations on 3200 FineWeb samples (400 steps). On Llama 3.2 1B, Qwen3 1.7B and SmolLM2 1.7B the method reports lower WikiText-2 perplexity and higher average zero-shot accuracy than homogeneous MXFP baselines and a KL-divergence layer-selection heuristic at intermediate average bit-widths, with ablations of regularization form, continuous vs ROUND+STE, and format pairs. Appendix A supplies closed-form STE gradients that match autograd.

Significance. If the empirical Pareto claim holds under modest robustness checks, the work is a useful and timely contribution: it is the first gradient-based bit-width allocation method specialized to MX floating-point formats, supplies a clean continuous parameterization plus annealing that avoids the oscillations of naive rounding+STE, and is compatible with existing PTQ pipelines (Brevitas, SpinQuant rotations). The closed-form gradients in Appendix A and the systematic ablations (target-aware vs simple penalty, continuous vs ROUND+STE, learned vs KL) are concrete strengths. The practical impact is currently limited by the 1–2B model scale and the use of average bit-width as the sole cost proxy, both of which the authors flag as future work.

major comments (3)
  1. [Tables 1-2, Figures 3-5] Tables 1–2 and Figures 3–5/7–8: the central Pareto-dominance claim (vs homogeneous MXFP and vs KL pre-selection at intermediate bit-widths ~5–7) rests on single calibration runs with no multi-seed error bars or reported variance. Given free parameters (T schedule, Tratio=60%, λ=5, SGD lr=1 for β, 400-step budget) and the discrete nature of the final assignment, modest seed or hyper-parameter variation could move the intermediate-bit-width advantage inside run-to-run noise. At least a small multi-seed study (or sensitivity sweeps) on one model is needed before the claim can be treated as robust.
  2. [Section 5, Abstract] Sec. 5 and experimental setting: all results are on models ≤1.7B. Cross-layer quantization interactions and the value of end-to-end β optimization may change at larger scale or for MoE architectures. The paper correctly lists scaling as future work, but the abstract and introduction state the Pareto claim without that qualifier; either add a clear scope statement or provide at least one larger-model check so the claim is not over-generalized.
  3. [Section 2.3] Sec. 2.3: average bit-width (simple or size-weighted) is used as the sole proxy for inference cost. On real MX hardware, latency/energy can depend on format-specific throughput, memory hierarchy, and activation vs weight traffic in ways that a scalar average does not capture. The paper acknowledges this as a coarse proxy; a short discussion of when the proxy is expected to be adequate (or a simple hardware-aware alternative) would strengthen the deployment claim.
minor comments (5)
  1. [Abstract, Section 2.1] Abstract and Sec. 2.1: repeated word 'format format' and occasional missing spaces around equations; a light copy-edit pass would help.
  2. [Figure 1] Figure 1 caption and surrounding text: the continuous-grid illustration is useful but the discrete E2M1 reference is only briefly mentioned; a short sentence clarifying what continuous e,m produce would aid readers unfamiliar with the construction.
  3. [Appendix B] Appendix B: hyper-parameter values (T from 8 to 400, Tratio=60%, λ=5, bit-width SGD lr=1) are listed; stating whether any of these were tuned per model or held fixed would improve reproducibility.
  4. [Section 4] Related Work (Sec. 4): MicroMix is cited for channel-wise MX assignment; a one-sentence contrast of layer-wise continuous β vs channel-wise heuristics would make the novelty claim sharper.
  5. [Tables 1-2] Tables 1–2 report only a subset of targets (4.5/5/6/8); the full curves are in the appendix. Cross-referencing the appendix figures more explicitly in the main text would help readers locate the complete Pareto fronts.

Circularity Check

0 steps flagged

No significant circularity: dMX is a continuous optimization procedure whose quality metrics are independent of the bit-width regularizer and parameterization.

full rationale

The paper formulates per-layer MXFP bit-width assignment as gradient-based minimization of task loss plus a target-aware average-bit-width regularizer (Eqs. 5–9), with continuous offset β and temperature annealing used only as a differentiable search mechanism that is later discretized. Final claims are empirical Pareto comparisons of held-out WikiText-2 perplexity and zero-shot accuracy against homogeneous MXFP baselines and a KL-preselection heuristic (Tables 1–2, Figs. 3–5, 7–8). These metrics are not defined by, nor statistically forced by, the fitted β values or the regularizer; the regularizer merely steers the cost axis. Appendix A derives STE gradients of the continuous FP quantizer from first principles (IEEE-style range and ULP) without circular appeal to the experimental outcomes. Self-citations (Brevitas library, SpinQuant-style rotations) supply implementation infrastructure only and do not underwrite uniqueness or force the Pareto results. No step reduces a claimed prediction or first-principles result to its own inputs by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

The paper is an empirical optimization method. It inherits standard STE/rounding assumptions and the OCP MX format definitions, introduces a continuous β parameterization and annealing schedule as design choices, and relies on several hand-set hyper-parameters and a bit-width proxy for cost. No new physical entities are postulated.

free parameters (4)
  • Temperature schedule (T_init, T_final, Tratio, exponential ramp)
    Chosen by hand (8→400, Tratio=60%) to approximate linear then step-like mapping; not derived from first principles and affects discretization behavior.
  • Regularization strength λ and target bit-width set
    λ fixed at 5 for target-aware penalty; 17 discrete targets from 4.1 to 8 scanned. Directly controls the reported operating points.
  • Bit-width optimizer learning rate (SGD lr=1) and calibration budget (3200 samples / 400 steps)
    Hand-chosen; paper notes highest targets sometimes fail to match and may need LR retuning.
  • Initial β configuration (all layers start at MXFP4)
    Fixed starting point for all mixed-precision runs; influences the optimization trajectory.
axioms (5)
  • domain assumption Straight-through estimator: derivative of round/floor is treated as 1 while forward values are preserved (Appendix A).
    Standard in QAT/PTQ literature; enables gradient flow through the continuous quantizer but is an approximation.
  • domain assumption OCP MXFP formats are restricted to the discrete set {E4M3, E2M3, E2M1} (and the paper’s E(2+β)M(1+β) interpolation).
    Hardware constraint taken from the OCP MX specification; defines the admissible endpoints of annealing.
  • ad hoc to paper Average bit-width (simple or size-weighted) is a sufficient coarse proxy for inference cost.
    Stated in Sec. 2.3; no latency/energy measurements on real MX hardware are provided.
  • ad hoc to paper Tying weight and activation bit-widths per layer and using a single scalar β is adequate.
    Design choice in Sec. 2.1 that collapses the search space; not proven optimal.
  • standard math Quantization equations remain well-defined and usefully differentiable for real-valued e, m (or β).
    Shown via continuous range bounds and STE gradients in Appendix A; verified against PyTorch autograd.
invented entities (2)
  • Continuous MXFP(β) offset parameterization E(2+β)M(1+β) no independent evidence
    purpose: Folds multi-variate format choice into one learnable scalar that interpolates between hardware MXFP formats during calibration.
    Core modeling device of the paper; independent evidence is only the empirical Pareto curves, not an external physical prediction.
  • Temperature-regulated sigmoid annealing F(β̂, T) for format discretization no independent evidence
    purpose: Smoothly transitions continuous offsets to discrete hardware formats without ROUND+STE oscillation.
    Design contribution; validated only by ablation against ROUND+STE inside this paper.

pith-pipeline@v1.1.0-grok45 · 22729 in / 3423 out tokens · 27979 ms · 2026-07-15T10:56:08.041335+00:00 · methodology

0 comments
read the original abstract

Quantizing large language models (LLMs) to low-precision floating-point representations is central to efficient deployment, yet applying a single bit-width uniformly across all layers is sub-optimal in terms of both performance and accuracy. This work introduces dMX, a differentiable mixed-precision quantization framework for learnable floating-point bit-width assignment. We study its application for the microscaling floating-point (MXFP) family of data types defined by the Open Compute Project (OCP) standard. The per-layer bit-width assignment is formulated as a continuous optimization problem in which each layer's floating-point format format is parameterized by a scalar parameter, folding the multi-variate design space into a single learnable offset. During training this offset takes continuous values, avoiding sudden oscillations between discrete quantization formats. A temperature-based annealing schedule progressively discretizes the learned offsets, ensuring that the final configuration maps to hardware-compatible MXFP formats without abrupt transitions between training and inference behavior. A target-aware regularization term steers the average bit-width toward a user-specified budget, serving as a coarse-grained proxy for inference cost and balancing model quality against deployment efficiency. We performed experiments on different families of LLM, such as Llama, Qwen3, and SmolLM2, evaluating perplexity on WikiText-2 and accuracy on four zero-shot reasoning benchmarks. Across these settings, dMX consistently yields Pareto-dominating models and improves over Kullback-Leibler (KL) divergence-based layer-selection heuristics, efficiently navigating trade-offs between model quality and average bit-width.

Figures

Figures reproduced from arXiv: 2606.04115 by Felix Marty, Giuseppe Franco, Ian Colbert, Nicholas Fraser, Pablo Monteagudo-Lago.

Figure 1
Figure 1. Figure 1: Comparison of the quantization grids when using continuous values for mantissa and [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the dMX pipeline. All blue elements highlight the main contributions of this work. The pre-trained LLM (left) contains a learned continuous offset βi for each layer i, which parameterizes the bit-width used in that layer. During the forward pass these offsets are mapped to discrete format assignments βˆ = F(β, T). A task loss and a user-defined regularization term R on β jointly drive the gradi… view at source ↗
Figure 3
Figure 3. Figure 3: Comparison of the target-aware penalty and the simple scaling penalty for MXFP8–MXFP4 [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Comparison of simple-average and tensor-size-weighted bit-width regularization for [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Learned bit-width optimization vs. KL divergence-based pre-selected layer precision for [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Forward and bitwidth-gradient transfer curves of the inner MX-FP quantizer at [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Comparison of the MXFP8/MXFP4 mixed precision quantization and the MXFP6/MXFP4 [PITH_FULL_IMAGE:figures/full_fig_p018_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Comparison between continuous forward-pass bit-width learning and a discretized forward [PITH_FULL_IMAGE:figures/full_fig_p019_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models

    cs.LG 2026-07 conditional novelty 7.0

    Learned per-group bit-widths yield a reusable low-bit recipe that makes language models simultaneously larger in parameters and smaller in storage than FP16 baselines, with growing decode speedups.

Reference graph

Works this paper leans on

47 extracted references · 16 linked inside Pith · cited by 1 Pith paper

  1. [1]

    GPTQ: Accurate post-training quantization for generative pre-trained transformers.arXiv preprint arXiv:2210.17323, 2022

    Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. GPTQ: Accurate post-training quantization for generative pre-trained transformers.arXiv preprint arXiv:2210.17323, 2022

  2. [2]

    AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. InProceedings of Machine Learning and Systems, 2024

  3. [3]

    SmoothQuant: Accurate and efficient post-training quantization for large language models

    Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. SmoothQuant: Accurate and efficient post-training quantization for large language models. In Proceedings of the 40th International Conference on Machine Learning, 2023

  4. [4]

    FP8 versus INT8 for efficient deep learning inference.arXiv preprint arXiv:2303.17951, 2023

    Mart van Baalen, Andrey Kuzmin, Suparna S Nair, Yuwei Ren, Eric Mahurin, Chirag Patel, Sundar Subramanian, Sanghyuk Lee, Markus Nagel, Joseph Soriaga, and Tijmen Blankevoort. FP8 versus INT8 for efficient deep learning inference.arXiv preprint arXiv:2303.17951, 2023

  5. [5]

    Microscaling data formats for deep learning.arXiv preprint arXiv:2310.10537, 2023

    Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, et al. Microscaling data formats for deep learning.arXiv preprint arXiv:2310.10537, 2023

  6. [6]

    OCP microscaling formats (MX) specification

    Bita Darvish Rouhani, Nitin Garegrat, Tom Savell, Ankit More, Kyung-Nam Han, Ritchie Zhao, Mathew Hall, Jasmine Klar, Eric Chung, Yuan Yu, Michael Schulte, Ralph Wittig, Ian Bratt, Nigel Stephens, Jelena Milanovic, John Brothers, Pradeep Dubey, Marius Cornea, Alexander Heinecke, Andres Rodriguez, Martin Langhammer, Summer Deng, Maxim Naumov, Paulius Micik...

  7. [7]

    Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh

    Vage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Bridging the gap between promise and performance for microscaling fp4 quantization, 2025

  8. [8]

    MixQuant: Pushing the limits of block rotations in post-training quantization

    Sai Sanjeet, Ian Colbert, Pablo Monteagudo-Lago, Giuseppe Franco, Yaman Umuroglu, and Nicholas J Fraser. MixQuant: Pushing the limits of block rotations in post-training quantization. arXiv preprint arXiv:2601.22347, 2026

  9. [9]

    Gradient-free training of quantized neural networks.arXiv preprint arXiv:2410.09734, 2024

    Noa Cohen, Omkar Joglekar, Dotan Di Castro, Vladimir Tchuiev, Shir Kozlovsky, and Michal Moshkovitz. Gradient-free training of quantized neural networks.arXiv preprint arXiv:2410.09734, 2024

  10. [10]

    Mixed precision quantization of ConvNets via differentiable neural architecture search.arXiv preprint arXiv:1812.00090, 2018

    Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer. Mixed precision quantization of ConvNets via differentiable neural architecture search.arXiv preprint arXiv:1812.00090, 2018

  11. [11]

    InfoQ: Mixed-precision quantization via global information flow

    Mehmet Emre Akbulut, Hazem Hesham Yousef Shalby, Fabrizio Pittorino, and Manuel Roveri. InfoQ: Mixed-precision quantization via global information flow. InProceedings of the AAAI Conference on Artificial Intelligence, 2026

  12. [12]

    Mix-QSAM: Mixed-precision quantization of the segment anything model

    Navin Ranjan and Andreas Savakis. Mix-QSAM: Mixed-precision quantization of the segment anything model. InIEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2025

  13. [13]

    Mahoney, and Kurt Keutzer

    Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. HAWQ: Hessian AWare quantization of neural networks with mixed-precision. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 293–302, 2019

  14. [14]

    Mahoney, and Kurt Keutzer

    Zhen Dong, Zhewei Yao, Yaohui Cai, Daiyaan Arfeen, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. HAWQ-V2: Hessian aware trace-weighted quantization of neural networks. InAdvances in Neural Information Processing Systems, volume 33, 2020

  15. [15]

    FracBits: Mixed precision quantization via fractional bit-widths

    Linjie Yang and Qing Jin. FracBits: Mixed precision quantization via fractional bit-widths. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10612–10620, 2021. 11

  16. [16]

    SDQ: Stochastic differentiable quantization with mixed precision

    Xijie Huang, Zhiqiang Shen, Shichao Li, Zechun Liu, Xianghong Hu, Jeffry Wicaksana, Eric Xing, and Kwang-Ting Cheng. SDQ: Stochastic differentiable quantization with mixed precision. InProceedings of the 39th International Conference on Machine Learning, pages 9295–9309, 2022

  17. [17]

    BSQ: Exploring bit-level sparsity for mixed-precision neural network quantization

    Huanrui Yang, Lin Duan, Yiran Chen, and Hai Li. BSQ: Exploring bit-level sparsity for mixed-precision neural network quantization. InProceedings of the International Conference on Learning Representations, 2021

  18. [18]

    Estimating or propagating gradients through stochastic neurons for conditional computation.arXiv preprint arXiv:1308.3432, 2013

    Yoshua Bengio, Nicholas L´eonard, and Aaron Courville. Estimating or propagating gradients through stochastic neurons for conditional computation.arXiv preprint arXiv:1308.3432, 2013

  19. [19]

    Spinquant: Llm quantization with learned rotations.arXiv preprint arXiv:2405.16406, 2024

    Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort. Spinquant: Llm quantization with learned rotations.arXiv preprint arXiv:2405.16406, 2024

  20. [20]

    FP8 formats for deep learning.arXiv preprint arXiv:2209.05433, 2022

    Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellem- pudi, Stuart Oberman, Mohammad Shoeybi, Michael Siu, and Hao Wu. FP8 formats for deep learning.arXiv preprint arXiv:2209.05433, 2022

  21. [21]

    The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

  22. [22]

    Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

  23. [23]

    SmolLM2: When smol goes big – data-centric training of a small language model.arXiv preprint arXiv:2502.02737, 2025

    Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Mart´ın Bl´azquez, Guilherme Penedo, Lewis Tunstall, Andr´es Marafioti, Hynek Kydl´ıˇcek, Agust´ın Piqueres Lajar´ın, Vaibhav Srivastav, et al. SmolLM2: When smol goes big – data-centric training of a small language model.arXiv preprint arXiv:2502.02737, 2025

  24. [24]

    The FineWeb datasets: Decanting the web for the finest text data at scale.arXiv preprint arXiv:2406.17557, 2024

    Guilherme Penedo, Hynek Kydl´ıˇcek, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro V on Werra, and Thomas Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale.arXiv preprint arXiv:2406.17557, 2024

  25. [25]

    Pointer sentinel mixture models

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. InProceedings of the International Conference on Learning Representations, 2017

  26. [26]

    Think you have solved question answering? try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try ARC, the AI2 reasoning challenge.arXiv preprint arXiv:1803.05457, 2018

  27. [27]

    HellaSwag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

  28. [28]

    WinoGrande: An adversarial Winograd schema challenge at scale

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: An adversarial Winograd schema challenge at scale. InProceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8732–8740, 2020

  29. [29]

    LightEval: A lightweight framework for LLM evaluation, 2023

    Nathan Habib, Cl ´ementine Fourrier, Hynek Kydl ´ıˇcek, Thomas Wolf, and Lewis Tunstall. LightEval: A lightweight framework for LLM evaluation, 2023

  30. [30]

    Xilinx/brevitas, 2025

    Giuseppe Franco, Alessandro Pappalardo, and Nicholas J Fraser. Xilinx/brevitas, 2025

  31. [31]

    HAQ: Hardware-aware automated quantization with mixed precision

    Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. HAQ: Hardware-aware automated quantization with mixed precision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8612–8620, 2019. 12

  32. [32]

    Mahoney, and Kurt Keutzer

    Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael W. Mahoney, and Kurt Keutzer. HAWQ-V3: Dyadic neural network quantization. InProceedings of the 38th International Conference on Machine Learning, pages 11875–11886, 2021

  33. [33]

    Towards mixed-precision quantization of neural networks via constrained optimization

    Weihan Chen, Peisong Wang, and Jian Cheng. Towards mixed-precision quantization of neural networks via constrained optimization. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5350–5359, 2021

  34. [34]

    APTQ: Attention- aware post-training mixed-precision quantization for large language models

    Ziyi Guan, Hantao Huang, Yupeng Su, Hong Huang, Ngai Wong, and Hao Yu. APTQ: Attention- aware post-training mixed-precision quantization for large language models. InProceedings of the 61st IEEE/ACM Design Automation Conference, 2024

  35. [35]

    ResQ: Mixed-precision quan- tization of large language models with low-rank residuals.arXiv preprint arXiv:2412.14363, 2024

    Utkarsh Saxena, Sayeh Sharify, Kaushik Roy, and Xin Wang. ResQ: Mixed-precision quan- tization of large language models with low-rank residuals.arXiv preprint arXiv:2412.14363, 2024

  36. [36]

    Rethinking differentiable search for mixed-precision neural networks

    Zhaowei Cai and Nuno Vasconcelos. Rethinking differentiable search for mixed-precision neural networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2346–2355, 2020

  37. [37]

    Q-ViT: Fully differentiable quantization for vision transformer.arXiv preprint arXiv:2201.07703, 2022

    Zhexin Li, Tong Yang, Peisong Wang, and Jian Cheng. Q-ViT: Fully differentiable quantization for vision transformer.arXiv preprint arXiv:2201.07703, 2022

  38. [38]

    Jennings, and Arnon Netzer

    Hai Victor Habi, Roy H. Jennings, and Arnon Netzer. HMQ: Hardware friendly mixed precision quantization block for CNNs. InComputer Vision – ECCV 2020, pages 448–463. Springer, 2020

  39. [39]

    Categorical reparameterization with gumbel-softmax

    Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016

  40. [40]

    Maddison, Andriy Mnih, and Yee Whye Teh

    Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. The concrete distribution: A continuous relaxation of discrete random variables. InProceedings of the International Conference on Learning Representations, 2017

  41. [41]

    Bayesian bits: Unifying quantization and pruning

    Mart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad, Ying Wang, Tijmen Blankevoort, and Max Welling. Bayesian bits: Unifying quantization and pruning. InAdvances in Neural Information Processing Systems, volume 33, 2020

  42. [42]

    Mixed precision DNNs: All you need is a good parametrization

    Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso Garcia, Stephen Tiedemann, Thomas Kemp, and Akira Nakamura. Mixed precision DNNs: All you need is a good parametrization. InProceedings of the International Conference on Learning Representations, 2020

  43. [43]

    Micromix: Efficient mixed-precision quantization with microscaling formats for large language models

    Wenyuan Liu, Haoqian Meng, Yilun Luo, Yafei Zhao, Peng Zhang, and Xindian Ma. Micromix: Efficient mixed-precision quantization with microscaling formats for large language models. arXiv preprint arXiv:2508.02343, 2025

  44. [44]

    Mixture compressor for mixture-of-experts LLMs gains more

    Wei Huang, Yue Liao, Jianhui Liu, Ruifei He, Haoru Tan, Shiming Zhang, Hongsheng Li, Si Liu, and Xiaojuan Qi. Mixture compressor for mixture-of-experts LLMs gains more. InThe Thirteenth International Conference on Learning Representations, 2025

  45. [45]

    Ieee standard for floating-point arithmetic.IEEE Std 754-2019 (Revision of IEEE 754-2008), pages 1–84, 2019

    IEEE. Ieee standard for floating-point arithmetic.IEEE Std 754-2019 (Revision of IEEE 754-2008), pages 1–84, 2019

  46. [46]

    Jain, Albert Gural, Michael Wu, and Chris H

    Sambhav R. Jain, Albert Gural, Michael Wu, and Chris H. Dick. Trained quantization thresholds for accurate and efficient fixed-point inference of deep neural networks. InProceedings of the 3rd Machine Learning and Systems (MLSys) Conference, 2020

  47. [47]

    PyTorch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K¨opf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, high-perf...