Pith. sign in

REVIEW 4 major objections 6 minor 2 cited by

AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read AMXFP4, an asymmetric microscaling 4-bit floating-point format that gives positive and negative elements their own shared scales, enables calibration-free direct-cast 4-bit LLM inference that matches or beats rotation-based INT4 pipelines…

desk verdict Solid format paper with a real insight; abstract overclaims and the emulator/hardware gap needs a check. read the letter →

arxiv 2411.09909 v2 pith:VL4S7YSV submitted 2024-11-15 cs.AI

classification cs.AI
keywords 4-bitquantizationmicroscalingasymmetricfloating-pointLLMinferenceactivationoutlierscalibration-freerotation-basedMXFP4
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AMXFP4 is a 4-bit floating-point format for LLM inference that splits each quantization group into two halves, each with its own shared scale: one for positive values and one for negative values. The paper's central argument is that standard microscaling formats such as MXFP4 suppress activation outliers by grouping elements into blocks of 32, but that this grouping makes each block's value distribution asymmetric, so a symmetric zero-centered representation is no longer optimal. By using asymmetric shared FP8 scales, AMXFP4 keeps the outlier suppression of microscaling while quantizing closer to the ideal Lloyd-Max limit, and it does this through direct casting without calibration. Reported results show AMXFP4 beating MXFP4 across language modeling, chatbot, long-context, and visual question-answering tasks, and matching or exceeding rotation-based INT4 pipelines such as QuaRot with GPTQ and SpinQuant while also quantizing attention matmuls. If correct, this would remove the multi-hour calibration step from accurate 4-bit LLM inference and make fully quantized attention practical at long context.

What carries the argument

The load-bearing mechanism is the pair of asymmetric shared scales. For each group of 32 elements, AMXFP4 stores a positive shared scale $2^{e_{sp}}\hat{M}_p$ and a negative shared scale $2^{e_{sn}}\hat{M}_n$, both FP8 with 5 exponent bits and 2 mantissa bits, and encodes each element as an FP4 value with sign, 2-bit exponent, and 1-bit mantissa. During multiplication the product of two elements takes one of four scale combinations, $S_{Xp}S_{Wp}$, $S_{Xp}S_{Wn}$, $S_{Xn}S_{Wp}$, or $S_{Xn}S_{Wn}$, selected by a small lookup on the two operand signs; because the scale mantissa is only 2 bits and the same scale serves the whole group, this sign-aware scaling is nearly free in hardware. The design also replaces MX's floor-based power-of-two scale decision with rounding, avoiding the clamping error that otherwise grows as group size shrinks. Together these choices give the format an asymmetric grid of representable values that matches the lopsided group distributions that microscaling itself creates.

What would settle it

Feed the same random and real activation tensors through the software emulation from Appendix B and a detailed simulation of the synthesized AMXFP4 MAC unit, and compare outputs bit by bit; any divergence in scale-pair selection, clamping, or rounding would show that the reported accuracy gains may not hold on the hardware whose cost is claimed.

Watch

Extended reading notes

Core claim

The central discovery is a trade-off hidden in microscaling: shrinking the quantization group to 32 elements tames activation outliers (kurtosis falls nearly to zero), but it scatters the group means, meaning each group is more asymmetric the finer the group granularity gets. MXFP4's symmetric representation therefore leaves error on the table, and data rotation, which helps at row-level group sizes, actively hurts when combined with microscaling because it adds still more asymmetry. AMXFP4 counters this with two shared FP8 (E5M2) scales per group, one for positive and one for negative elements, with the scale chosen at multiply time from the signs of the two operands. On Wikitext-2 this lowers perplexity from 6.49 to 6.22 for LLaMA2-7B relative to MXFP4, lifts ChartQA from 46.20 to 49.48, and in Table 10 reaches a WinoGrande accuracy of 67.32 against a best rotation-baseline value of 66.22, without any calibration. The format also narrows the gap to the 16-bit baseline enough that MT-Bench conversational scores recover close to baseline, and a synthesized MAC unit implementing the sign-dependent scale selection is reported at only about 10% area overhead over a compatible MX MAC.

Load-bearing premise

The load-bearing premise is that the software emulator used for all accuracy measurements behaves exactly like the separately synthesized hardware MAC unit, including the sign-dependent choice of shared scale and the treatment of clamping and rounding; the paper reports the two implementations separately and does not show bit-level agreement between them.

Editorial extensions

If this is right

  • All attention matrix multiplications, including the softmax-output and query-key products that rotation methods leave in FP16, can be run in 4-bit, which matters most as context length grows because attention FLOPs scale quadratically.
  • Models can be deployed at 4-bit by direct casting with no calibration run, removing the multi-hour overhead and the calibration-set overfitting that rotation-based pipelines exhibit.
  • A 3-bit variant, AMXFP3, degrades Wikitext-2 perplexity by only about 1.7 on LLaMA2-7B, whereas QuaRot with GPTQ degrades by more than 30, suggesting the approach extends below 4 bits.
  • AMXFP4 is compatible with other compression methods: applied to a 20%-pruned LLaMA-7B model it recovers most of the pruning accuracy drop, so its benefits are additive.
  • An asymmetric version of the recently deployed commercial MXFP4 variant NVFP4, called ANVFP4, also beats the commercial format, particularly at group size 16, indicating that per-sign scales are a generally useful addition to MX-family formats.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own group-size sweeps show AMXFP4's advantage over symmetric MXFP4 widening as groups shrink, so a natural extension is to push the asymmetric shared-scale idea to even smaller groups (for example, 16 or 8 elements) or to per-channel granularity.
  • Because the mechanism targets distribution shape rather than LLM-specific structure, the same asymmetric shared-scale design could plausibly transfer to other outlier-heavy workloads such as vision transformers, diffusion models, or training-side rescaling, though the paper does not test these.
  • The emulator-versus-hardware gap could be closed by a bit-exact co-verification harness; without that artifact, the accuracy story and the 10% overhead story remain two claims about two different implementations.
  • A testable prediction of the paper's analysis is that combining row-level rotation with fine-grained asymmetric microscaling would recover the row-level rotation benefit without the destructive interaction seen at group size 32, something the paper's Table 5 suggests but does not directly explore.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes AMXFP4, a 4-bit asymmetric microscaling floating-point format for LLM inference. The format keeps FP4 (E2M1) elements but uses separate FP8 E5M2 shared scales for positive and negative values, with a sign-dependent scale-selection rule during multiplication. The authors argue that microscaling suppresses activation outliers but increases group-wise asymmetry, and that the asymmetric shared scale addresses this without calibration. They evaluate AMXFP4 on Wikitext-2 perplexity, MT-Bench, visual question answering, LongBench-E, MMLU, CSQA, and attention-only settings, reporting consistent improvements over MXFP4 and competitive or better accuracy than rotation-based methods such as QuaRot and SpinQuant. They also implement a custom AMXFP4 MAC unit and claim roughly 10% hardware overhead over an MX-compatible MAC. The paper includes ablations on shared-scale format, group size, rotation interaction, QAT, 3-bit extension, pruning, and a 70B model.

Significance. If the claims hold, AMXFP4 is a practically useful format: it enables calibration-free direct-cast 4-bit inference with fully quantized attention, which is a gap relative to rotation-based methods that require calibration and leave softmax outputs in FP16. The paper is creditable for releasing code, evaluating across a broad set of tasks and model families (including a 70B model and an encoder-decoder model), and for providing extensive ablations, including the interaction of rotation with microscaling and the choice of shared-scale encoding. The central accuracy comparisons (Tables 1, 3, 4, 10, 12) support the qualitative claim that asymmetric shared scales improve over symmetric MXFP4 and often match or exceed rotation-based methods. However, the hardware-cost claim is currently not supported by the reported numbers, and the accuracy results come from a software emulator that is not shown to be bit-exact with the synthesized hardware. These gaps are load-bearing for the paper's central accuracy-plus-cost claim.

major comments (4)
  1. [§5.5, Table 9] The hardware-cost claim is not supported by the displayed numbers: Table 9 reports AMXFP4 at 8.32x versus MXFP4 at 9.23x on Area-Memory, and AMXFP4 is also lower on Power-Area and Power-Area-Memory. This implies AMXFP4 is roughly 10% cheaper, not 10% more expensive, as the text states in Section 4.3 and Section 5.5 ("adds only 10% overhead"). Please correct the rows, units, or the baseline for the overhead statement, and re-run the synthesis analysis if necessary, because the claimed cost is a central part of the contribution.
  2. [§5.5 vs. Appendix B.3] The accuracy results are produced by a software emulation path (quantize_mx_op with fp4_e2m1_asym and scale_mode=152 in Appendix B.3), while the hardware cost is measured on a separately synthesized AMXFP4 MAC unit in Section 5.5. The paper provides no bit-exact comparison between these two implementations for the same tensors, including the sign-dependent selection between the positive and negative FP8 E5M2 shared scales, the exponent-rounding rule, and the clamping behavior. Please add such a check, or characterize the numerical divergence if the implementations are not bit-identical, so that the reported accuracy improvements can be attributed to the hardware design whose cost is claimed.
  3. [Abstract and Tables 3, 4, 10] The abstract's quantitative claims exceed what the tables show. "Outperforms MXFP4 by 3% on VQA" does not match Table 3, where the gains are 1.25 points on VQA-T, 2.72 on DocVQA, 0.50 on OCRBench, and 3.28 on ChartQA; there is no single 3% improvement across the VQA benchmarks. Similarly, "exceeds rotation-based methods by 1.6% on CSQA" is not directly shown: Table 10 reports ARC-Challenge and WinoGrande, and Table 4 compares against NVFP4 rather than rotation-based methods. Please revise the abstract to cite the exact table and metric, or add the missing CSQA comparison.
  4. [§4.2, Fig. 5, Table 18] The choice of shared-scale format (E5M2 vs. E4M3 vs. PoT with floor or round) is made by minimizing Wikitext-2 perplexity on LLaMA2-7B (Fig. 5, Table 18), and Wikitext-2 perplexity is also a headline evaluation metric throughout the paper (Tables 1, 12, 14). The paper should either report whether the selected configuration remains optimal on a held-out task or model that was not used for selection, or explicitly discuss the potential selection bias. As written, part of the reported advantage is tuned to the evaluation objective, which weakens the claim that E5M2 is intrinsically the best shared-scale choice.
minor comments (6)
  1. [§4.3] The text refers to "Appendix 5.5" when describing the hardware evaluation; this should be Section 5.5.
  2. [Appendix B.1, Algorithm 1] Algorithm 1 still describes the original floor-based MX quantization, but Section 4.2 introduces a modified rounding rule for the PoT scale; please present the updated algorithm or state explicitly that the emulator implements the rounding variant.
  3. [Fig. 1] The caption lists two subfigures labeled "(d)" and the subfigure letters do not match the order in which they are discussed in the text; please renumber the subfigures.
  4. [Table 9] The column headings "Area-Memory", "Power-Area", and "Power-Area-Memory" need explicit definitions (e.g., whether these are products of normalized ratios) so the reader can interpret the reported multipliers.
  5. [§5.2] The term "MXFP4-PoT" is used before it is defined in Section 4.2; please define it at first use.
  6. [§5.1, Table 2] The paper does not report the number of random seeds or trials for the rotation-based comparisons; adding this information would improve the reliability of the overfitting analysis.

Circularity Check

1 steps flagged · score 3.0 of 10

Format choices are selected on the same LLaMA2-7B/Wikitext-2 perplexity that is later reported as evidence, but the central accuracy claim is independently supported on held-out tasks.

  1. other [Sec. 4.1-4.2 (Fig. 5(a), Table 18); reported in Tables 1 and 12]
    "To evaluate the benefits of asymmetric formats, we compare the mean-square error (MSE) on activation samples from LLaMA2-7B’s QKV-Proj at layer 5 ... This finding supports the selection of AsymFP4 as the element-wise format, further validated empirically in Table 1. ... Therefore, we select FP8 with a 5-bit exponent (E5M2) as the shared scale, as these scales largely mitigate accuracy degradation caused by the limited resolution and narrower dynamic range (see Table 18 for ablation studies)."

    The element format (AsymFP4) and shared scale (E5M2) are chosen by minimizing quantization MSE on LLaMA2-7B activations and by minimizing LLaMA2-7B Wikitext-2 perplexity across candidate formats (Fig. 5(a), Table 18). Tables 1 and 12 then present the lower Wikitext-2 perplexity of AMXFP4 on LLaMA2-7B as empirical validation. On this specific row the comparison is the selection objective: the chosen format's perplexity is lower than the rejected alternatives by construction of the argmin, so it is not an independent confirmation. The circularity is limited because the paper's broader claims are reproduced on held-out models/tasks (VQA, CSQA, MT-Bench, LongBench, LLaMA3-70B) that were not used for format selection.

full rationale

Aside from the in-sample format-selection issue above, the derivation chain is not circular. The proposed AMXFP4 encoding (Eq. 1) is a concrete extension of the cited AsymFP/AFPQ idea to group-wise FP8 shared scales; it is not a renaming of a known result, and it does not import any uniqueness theorem from the authors' prior work. The accuracy comparisons against MXFP4, QuaRot, SpinQuant, and NVFP4 are empirical and mostly on tasks/models outside the selection data. The hardware claim in Sec. 5.5 is a validation gap rather than a circularity: Table 9's numbers appear inconsistent with the stated ~10% overhead, and no bit-exact check links the software emulator (Appendix B.3) to the synthesized MAC, but these are correctness/reporting concerns, not a reduction of the result to its inputs. The limitation paragraph explicitly acknowledges that only MAC-level hardware evaluation was performed.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The format has no continuous fitted parameters, but three discrete design choices (element format, shared-scale bit split, PoT rounding rule) were tuned on the same benchmark types used for the headline results. Two domain assumptions about emulation fidelity and hardware-level generality are load-bearing for the accuracy and cost claims.

free parameters (3)
  • Shared-scale FP8 exponent/mantissa split = E5M2 (1-5-2)
    Selected by Wikitext-2 perplexity on LLaMA2-7B (Fig. 5, Table 18); E4M3 and E5M2 variants were compared and E5M2 chosen, which sets the dynamic range and rounding granularity of the format.
  • Power-of-two shared-scale rounding rule = Round instead of floor
    The PoT shared scale in the MX spec uses floor on the exponent; the paper proposes rounding to reduce clamping error, a design choice validated by perplexity and error decomposition (Appendix B.2, Fig. 9).
  • Element-wise format = AsymFP4 (E2M1 with sign-dependent scale)
    Selected by MSE comparison against Lloyd-Max on LLaMA2-7B layer 5 activations (Sec. 4.1, Fig. 4); other formats (INT4, FP4, NF4, SF4) were tested and AsymFP4 had lowest MSE.
assumptions (4)
  • domain assumption Group-wise kurtosis and mean are sufficient to characterize quantization difficulty under microscaling
    Sec. 3 uses these two statistics to claim that microscaling suppresses outliers but introduces asymmetry; no theorem or additional statistics are used to justify the conclusion.
  • domain assumption Reducing element-wise MSE on sampled activations improves end-task LLM accuracy
    Sec. 4.1 selects AsymFP4 based on MSE on LLaMA2-7B activations; the paper does not prove that MSE ordering transfers to perplexity/accuracy for all models.
  • domain assumption The MX emulation library faithfully models the numerical behavior of the proposed AMXFP4 hardware
    All accuracy numbers come from software emulation (Appendix B) while the hardware claim comes from a synthesized MAC (Sec. 5.5); no bit-exact comparison is reported.
  • domain assumption MAC-level hardware cost is a valid proxy for inference-system-level cost
    The 'negligible hardware cost' claim is based on a single MAC unit synthesis on 4nm (Sec. 5.5); the paper's Limitations section acknowledges that full system-level throughput and energy were not evaluated.
invented entities (1)
  • AMXFP4 format: asymmetric microscaling FP4 with separate FP8 E5M2 shared scales for positive and negative values
    purpose: Reduces activation quantization error in 4-bit LLM inference to approach Lloyd-Max performance without calibration
    Support for the format comes from the authors' own emulation experiments and MAC synthesis; no external replication, silicon, or formal proof is reported. The released code could provide independent verification, but that has not been demonstrated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference." pith.science (2026). https://pith.science/paper/VL4S7YSV

@misc{pith2026241109909,
  author       = {Pith},
  title        = {Pith review of: AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VL4S7YSV}},
  note         = {Machine review of arXiv:2411.09909}
}
read the original abstract

As large language models (LLMs) grow in parameter size and context length, computation precision has been reduced from 16-bit to 4-bit to improve inference efficiency. However, this reduction causes accuracy degradation due to activation outliers. Rotation-based INT4 methods address this via matrix calibration, but they introduce multi-hour overheads and leave key computations in full precision. Microscaling (MX) floating-point (FP) formats offer fine-grained representation with a shared scale, enabling fully quantized matrix multiplications through direct casting without calibration. However, existing research shows unsatisfactory empirical results for MXFP4 inference, and the robustness of MX formats remains largely unexplored. In this work, we uncover the fundamental tradeoffs of the MX format: while it effectively suppresses activation outliers, it does so at the cost of increased group-wise asymmetry. To address this, we propose AMXFP4, a 4-bit asymmetric FP format that handles both issues using asymmetric shared scales, without requiring calibration. Our custom MAC engine adds negligible hardware cost while improving accuracy: AMXFP4 outperforms MXFP4 by 3% on VQA and exceeds rotation-based methods by 1.6% on CSQA. It also surpasses recently deployed commercial MXFP4 variants. Code: https://github.com/aiha-lab/MX-QLLM

Figures

Figures reproduced from arXiv: 2411.09909 by the authors.

Figure 1
Figure 1. (a) FLOPS across context length and model sizes. (b) Precision scaling in NVIDIA Tensor Cores. (d) [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Data characteristics based on (a-d) types of [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Cluster-wise Lloyd-Max quantization and quantization error across data formats (LLaMA2-7B layer 5 QKV-Proj input activation). Detailed cluster￾wise error statistics and results from other layers are provided in [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figures from the paper (9 more)
Figure 5
Figure 5. Figure 5: (a) Illustration of AMXFP4 and LLaMA2- 7B Wikitext-2 perplexity across shared scale types. (b) Multiplication between two AMXFP4 datas. define AsymFP such that an exponent-bit-shifted mantissa represents a value, which is then scaled by a shared factor with sign-depend…
Figure 6
Figure 6. Figure 6: Normalized single score of MT-Bench (LLaMA2-Chat-7B). Absolute accuracies are in Ta￾ble 16 in Appendix. mains unaffected by the calibration set and no￾tably improves results and surpasses conventional calibration-based methods. 5.2 Enhancing MX Performance In this sect…
Figure 7
Figure 7. Figure 7: LongBench-E results on LLaMA2-Chat-7B. 2022), highlighting the significant advantages of asymmetric data representation in VLMs (example is shown in [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: (a) Illustration of where reduced-precision ma [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Impact of shared scale (LLaMA2-7B). More [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: MX dot-product architecture. entire quantization procedure, MX considers the maximum data value to determine the shared scale, performing a floor operation after extracting the exponent of the element’s maximum value with log2. B.2 Determining PoT Shared Scale: Floor …
Figure 11
Figure 11. Figure 11: Example of chatbot interactions from MT [PITH_FULL_IMAGE:figures/full_fig_p017_11.png]
Figure 12
Figure 12. Figure 12: Comparison between responses from MXFP4-PoT and AMXFP4 in ChartQA example. texts exceeding 8K. While MXFP4 substantially improves over MXFP4-PoT, it still experiences a score reduction of over 6 when handling contexts above 8K. AMXFP4 increases the average score by mo…
Figure 13
Figure 13. Figure 13: MT-Bench example (LLaMA2-Chat-7B) [PITH_FULL_IMAGE:figures/full_fig_p019_13.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference

    cs.AR 2026-07 conditional novelty 6.0 of 10

    LightRot uses grouped local rotation plus outlier alignment to make 4-bit LLaMA inference accurate and cheap, claiming 27.4 TOPS/W on a 28nm accelerator.

  2. GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference

    cs.AR 2026-07 conditional novelty 6.0 of 10

    Rotation and fine-grained group quantization can work together if rotation spans several quantization groups and outlier channels are permuted onto harmonic Hadamard rows, enabling 4-bit LLM inference with integer-onl...

Reference graph

Works this paper leans on

79 extracted references · 23 canonical work pages · cited by 2 Pith papers

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card

  4. [4]

    AMD. 2024. Amd instinct™ mi325x accelerators. https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/product-briefs/instinct-mi325x-datasheet.pdf

  5. [5]

    Michael Andersch, Greg Palmer, Ronny Krashinsky, Nick Stam, Vishal Mehta, Gonzalo Brito, and Sridhar Ramaswamy. 2022. Nvidia hopper architecture in-depth. https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/

  6. [6]

    Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. 2024. Quarot: Outlier-free 4-bit inference in rotated llms. arXiv preprint arXiv:2404.00456

  7. [7]

    AzureAI. 2024. Azure maia for the era of ai: From silicon to software to systems. https://azure.microsoft.com/en-us/blog/azure-maia-for-the-era-of-ai-from-silicon-to-software-to-systems/

  8. [8]

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609

Show all 79 references
  1. [9]

    Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. https://doi.org/10.18653/v1/2024.acl-long.172 L ong B ench: A bilingual, multitask benchmark for long context u...

  2. [10]

    Jihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang, and Simyung Chang. 2024. https://doi.org/10.18653/v1/2024.acl-long.204 Crayon: Customized on-device LLM via instant adapter blending and edge-server hybrid inference . In Proceedings of the 62nd Annual Meeting of the Associ...

  3. [11]

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. https://arxiv.org/abs/1911.11641 Piqa: Reasoning about physical commonsense in natural language . Preprint, arXiv:1911.11641

  4. [12]

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence

  5. [13]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\

  6. [14]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311

  7. [15]

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416

  8. [16]

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1

  9. [17]

    Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e . 2022. Flash A ttention: Fast and memory-efficient exact attention with IO -awareness. In Advances in Neural Information Processing Systems (NeurIPS)

  10. [18]

    Bita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu, Jeremy Fowers, Kalin Ovtcharov, Anna Vinogradsky, Sarah Massengill, Lita Yang, Ray Bittner, et al. 2020. Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point. Advances in neural...

  11. [19]

    Bita Darvish Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, et al. 2023. With shared microexponents, a little shifting goes a long way. In Proceedings of the 50th Annua...

  12. [20]

    Smith, and Matt Gardner

    Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers

  13. [21]

    Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. Llm.int8(): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339

  14. [22]

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. https://openreview.net/forum?id=OUIFPHEgJU QL o RA : Efficient finetuning of quantized LLM s . In Thirty-seventh Conference on Neural Information Processing Systems

  15. [23]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. https://openreview.net/forum?id=YicbFdNTTy An image is worth 1...

  16. [24]

    Abdelfattah, and Zhiru Zhang

    Jordan Dotzel, Yuzong Chen, Bahaa Kotb, Sushma Prasad, Gang Wu, Sheng Li, Mohamed S. Abdelfattah, and Zhiru Zhang. 2024. Learning from students: Applying t-distributions to explore accurate and efficient formats for llms. International Conference on Machine Learning

  17. [25]

    Mario Drumond, Tao Lin, Martin Jaggi, and Babak Falsafi. 2018. Training dnns with hybrid block floating point. Advances in Neural Information Processing Systems, 31

  18. [26]

    Maxim Fishman, Brian Chmiel, Ron Banner, and Daniel Soudry. 2024. https://arxiv.org/abs/2409.12517 Scaling fp8 training to trillion-token llms . Preprint, arXiv:2409.12517

  19. [27]

    Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323

  20. [28]

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027

  21. [29]

    Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. https://doi.org/10.5281/zenodo.53...

  22. [30]

    Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. https://doi.org/10.18653/v1/D19-5409 SAMS um corpus: A human-annotated dialogue dataset for abstractive summarization . In Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 70--79, Ho...

  23. [31]

    Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley. 2023. https://arxiv.org/abs/2306.14893 Longcoder: A long-range pre-trained language model for code completion . Preprint, arXiv:2306.14893

  24. [32]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. https://arxiv.org/abs/2009.03300 Measuring massive multitask language understanding . CoRR, abs/2009.03300

  25. [33]

    Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. https://doi.org/10.18653/v1/2020.coling-main.580 Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps . In Proceedings of the 28th International Conference on Computational Li...

  26. [34]

    Mark Horowitz. 2014. Energy table for 45nm process

  27. [35]

    Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. 2021. https://doi.org/10.18653/v1/2021.naacl-main.112 Efficient attentions for long document summarization . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computati...

  28. [36]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825

  29. [37]

    Mandar Joshi , Eunsol Choi , Daniel Weld , and Luke Zettlemoyer . 2017. https://arxiv.org/abs/1705.03551 triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension . arXiv e-prints, arXiv:1705.03551

  30. [38]

    Bryan Klimt and Yiming Yang. 2004. https://api.semanticscholar.org/CorpusID:13451873 The enron corpus: A new dataset for email classi(cid:12)cation research

  31. [39]

    Janghwan Lee, Minsoo Kim, Seungcheol Baek, Seok Hwang, Wonyong Sung, and Jungwook Choi. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.910 Enhancing computation efficiency in large language models through weight and activation quantization . In Proceedings of the 2023 Confe...

  32. [40]

    Janghwan Lee, Seongmin Park, Sukjin Hong, Minsoo Kim, Du-Seong Chang, and Jungwook Choi. 2024. https://doi.org/10.18653/v1/2024.acl-long.612 Improving conversational abilities of quantized large language models via direct preference alignment . In Proceedings of the 62nd Annua...

  33. [41]

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461

  34. [42]

    Xin Li and Dan Roth. 2002. https://aclanthology.org/C02-1150 Learning question classifiers . In COLING 2002: The 19th International Conference on Computational Linguistics

  35. [43]

    Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics

  36. [44]

    Haokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui, Yingtao Zhang, Linzhan Mou, Linqi Song, Zhenan Sun, and Ying Wei. 2024. https://arxiv.org/abs/2406.01721 Duquant: Distributing outliers via dual transformation makes stronger quantized llms . Preprint, arXiv:2406.01721

  37. [45]

    Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv

  38. [46]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 a . https://proceedings.neurips.cc/paper_files/paper/2023/file/6dcf277ea32ce3288914faf369fe6de0-Paper-Conference.pdf Visual instruction tuning . In Advances in Neural Information Processing Systems, volume 36, pages...

  39. [47]

    Tianyang Liu, Canwen Xu, and Julian McAuley. 2023 b . https://arxiv.org/abs/2306.03091 Repobench: Benchmarking repository-level code auto-completion systems . Preprint, arXiv:2306.03091

  40. [48]

    Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xucheng Yin, Cheng lin Liu, Lianwen Jin, and Xiang Bai. 2024 a . https://arxiv.org/abs/2305.07895 Ocrbench: On the hidden mystery of ocr in large multimodal models . Preprint, arXiv:2305.07895

  41. [49]

    Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort. 2024 b . Spinquant--llm quantization with learned rotations. arXiv preprint arXiv:2405.16406

  42. [50]

    S. Lloyd. 1982. https://doi.org/10.1109/TIT.1982.1056489 Least squares quantization in pcm . IEEE Transactions on Information Theory, 28(2):129--137

  43. [51]

    Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. Llm-pruner: On the structural pruning of large language models. In Advances in Neural Information Processing Systems

  44. [52]

    Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz

    Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993. https://aclanthology.org/J93-2004 Building a large annotated corpus of E nglish: The P enn T reebank . Computational Linguistics, 19(2):313--330

  45. [53]

    Ahmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. https://doi.org/10.18653/v1/2022.findings-acl.177 C hart QA : A benchmark for question answering about charts with visual and logical reasoning . In Findings of the Association for Computational Linguisti...

  46. [54]

    Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021. https://arxiv.org/abs/2007.00398 Docvqa: A dataset for vqa on document images . Preprint, arXiv:2007.00398

  47. [55]

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. https://arxiv.org/abs/1609.07843 Pointer sentinel mixture models . Preprint, arXiv:1609.07843

  48. [56]

    Nvidia. 2017. https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf Nvidia tesla v100 gpu architecture

  49. [57]

    Nvidia. 2020. Nvidia a100 tensor core gpu architecture. https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/nvidia-ampere-architecture-whitepaper.pdf

  50. [58]

    Nvidia. 2024. https://resources.nvidia.com/en-us-blackwell-architecture Nvidia blackwell architecture technical brief

  51. [59]

    NVIDIA. 2024. Tensorrt-llm. https://github.com/NVIDIA/TensorRT-LLM

  52. [60]

    National Library of Medicine

    Courtesy of the U.S. National Library of Medicine. 2023. Pubmed. https://huggingface.co/datasets/ncbi/pubmed

  53. [61]

    OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  54. [62]

    Bita Darvish Rouhani, Nitin Garegrat, Tom Savell, Ankit More, Kyung-Nam Han, Ritchie Zhao, Mathew Hall, Jasmine Klar, Eric Chung, Yuan Yu, Michael Schulte, Ralph Wittig, Ian Bratt, Nigel Stephens, Jelena Milanovic, John Brothers, Pradeep Dubey, Marius Cornea, Alexander Heineck...

  55. [63]

    Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, et al. 2023 b . Microscaling data formats for deep learning. arXiv preprint arXiv:2310.10537

  56. [64]

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. https://arxiv.org/abs/1907.10641 Winogrande: An adversarial winograd schema challenge at scale . Preprint, arXiv:1907.10641

  57. [65]

    Liu, and Christopher D

    Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. https://doi.org/10.18653/v1/P17-1099 Get to the point: Summarization with pointer-generator networks . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers...

  58. [66]

    Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. 2024. https://openreview.net/forum?id=8Wuvhh0LYW Omniquant: Omnidirectionally calibrated quantization for large language models . In The Twelfth Internat...

  59. [67]

    Amanpreet Singh, Vivek Natarjan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards vqa models that can read. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8317--8326

  60. [68]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://arxiv.org/abs/1811.00937 Commonsenseqa: A question answering challenge targeting commonsense knowledge . Preprint, arXiv:1811.00937

  61. [69]

    Hugo Touvron et al. 2023. https://arxiv.org/abs/2307.09288 Llama 2: Open foundation and fine-tuned chat models . Preprint, arXiv:2307.09288

  62. [70]

    Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, and Song Han. 2022. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438

  63. [71]

    Jaewoo Yang, Hayun Kim, and Younghoon Kim. 2024. https://arxiv.org/abs/2405.14428 Mitigating quantization errors due to activation spikes in glu-based llms . Preprint, arXiv:2405.14428

  64. [72]

    Cohen, Ruslan Salakhutdinov, and Christopher D

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://arxiv.org/abs/1809.09600 Hotpotqa: A dataset for diverse, explainable multi-hop question answering . Preprint, arXiv:1809.09600

  65. [73]

    Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, and Jianfei Chen. 2025 a . Sageattention2: Efficient attention with thorough outlier smoothing and per-thread int4 quantization. In International Conference on Machine Learning (ICML)

  66. [74]

    Jintao Zhang, Jia Wei, Pengle Zhang, Xiaoming Xu, Haofeng Huang, Haoxu Wang, Kai Jiang, Jun Zhu, and Jianfei Chen. 2025 b . Sageattention3: Microscaling fp4 attention for inference and an exploration of 8-bit training. arXiv preprint arXiv:2505.11594

  67. [75]

    Jintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu, and Jianfei Chen. 2025 c . Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration. In International Conference on Learning Representations (ICLR)

  68. [76]

    Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, and Ziwei Liu. 2024 a . https://arxiv.org/abs/2407.12772 Lmms-eval: Reality check on the evaluation of large multimodal models . Preprint, arX...

  69. [77]

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...

  70. [78]

    Yijia Zhang, Sicheng Zhang, Shijie Cao, DaYou Du, Jianyu Wei, Ting Cao, and Ningyi Xu. 2024 b . https://doi.org/10.18653/v1/2024.findings-acl.3 AFPQ : Asymmetric floating point quantization for LLM s . In Findings of the Association for Computational Linguistics ACL 2024, page...

  71. [79]

    Gonzalez, and Ion Stoica

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. https://openreview.net/forum?id=uccHPGDlao Judging LLM -as-a-judge with MT -bench and chatbot ...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.