Pith. sign in

Integer quantization for deep learning inference: Principles and empirical evaluation

10 Pith papers cite this work, alongside 218 external citations. Polarity classification is still indexing.

10 Pith papers citing it
218 external citations · external index

years

2026 8 2022 2

representative citing papers

Learning through Internalization

cs.LG · 2026-06-18 · unverdicted · novelty 7.0

A simplified one-layer transformer provably learns parities first with explicit CoT supervision then internalizes to direct computation as CoT tokens are removed.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

cs.LG · 2022-08-15 · conditional · novelty 7.0

LLM.int8() performs 8-bit inference for transformers up to 175B parameters with no accuracy loss by combining vector-wise quantization for most features with 16-bit mixed-precision handling of systematic outlier dimensions.

FP8 Formats for Deep Learning

cs.LG · 2022-09-12 · unverdicted · novelty 6.0

FP8 formats E4M3 and E5M2 match 16-bit training accuracy on CNNs, RNNs, and Transformers up to 175B parameters without hyperparameter changes.

citing papers explorer

Showing 10 of 10 citing papers.