TWLA is a PTQ method using E2M-ATQ, KOTMS, and ILA-AMP to enable W1.58A4 quantization for LLMs with maintained accuracy.
Bitnet v2: Native 4-bit activations with hadamard transformation for 1-bit llms
5 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
UniSVQ unifies scalar and vector quantization for 2-bit LLM compression via affine transforms on integer lattices plus block-wise fine-tuning, outperforming SQ methods and matching advanced VQ with higher throughput.
LBLLM achieves better accuracy than prior binarization methods for LLMs by decoupling weight and activation quantization through initialization, layer-wise distillation, and learnable activation scaling.
RobuQ delivers the first stable DiT image generation at W1.58A2 average bits via Hadamard-based robust activation quantization and layer-wise mixed-precision activations.
CAT-Q performs post-training ternary quantization of 1.7B-235B LLMs with 512 samples via learnable modulation and softened ternarization, outperforming BitNet v1/v2 models trained on 100B tokens.
citing papers explorer
-
TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization
TWLA is a PTQ method using E2M-ATQ, KOTMS, and ILA-AMP to enable W1.58A4 quantization for LLMs with maintained accuracy.
-
UniSVQ: 2-bit Unified Scalar-Vector Quantization
UniSVQ unifies scalar and vector quantization for 2-bit LLM compression via affine transforms on integer lattices plus block-wise fine-tuning, outperforming SQ methods and matching advanced VQ with higher throughput.
-
LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation
LBLLM achieves better accuracy than prior binarization methods for LLMs by decoupling weight and activation quantization through initialization, layer-wise distillation, and learnable activation scaling.
-
RobuQ: Pushing DiTs to W1.58A2 via Robust Activation Quantization
RobuQ delivers the first stable DiT image generation at W1.58A2 average bits via Hadamard-based robust activation quantization and layer-wise mixed-precision activations.
-
CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
CAT-Q performs post-training ternary quantization of 1.7B-235B LLMs with 512 samples via learnable modulation and softened ternarization, outperforming BitNet v1/v2 models trained on 100B tokens.