Huff-LLM splits FP16/BF16 LLM weights into small bit groups, Huffman-compresses each group, and uses custom hardware decoders so weights stay compressed through the memory hierarchy during inference.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Huff-LLM splits FP16/BF16 LLM weights into small bit groups, Huffman-compresses each group, and uses custom hardware decoders so weights stay compressed through the memory hierarchy during inference.