Quantizing LLMs from 32 to 8 bits cuts GPU memory by about 76%, shifts affective-generation F1 by up to 10 percentage points in either direction, and lets larger quantized models rival smaller full-precision models in text quality.
Title resolution pending
1 Pith paper cite this work, alongside 109 external citations. Polarity classification is still indexing.
1
Pith paper citing it
109
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLM-based Affective Text Generation Quality Based on Different Quantization Values
Quantizing LLMs from 32 to 8 bits cuts GPU memory by about 76%, shifts affective-generation F1 by up to 10 percentage points in either direction, and lets larger quantized models rival smaller full-precision models in text quality.