ITERA-LLM shows that iteratively decomposing quantized LLM weight matrices into low-rank factors and allocating ranks by BLEU sensitivity yields better accuracy-latency trade-offs than quantization alone on FPGAs.
Democratizing neural ma- chine translation with OPUS-MT
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ITERA-LLM: Boosting Sub-8-Bit Large Language Model Inference via Iterative Tensor Decomposition
ITERA-LLM shows that iteratively decomposing quantized LLM weight matrices into low-rank factors and allocating ranks by BLEU sensitivity yields better accuracy-latency trade-offs than quantization alone on FPGAs.