Initializing tensor adapters from the MPO decomposition of pretrained weights boosts fine-tuning accuracy and parameter efficiency over random and SVD-based initialization in LLaMA2-7B and LLaMA3-8B.
Advances in Neural Information Processing Systems 36 (2024)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
Initializing tensor adapters from the MPO decomposition of pretrained weights boosts fine-tuning accuracy and parameter efficiency over random and SVD-based initialization in LLaMA2-7B and LLaMA3-8B.