A new LoRA-MoE variant that shares the up-projection matrix across experts and adds dropout on it improves multi-task fine-tuning accuracy on six commonsense reasoning datasets.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MoSLD: An Extremely Parameter-Efficient Mixture-of-Shared LoRAs for Multi-Task Learning
A new LoRA-MoE variant that shares the up-projection matrix across experts and adds dropout on it improves multi-task fine-tuning accuracy on six commonsense reasoning datasets.