Fine-tuning Llama-3-8B on GPT-4-rewritten 'silly' versions of MMLU questions gives at most a 0.54% overall gain and no improvement over seed-only fine-tuning.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning from "Silly" Questions Improves Large Language Models, But Only Slightly
Fine-tuning Llama-3-8B on GPT-4-rewritten 'silly' versions of MMLU questions gives at most a 0.54% overall gain and no improvement over seed-only fine-tuning.