Fine-tuning Qwen2.5-Math-1.5B with LoRA and GRPO on one RTX 3080 Ti improves GSM8K accuracy from 71.65 to 73.69, matching or beating several larger base models.
https://novasky- ai.github.io/posts/sky-t1
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can A Gamer Train A Mathematical Reasoning Model?
Fine-tuning Qwen2.5-Math-1.5B with LoRA and GRPO on one RTX 3080 Ti improves GSM8K accuracy from 71.65 to 73.69, matching or beating several larger base models.