Long chain-of-thought reasoning can be distilled into Qwen2.5-32B-Instruct with 17k samples and LoRA, and performance degrades far more when reasoning steps are shuffled or deleted than when step content is corrupted.
Using logarithmic properties, we can rewrite log(n2) as 2 logn
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
Long chain-of-thought reasoning can be distilled into Qwen2.5-32B-Instruct with 17k samples and LoRA, and performance degrades far more when reasoning steps are shuffled or deleted than when step content is corrupted.