A self-alignment prompt pipeline (SHARP) generated roughly 190,000 verifiable STEM problems that, when used to distill or RL-train 7B models, improved GPQA Diamond accuracy in the reported runs.
It does not appear to present new theoretical results in the form of theorems or mathematical proofs that would require a separate section for assumptions and proofs
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning
A self-alignment prompt pipeline (SHARP) generated roughly 190,000 verifiable STEM problems that, when used to distill or RL-train 7B models, improved GPQA Diamond accuracy in the reported runs.