Training a 7B LLM on 10K formatted demonstrations plus large-scale RL with a restart-and-explore strategy produces a single model that searches over reasoning steps and improves math and out-of-domain benchmarks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Training a 7B LLM on 10K formatted demonstrations plus large-scale RL with a restart-and-explore strategy produces a single model that searches over reasoning steps and improves math and out-of-domain benchmarks.