Training a 7B LLM on 10K formatted demonstrations plus large-scale RL with a restart-and-explore strategy produces a single model that searches over reasoning steps and improves math and out-of-domain benchmarks.
( s + 2) ( 2.4− t 60 ) = 9 Let’s solve these equations step by step
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search
Training a 7B LLM on 10K formatted demonstrations plus large-scale RL with a restart-and-explore strategy produces a single model that searches over reasoning steps and improves math and out-of-domain benchmarks.