L0 combines a code-as-action agent scaffold with multi-turn RLVR, lifting Qwen2.5-7B HotpotQA EM from 22 to 41 and SimpleQA judge accuracy from 30 to 80.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
L0: Reinforcement Learning to Become General Agents
L0 combines a code-as-action agent scaffold with multi-turn RLVR, lifting Qwen2.5-7B HotpotQA EM from 22 to 41 and SimpleQA judge accuracy from 30 to 80.