R1-style reinforcement learning on single-step web actions lifts open-source agents above gpt-4o on WorkArena while avoiding the reward hacking seen with dense rewards.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
WorkForceAgent-R1: Incentivizing Reasoning Capability in LLM-based Web Agents via Reinforcement Learning
R1-style reinforcement learning on single-step web actions lifts open-source agents above gpt-4o on WorkArena while avoiding the reward hacking seen with dense rewards.