EvoPolicyGym is a new benchmark suite of 16 compact RL environments that evaluates autonomous policy evolution, with GPT-5.5 achieving the top aggregate rank and top-two performance on all tasks.
Nature , volume=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.AI 2years
2026 2representative citing papers
citing papers explorer
-
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
EvoPolicyGym is a new benchmark suite of 16 compact RL environments that evaluates autonomous policy evolution, with GPT-5.5 achieving the top aggregate rank and top-two performance on all tasks.
- When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning