On a new 2D benchmark (WorldGen), LLM success falls from 100% on trivial worlds to 4% on medium worlds; ACE, an actor-critic-synthesizer loop, raises GPT-4's success to 88% on simple worlds but leaves medium worlds mostly unsolved.
] at the agent’s selected points are revealed, the agent can refine its strategy using this feedback
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Are Language Models Up to Sequential Optimization Problems? From Evaluation to a Hegelian-Inspired Enhancement
On a new 2D benchmark (WorldGen), LLM success falls from 100% on trivial worlds to 4% on medium worlds; ACE, an actor-critic-synthesizer loop, raises GPT-4's success to 88% on simple worlds but leaves medium worlds mostly unsolved.