An LLM can generate executable world-model components that, combined with MCTS, outperform LLM-as-policy in GOPS and match it in Taboo, though the evaluation under-supports the partial-observability claim.
Monte carlo sampling for regret minimization in extensive games
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2024 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making
An LLM can generate executable world-model components that, combined with MCTS, outperform LLM-as-policy in GOPS and match it in Taboo, though the evaluation under-supports the partial-observability claim.