LLMs encode accurate but brittle internal beliefs about latent game states and convert them poorly into actions, creating systematic gaps that explain strategic failures.
Is gpt-oss good? a comprehensive evaluation of openai’s latest open source models
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
An experiment finds that instructions and binary automated feedback can raise Singleton pattern adherence to 99-100% in models like Llama 3.3 and Qwen 3 while preserving or increasing functional test pass rates.
citing papers explorer
-
Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions
LLMs encode accurate but brittle internal beliefs about latent game states and convert them poorly into actions, creating systematic gaps that explain strategic failures.
-
Strategies for Guiding LLMs to Use Software Design Patterns: A Case of Singleton
An experiment finds that instructions and binary automated feedback can raise Singleton pattern adherence to 99-100% in models like Llama 3.3 and Qwen 3 while preserving or increasing functional test pass rates.