REVIEW 2 cited by
Modeling the Mistakes of Boundedly Rational Agents Within a Bayesian Theory of Mind
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
When inferring the goals that others are trying to achieve, people intuitively understand that others might make mistakes along the way. This is crucial for activities such as teaching, offering assistance, and deciding between blame or forgiveness. However, Bayesian models of theory of mind have generally not accounted for these mistakes, instead modeling agents as mostly optimal in achieving their goals. As a result, they are unable to explain phenomena like locking oneself out of one's house, or losing a game of chess. Here, we extend the Bayesian Theory of Mind framework to model boundedly rational agents who may have mistaken goals, plans, and actions. We formalize this by modeling agents as probabilistic programs, where goals may be confused with semantically similar states, plans may be misguided due to resource-bounded planning, and actions may be unintended due to execution errors. We present experiments eliciting human goal inferences in two domains: (i) a gridworld puzzle with gems locked behind doors, and (ii) a block-stacking domain. Our model better explains human inferences than alternatives, while generalizing across domains. These findings indicate the importance of modeling others as bounded agents, in order to account for the full richness of human intuitive psychology.
Forward citations
Cited by 2 Pith papers
-
Belief Attribution as Mental Explanation: The Role of Accuracy, Informativity, and Causality
When people say what an agent believes, they prefer beliefs that are causally relevant to the agent's actions, more than beliefs that are merely accurate or informative.
-
Machine Theory of Mind and the Structure of Human Values
Human values are claimed to have a rational instrumental structure that lets AI infer unseen values from known ones, framing this as the 'value generalization problem' in AI safety.
Discussion (0). Sign in to comment.