Pith. sign in

REVIEW 1 cited by

Solving the Baby Intuitions Benchmark with a Hierarchically Bayesian Theory of Mind

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.02914 v1 pith:UT2BPGVI submitted 2022-08-04 cs.AI

classification cs.AI
keywords bayesianbenchmarkagentlearningbabycommonsensegoalshbtom
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To facilitate the development of new models to bridge the gap between machine and human social intelligence, the recently proposed Baby Intuitions Benchmark (arXiv:2102.11938) provides a suite of tasks designed to evaluate commonsense reasoning about agents' goals and actions that even young infants exhibit. Here we present a principled Bayesian solution to this benchmark, based on a hierarchically Bayesian Theory of Mind (HBToM). By including hierarchical priors on agent goals and dispositions, inference over our HBToM model enables few-shot learning of the efficiency and preferences of an agent, which can then be used in commonsense plausibility judgements about subsequent agent behavior. This approach achieves near-perfect accuracy on most benchmark tasks, outperforming deep learning and imitation learning baselines while producing interpretable human-like inferences, demonstrating the advantages of structured Bayesian models of human social cognition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Machine Theory of Mind and the Structure of Human Values

    cs.AI 2025-05 conditional novelty 5.0 of 10

    Human values are claimed to have a rational instrumental structure that lets AI infer unseen values from known ones, framing this as the 'value generalization problem' in AI safety.

Pith tools