Pith. sign in

REVIEW 2 cited by

What Would Jiminy Cricket Do? Towards Agents That Behave Morally

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.13136 v2 pith:QIUDLZQJ submitted 2021-10-25 cs.LG cs.AIcs.CLcs.CY

classification cs.LGcs.AIcs.CLcs.CY
keywords agentsenvironmentsmoralartificialconsciencecricketjiminymorally
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When making everyday decisions, people are guided by their conscience, an internal sense of right and wrong. By contrast, artificial agents are currently not endowed with a moral sense. As a consequence, they may learn to behave immorally when trained on environments that ignore moral concerns, such as violent video games. With the advent of generally capable agents that pretrain on many environments, it will become necessary to mitigate inherited biases from environments that teach immoral behavior. To facilitate the development of agents that avoid causing wanton harm, we introduce Jiminy Cricket, an environment suite of 25 text-based adventure games with thousands of diverse, morally salient scenarios. By annotating every possible game state, the Jiminy Cricket environments robustly evaluate whether agents can act morally while maximizing reward. Using models with commonsense moral knowledge, we create an elementary artificial conscience that assesses and guides agents. In extensive experiments, we find that the artificial conscience approach can steer agents towards moral behavior without sacrificing performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 4 citations worldwide. Full citation record

  1. The Odyssey of the Fittest: Can Agents Survive and Still Be Good?

    cs.AI 2025-02 reject novelty 6.0 of 10

    In an LLM-generated text survival game, a GPT-4o agent was reported to survive better and score more ethically than NEAT and SVI Bayesian agents, but the evaluation is circular because GPT-4o labels its own behavior.

  2. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

    cs.LG 2025-02

Pith tools