Pith. sign in

Embedded Agency

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

fields

cs.AI 3 cs.LG 1

years

2026 2 2019 2

representative citing papers

Safety from Honesty in a Disinterested AI Predictor

cs.AI · 2026-06-28 · conditional · novelty 7.0

Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.

Towards Empathic Deep Q-Learning

cs.LG · 2019-06-26 · unverdicted · novelty 6.0

Empathic DQN augments DQN value estimates with an empathy term computed by swapping the learning agent into other agents' situations, reducing collateral harms in two gridworld proof-of-concept environments.

Categorizing Wireheading in Partially Embedded Agents

cs.AI · 2019-06-21 · unverdicted · novelty 6.0

Presents a taxonomy of wireheading in partially embedded agents, defines wirehead-vulnerable agents, demonstrates via AIXIjs simulation, and conjectures that specification gaming is the only other misalignment type.

citing papers explorer

Showing 4 of 4 citing papers.

  • Safety from Honesty in a Disinterested AI Predictor cs.AI · 2026-06-28 · conditional · none · ref 25

    Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.

  • Towards Empathic Deep Q-Learning cs.LG · 2019-06-26 · unverdicted · none · ref 3

    Empathic DQN augments DQN value estimates with an empathy term computed by swapping the learning agent into other agents' situations, reducing collateral harms in two gridworld proof-of-concept environments.

  • Categorizing Wireheading in Partially Embedded Agents cs.AI · 2019-06-21 · unverdicted · none · ref 6

    Presents a taxonomy of wireheading in partially embedded agents, defines wirehead-vulnerable agents, demonstrates via AIXIjs simulation, and conjectures that specification gaming is the only other misalignment type.

  • The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self cs.AI · 2026-06-18 · unverdicted · none · ref 66

    Autotelic AI requires agents to generate and relativize their own self-boundaries in embedded settings, with the paper consolidating this into a framework extended to quantum, philosophical, and LLM contexts.