Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.
Embedded Agency
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Empathic DQN augments DQN value estimates with an empathy term computed by swapping the learning agent into other agents' situations, reducing collateral harms in two gridworld proof-of-concept environments.
Presents a taxonomy of wireheading in partially embedded agents, defines wirehead-vulnerable agents, demonstrates via AIXIjs simulation, and conjectures that specification gaming is the only other misalignment type.
Autotelic AI requires agents to generate and relativize their own self-boundaries in embedded settings, with the paper consolidating this into a framework extended to quantum, philosophical, and LLM contexts.
citing papers explorer
-
Safety from Honesty in a Disinterested AI Predictor
Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.
-
Towards Empathic Deep Q-Learning
Empathic DQN augments DQN value estimates with an empathy term computed by swapping the learning agent into other agents' situations, reducing collateral harms in two gridworld proof-of-concept environments.
-
Categorizing Wireheading in Partially Embedded Agents
Presents a taxonomy of wireheading in partially embedded agents, defines wirehead-vulnerable agents, demonstrates via AIXIjs simulation, and conjectures that specification gaming is the only other misalignment type.
-
The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self
Autotelic AI requires agents to generate and relativize their own self-boundaries in embedded settings, with the paper consolidating this into a framework extended to quantum, philosophical, and LLM contexts.