Seven disorder-like phenotypes in PPO agents emerge from single appraisal/reward knobs, organize into a two-dimensional affective space, and split into remitting vs treatment-resistant classes.
Appraisal-Guided Proximal Policy Optimization: Modeling Psychological Disorders in Dynamic Grid World
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The integration of artificial intelligence across multiple domains has emphasized the importance of replicating human-like cognitive processes in AI. By incorporating emotional intelligence into AI agents, their emotional stability can be evaluated to enhance their resilience and dependability in critical decision-making tasks. In this work, we develop a methodology for modeling psychological disorders using Reinforcement Learning (RL) agents. We utilized Appraisal theory to train RL agents in a dynamic grid world environment with an Appraisal-Guided Proximal Policy Optimization (AG-PPO) algorithm. Additionally, we investigated numerous reward-shaping strategies to simulate psychological disorders and regulate the behavior of the agents. A comparison of various configurations of the modified PPO algorithm identified variants that simulate Anxiety disorder and Obsessive-Compulsive Disorder (OCD)-like behavior in agents. Furthermore, we compared standard PPO with AG-PPO and its configurations, highlighting the performance improvement in terms of generalization capabilities. Finally, we conducted an analysis of the agents' behavioral patterns in complex test environments to evaluate the associated symptoms corresponding to the psychological disorders. Overall, our work showcases the benefits of the appraisal-guided PPO algorithm over the standard PPO algorithm and the potential to simulate psychological disorders in a controlled artificial environment and evaluate them on RL agents.
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Transdiagnostic Space of Disorder Like Phenotypes in Reinforcement Learning Agents
Seven disorder-like phenotypes in PPO agents emerge from single appraisal/reward knobs, organize into a two-dimensional affective space, and split into remitting vs treatment-resistant classes.