Pith. sign in

REVIEW 5 major objections 4 minor 27 references

A single appraisal-based reward knob per disorder — anxiety, mania, OCD checking, depression, impulsivity, addiction, PTSD — produces graded, dose-dependent disorder-like behaviour in a reinforcement-learning agent, with an emergent two-dim

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 08:03 UTC pith:YI5NJHUT

load-bearing objection A serious, unusually systematic RL study of disorder-like phenotypes whose useful core (dose-response, rescue, comorbidity) is undermined by over-claimed monotonicity, a partly constructed 'emergent' space, and an unresolved seed-count inconsistency. the 5 major comments →

arxiv 2607.07753 v2 pith:YI5NJHUT submitted 2026-07-08 cs.LG cs.AI

A Transdiagnostic Space of Disorder Like Phenotypes in Reinforcement Learning Agents

classification cs.LG cs.AI
keywords appraisal theoryreinforcement learningcomputational psychiatryreward shapingdisorder phenotypesdose-responsetransdiagnosticexposure therapy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that seven psychiatric disorder-like phenotypes can be induced, measured, and reversed in a reinforcement-learning agent by manipulating a single cognitive appraisal signal per disorder, and that this knob space has emergent structure not written into the rewards. Across more than 1,375 runs with confidence intervals and four controls, every disorder shows a graded, monotone dose-response that no control reproduces. Three findings emerge without being encoded in the reward: disorders self-organize into a two-dimensional affective space where mania mirrors anxiety; removing the knob remits reward-distortion disorders but not avoidance disorders, which recover only under graded exposure; and pairs of knobs interact nonadditively, yielding testable comorbidity predictions. The framework generalizes beyond grid worlds: depression and addiction show a double dissociation in a 3D pixel environment with a standard convolutional agent and no appraisal critic. A sympathetic reader would care because this offers a controllable, falsifiable testbed for computational psychiatry rather than post-hoc description.

Core claim

The central claim is that a single reward-shaping term or discount change, grounded in a named computational psychiatry account, produces a graded, monotone dose-response for each of seven disorders, and no control condition reproduces any phenotype. Three results are emergent rather than induced: the seven disorders occupy a two-dimensional affective space recovered by data-driven embedding in which mania is the behavioural mirror of anxiety; removing the knob remits reward-distortion disorders (mania, checking, addiction) but not avoidance disorders (anxiety, PTSD), which recover only under a graded exposure curriculum with response prevention; and two simultaneous knobs interact nonadditi

What carries the argument

AG-PPO, an appraisal-guided PPO agent whose critic consumes a six-dimensional appraisal vector (motivational relevance, goal congruence, certainty, novelty, coping potential, anticipation) computed from geometric, information-theoretic, and predictive signals. Each disorder is isolated to a single 'knob': a weight in the reward-shaping function or a discount-factor change, grounded in a named account (e.g., Redish for addiction, Treadway–Zald for depression). Behavioural symptoms are measured by preregistered assays, and four control conditions (standard PPO, critic-noise, PPO+RND, and the appraisal critic without shaping) anchor every comparison.

Load-bearing premise

The load-bearing premise is that each psychiatric disorder can be faithfully represented by a single hand-calibrated reward-shaping term (or discount change) and that the resulting behavioral assays are valid proxies for clinical symptoms — the paper admits these labels have not been validated against human or animal behavior data; if this mapping is false, the transdiagnostic space and treatment resistance describe the reward-shaping scheme, not the disorders.

What would settle it

A concrete check: if any matched-magnitude corruption of the appraisal signal (shuffle, random, or shift) also produced the avoidance phenotype in the Approach–Avoidance experiment, the claim that phenotypes require appraisal-contingent shaping would collapse. A second: a human or animal study of manic episodes with comorbid impulsivity that found impulsivity compounds rather than reduces risk-taking lethality would falsify the mania×impulsivity prediction. A third: if a non-contingent penalty of equal magnitude and frequency produced the same dose-response curves, the specificity claim would

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Disorder modelling becomes dose-controllable and falsifiable: any proposed mechanism must produce a graded, monotone phenotype that no matched control reproduces.
  • The remit-versus-resist dissociation predicts that removing a stressor suffices for reward-distortion disorders, whereas avoidance disorders require active graded exposure with response prevention, paralleling clinical practice.
  • The mania×impulsivity interaction yields a concrete, falsifiable prediction: comorbid impulsivity should reduce, not compound, the lethality of manic risk-taking, because dangerous acts require sustained multi-step pursuit that steep discounting undercuts.
  • The rapid-induction result implies the pathological policy is latent in a healthy agent and reachable along a low-cost path, supporting a manifold view of psychopathology rather than a collection of unrelated reward hacks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • This reader's inference: the same appraisal-knob methodology could be applied to other RL algorithms and environments to test whether the dose-response and emergent geometry are properties of value learning generally or of PPO specifically.
  • This reader's inference: the data-driven embedding of behavioral assays, here a first-step PCA, points toward a research program that replaces diagnostic labels with continuous transdiagnostic dimensions, consistent with dimensional frameworks in psychiatry.
  • This reader's inference: the exposure-curriculum result suggests an agent-based optimisation problem — how to schedule response-prevention prompts to minimise recovery time — which the paper's annealed-penalty curriculum only begins to explore.
  • This reader's inference: if the knob-to-disorder mapping is ever validated against human or animal data, the dose-response curves could calibrate individualised computational models, but the paper itself notes that validation is still missing.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper introduces AG-PPO, an appraisal-guided PPO agent in which seven psychiatric-disorder-like phenotypes are induced by single reward-shaping knobs (anxiety, mania, OCD checking, depression, impulsivity, addiction, PTSD). Each disorder has a preregistered primary assay, and the authors report dose–response curves across 1,375 runs with 10 seeds, four control conditions, and 95% CIs. The central claimed contributions are: (i) every disorder shows a graded, monotone dose–response that no control reproduces; (ii) the disorders self-organise into a two-dimensional affective space, with mania mirroring anxiety, recovered without hand-chosen axes via PCA; (iii) a remit-versus-resist treatment dissociation under passive knob removal, with avoidance disorders recovering under a graded exposure curriculum; (iv) nonadditive comorbidity interactions; and (v) a 3D pixel-environment double dissociation for depression and addiction using a standard CNN with no appraisal critic.

Significance. If the results hold, the paper would provide a unusually systematic, reproducible framework for computational psychiatry: 10 seeds, confidence intervals, preregistered assays, four controls, an explicit negative result (the OCD certainty mechanism), a reported 3D null, and a single analysis script regenerating all tables and figures. These are genuine strengths. However, the paper's primary contribution is the set of 'emergent' findings, and both the dose–response support and the emergence claims contain internal contradictions that materially affect the conclusions. The stress-test concern about circularity lands in part: the mania–anxiety mirror is at least partially imposed by the opposite signs of the same shaping term, and the PCA input is the pre-registered assay battery whose components are the primary symptom definitions for the corresponding disorders. The silhouette coefficient of 0.31 is modest, so the 'data-driven affective space' is not yet established as an independent discovery.

major comments (5)
  1. [Table 2 and Abstract] The central claim that 'every disorder shows a graded, monotone dose-response' is contradicted by the reported numbers. Impulsivity Near-reward choice: PPO 0.46, ϵ1 0.61, ϵ2 0.56, ϵ3 1.00, ϵ4 0.94 — a decrease from ϵ1 to ϵ2 and from ϵ3 to ϵ4. Addiction Drug occupancy: 0.00, 0.02, 0.00, 0.50, 0.76 — a decrease from ϵ1 to ϵ2. These are not monotone. Since the dose–response is the paper's validating induced-effect claim, this needs either corrected data/error bars or a revised, weaker formulation (e.g., 'graded with a threshold' for addiction and 'non-monotone at low doses' for impulsivity).
  2. [Eq. 2, Table 1, Fig. 4, Fig. 8] The claimed emergent affective space is partly constructed. Eq. 2 implements anxiety and mania as opposite signs of the same coping-potential penalty (w_lo vs w_hi on ζ_CP), and Table 1 explicitly labels mania as the 'opposite behavioural pole' of anxiety. The PCA in Fig. 8 uses as input the pre-registered assay vector whose dimensions were selected as the primary symptom for each disorder (risky-goal choice, forward-action fraction, drug occupancy, near-reward choice, etc.). Under those inputs, separating the disorders along axes tied to those same assays is expected. The claim that the mirror is 'not written into the reward' conflates reward-shaping with assay selection. Silhouette 0.31 is modest. Please reframe the affective space as a consistency check of the assay choice, or provide a projection that does not use the disorder-defining assays as input.
  3. [3D Pixel Generalisation, Table 15, Fig. 5] The text states that for addiction 'forward fraction is unchanged from baseline (0.73→0.76)', and Fig. 5 reports 'no locomotion suppression (≈0.10, 95% CI overlapping zero)'. Table 15, however, shows baseline fwd_frac 0.728, and addiction conditions 0.468 (ϵ=0.1), 0.488 (ϵ=0.3), 0.648 (ϵ=0.6) — a 0.26–0.08 reduction, not unchanged. This is an internal inconsistency in the central double-dissociation evidence. The 3D result may still hold, but the reported numbers must be corrected and the claim adjusted.
  4. [Table 4 and Table 11] The 'resist passive removal' claim for anxiety and PTSD is not supported by the reported statistics. The passive-removal endpoints for Anxiety risky choice and PTSD short-route use are both 0.20 with 95% CI ±0.39 at n=5 (Table 11); these intervals include zero, so there is no significant difference from the severe untreated model (0.00) or from full suppression. The dissociation in Table 4 therefore rests on means with very wide CIs for the two avoidance disorders. More seeds or a proper significance test comparing passive vs. control vs. severe are needed before claiming a clean remit-versus-resist split.
  5. [Table 10 and Fig. 4] The claim that 'mania is the mirror of anxiety across the origin' is not borne out by the coordinates used. Table 10 lists Anxiety at (reward–approach +0.0, threat–avoidance +1.0) and Mania at (+1.0, −0.2). A reflection across the origin would require Mania at (0.0, −1.0). The two are not antipodal. The figure should either use coordinates consistent with the text or describe the relationship as 'opposite on the threat–avoidance dimension' rather than a mirror across the origin.
minor comments (4)
  1. [Abstract/Supplement] The abstract says '1,375 configurations' and the supplement also uses '1,375 runs'; be consistent about whether the number includes the 36 MiniWorld runs.
  2. [Eq. 4] The denominator of ζ_GC, √(((n−1)/2)^2+n^2), is described as the half-width-and-full-height diagonal of the egocentric view; please double-check this normalization, as it is not immediately obvious and the maximum distance inside the view may be different for the adopted grid.
  3. [Fig. 2] The color scale is described as 'brighter is more visited', but the panel for depression (stationary at start) appears qualitatively different from the quantitative forward-action fraction. A color bar or a note on normalization would improve readability.
  4. [Rescue section] The phrase 'the answer splits cleanly along mechanistic lines' is stronger than the data justify given the wide CIs in Table 11; consider using more cautious wording such as 'the pattern is consistent with…'.

Circularity Check

3 steps flagged

The 'emergent' affective geometry is largely constructed: mania and anxiety are opposite signs of a single coping-potential knob, and both the hand-chosen axes and the PCA input are the same disorder-defining assay battery.

specific steps
  1. self definitional [Method, Eq. 2; 'Disorder mechanisms' / Table 1]
    "the anxiety and mania knobs act on coping potential ζCP (penalising low or high values respectively) ... Anxiety penalises low coping potential (a threat in view), inducing avoidance; mania is its opposite behavioural pole, penalising high coping potential and thus seeking threat."

    The two disorders are the same reward-shaping term with opposite sign on the same appraisal ζ_CP, and ζ_CP is defined as the complement of threat visibility (Eq. 7). Therefore 'mania mirrors anxiety' is an algebraic consequence of the sign of one hand-set weight, not an emergent property of value learning. The paper's own language ('opposite behavioural pole') states the construction. The PCA defense does not change this: it is applied to an assay vector whose central directions are the same threat-approach/avoidance readouts used to define the two disorders.

  2. renaming known result [Supplementary Material, 'Affective-Space Construction' / Table 10]
    "The reward-approach axis x combines reward-pursuit markers (forward-action fraction for depression, drug occupancy for addiction, near-reward choice for impulsivity, risk-taking for mania); the threat-avoidance axis y uses threat and trauma distance (positive for anxiety and PTSD, negative for mania)."

    The two axes of Fig. 4 are literally composed of the pre-registered primary assays for each disorder (Table 2). Plotting each disorder in coordinates built from its own defining symptom is a re-plot of the definitions, not an independently discovered organization. The coordinate table (mania y=-0.2 vs anxiety y=+1.0) is determined by the sign of the shared coping-potential knob, so the claimed 'reflection across the origin' is built into the axes.

  3. renaming known result ['A transdiagnostic affective space' / Supplementary 'Data-Driven Recovery of the Affective Space']
    "we run principal-component analysis on the full vector of pre-registered behavioural assays, with no hand-chosen axes and no disorder labels supplied to the projection."

    A PCA can only rotate the variables it is given. Those variables are the disorder-defining assays fixed in advance (risky-goal choice, mean threat distance, checking rate, forward-action fraction, near-reward choice, drug occupancy, trauma distance). Separating disorders along linear combinations of the very measures used to define and verify each knob is expected; the modest silhouette (0.31) and the appendix admission that low doses sit 'at the healthy origin by construction' show the embedding tracks the knob-induced assay shifts rather than an independently latent psychiatric space.

full rationale

The paper is careful to separate induced dose-response effects from emergent ones, and the dose-response, control, rescue, comorbidity, representational, and 3D-transfer results are not themselves circular. The circularity is concentrated in the paper's primary emergent claim, the transdiagnostic affective space. Eq. 2 makes mania and anxiety opposite signs of one coping-potential penalty, and the affective-space axes/PCA input are drawn from the same pre-registered assay battery that defines each disorder, so the mania-anxiety mirror and the overall geometry are substantially constructed rather than discovered. The rescue dissociation and nonadditive comorbidity are genuinely emergent from value learning and keep the paper from being wholly reducible to its inputs; the self-citation to prior appraisal-guided PPO work is descriptive rather than load-bearing. The paper's own limitation statement ('labels have not been validated against human or animal behaviour data') is an external-validity caveat, not a circularity, but it does not repair the constructed geometry.

Axiom & Free-Parameter Ledger

9 free parameters · 5 axioms · 0 invented entities

The central results rest on hand-calibrated reward terms, hand-chosen appraisal operationalizations, and a convergence filter. None of these are anchored to independent external benchmarks or human/animal data, and several were explicitly tuned via sweeps to produce graded dose-responses.

free parameters (9)
  • anxiety/mania coping-potential penalty weight w_CP = doses 0.01, 0.03, 0.1, 0.3
    Calibrated by hand so the agent produces a graded dose-response rather than degenerate behavior.
  • OCD checking bonus b_chk with habituation lambda=0.5 = doses 0.1, 0.2, 0.4, 0.6; lambda=0.5
    Selected by sweeping candidate values to produce rising checking while maintaining task success.
  • depression effort cost c_eff = doses 0.01, 0.02, 0.05, 0.1
    Hand-calibrated to collapse forward action at severe doses.
  • impulsivity discount reduction 1-gamma = doses 0.05, 0.1, 0.2, 0.4
    Chosen to span the expected discounting crossover.
  • addiction drug bonus = doses 0.01, 0.02, 0.05, 0.1
    Calibrated by locating the threshold at which drug occupancy overtakes goal reward.
  • PTSD trauma shock b_shk = doses 0.05, 0.1, 0.2, 0.4
    Hand-set to produce graded avoidance without collapsing task success.
  • exposure penalty coefficient = 0.5 after a short sweep
    Selected by sweeping to make avoided-route penalty effective.
  • stress-index weights = (0.25, 0.05, 0.1, 0.2, 0.35, 0.05)
    Carried over from the base model by hand; no independent derivation.
  • learning rate and entropy coefficient = lr=1e-3, entropy=0.03
    Adjusted from default PPO after baseline freezing; selected to make all environments solvable.
axioms (5)
  • domain assumption The six appraisal equations meaningfully operationalize relevance, congruence, certainty, novelty, coping potential, and anticipation.
    The paper maps appraisal-theory constructs to geometric and information-theoretic quantities (Eqs. 3-8) with no external validation of the mapping.
  • domain assumption A single reward-shaping term is sufficient to model each disorder.
    Table 1 assigns each disorder one knob; the paper explicitly notes real disorders are "multiply determined" and this is a simplification.
  • domain assumption The convergence criterion (success >= 50%) selects an unbiased subset of runs for symptom analysis.
    The paper says the criterion was fixed before analysis, but no preregistration artifact is provided to verify this.
  • domain assumption Simulated treatment responses in these agents are informative about human exposure therapy and comorbidity.
    The interpretation of the rescue and exposure experiments relies on this analogy, which the paper does not validate against clinical data.
  • standard math Standard PPO and convolutional-network training converge to the reported behaviors.
    Uses standard PPO (Eq. 1) and CNN training; accepted as background.

pith-pipeline@v1.3.0-alltime-deepseek · 223 in / 8702 out tokens · 124412 ms · 2026-08-02T08:03:33.655997+00:00 · methodology

0 comments
read the original abstract

Modelling psychological disorders in artificial agents offers a testbed for computational psychiatry and a lens on affective-control failure modes. Prior work induces one or two disorders by hand-tuned reward shaping, labels the behaviour post hoc, and reports single runs. We recast disorder modelling as dose-controllable manipulation of cognitive appraisal signals in an appraisal-guided PPO agent, expressing seven disorders (anxiety, mania, obsessive-compulsive checking, depression, impulsivity, addiction, and post-traumatic stress) each as a single knob grounded in a computational psychiatry account, with each symptom measured by a preregistered assay. Across more than a thousand runs (10 seeds, four controls, 95% confidence intervals) every disorder shows a graded, monotone dose-response that no control reproduces. Beyond these induced effects, three findings emerge that were not written into the reward: disorders self-organise into a two-dimensional affective space in which mania mirrors anxiety; removing a knob remits reward-distortion disorders (mania, checking, addiction) but not avoidance disorders (anxiety, PTSD), which recover under a graded exposure curriculum; and two simultaneous knobs interact nonadditively, yielding testable comorbidity predictions. The depression and addiction knobs further reproduce their double dissociation in a 3D pixel environment (MiniWorld) with a standard convolutional agent and no appraisal critic, showing the framework generalises beyond grid worlds.

Figures

Figures reproduced from arXiv: 2607.07753 by Hari Prasad.

Figure 1
Figure 1. Figure 1: AG-PPO. A shared convolutional encoder feeds [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Phenotype gallery: state-occupancy over 40 episodes at the severe dose of each disorder, overlaid on the environment [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Dose–response of the primary symptom assay per disorder (mean, shaded 95% CI, 10 seeds); dashed lines are the [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The transdiagnostic affective space. Each disorder [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 4
Figure 4. Figure 4: The transdiagnostic affective space. Each disorder [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 2
Figure 2. Figure 2: the anxiety avoidance band, the mania approach to [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 5
Figure 5. Figure 5: Rescue trajectories: primary symptom over continued training with the knob removed (treatment) vs. kept (control), [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Comorbidity: joint dose grids for two knob pairs [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 6
Figure 6. Figure 6: The seven environments: four threat/spatial grids and three custom environments realising delay discounting, drug [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Block-diagonal cross-assay dissociation across all seven disorders. Rows are disorder conditions (at high dose); [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗
Figure 7
Figure 7. Figure 7: Dose–response of the primary symptom assay per [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: The seven environments: four threat/spatial grids and three custom environments realising delay discounting, drug [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 8
Figure 8. Figure 8: Data-driven affective space. PCA of the pre [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 2
Figure 2. Figure 2: The signatures are disorder-specific rather than task￾specific: the four control agents (PPO, critic-noise, PPO+RND, appraisal critic without shaping) run on the same environments and produce uniform or random oc￾cupancy patterns with no spatial clustering. The pheno￾type gallery therefore provides face validity at a glance: an anxiety-like agent crowds the periphery, a PTSD-like agent takes the long safe … view at source ↗
Figure 10
Figure 10. Figure 10: Rapid induction. Dashed: phenotype induced by [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 9
Figure 9. Figure 9: Grouped bar chart of all three assay metrics across [PITH_FULL_IMAGE:figures/full_fig_p015_9.png] view at source ↗
Figure 13
Figure 13. Figure 13: Rescue trajectories: primary symptom over con [PITH_FULL_IMAGE:figures/full_fig_p014_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Comorbidity: joint dose grids for two knob pairs [PITH_FULL_IMAGE:figures/full_fig_p015_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Block-diagonal cross-assay dissociation across all [PITH_FULL_IMAGE:figures/full_fig_p016_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Grouped bar chart of all three assay metrics across [PITH_FULL_IMAGE:figures/full_fig_p016_16.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

27 extracted references · 3 linked inside Pith

  1. [1]

    H.; and Koob, G

    Ahmed, S. H.; and Koob, G. F. 1998. Transition from moderate to excessive drug intake: change in hedonic set point. Science 282(5387):298--300

  2. [2]

    Ainslie, G. 1975. Specious reward: A behavioral theory of impulsiveness and impulse control. Psychological Bulletin 82(4):463--496

  3. [3]

    Prasad, H.; Jacob, C.; and Ahamed, I. 2024. Appraisal-Guided Proximal Policy Optimization: Modeling Psychological Disorders in Dynamic Grid World. arXiv:2407.20383

  4. [4]

    G.; Treanor, M.; Conway, C

    Craske, M. G.; Treanor, M.; Conway, C. C.; Zbozinek, T.; and Vervliet, B. 2014. Maximizing exposure therapy: An inhibitory learning approach. Behaviour Research and Therapy 58:10--23

  5. [5]

    Huys, Q. J. M.; Maia, T. V.; and Frank, M. J. 2016. Computational psychiatry as a bridge from neuroscience to clinical applications. Nature Neuroscience 19(3):404--413

  6. [6]

    L.; Edge, M

    Johnson, S. L.; Edge, M. D.; Holmes, M. K.; and Carver, C. S. 2012. The behavioral activation system and mania. Annual Review of Clinical Psychology 8:243--267

  7. [7]

    N.; Petry, N

    Kirby, K. N.; Petry, N. M.; and Bickel, W. K. 1999. Heroin addicts have higher discount rates for delayed rewards than non-drug-using controls. Journal of Experimental Psychology: General 128(1):78--87

  8. [8]

    Kullback, S.; and Leibler, R. A. 1951. On information and sufficiency. Annals of Mathematical Statistics 22(1):79--86

  9. [9]

    Lazarus, R. S. 1991. Emotion and Adaptation. Oxford University Press

  10. [10]

    V.; and Frank, M

    Maia, T. V.; and Frank, M. J. 2011. From reinforcement learning models to psychiatric and neurological disorders. Nature Neuroscience 14(2):154--162

  11. [11]

    R.; and Quirk, G

    Milad, M. R.; and Quirk, G. J. 2012. Fear extinction as a model for translational neuroscience: Ten years of progress. Annual Review of Psychology 63:129--151

  12. [12]

    M.; Broekens, J.; and Jonker, C

    Moerland, T. M.; Broekens, J.; and Jonker, C. M. 2018. Emotion in reinforcement learning agents and robots: A survey. Machine Learning 107(2):443--480

  13. [13]

    R.; Dolan, R

    Montague, P. R.; Dolan, R. J.; Friston, K. J.; and Dayan, P. 2012. Computational psychiatry. Trends in Cognitive Sciences 16(1):72--80

  14. [14]

    Rachman, S. 2002. A cognitive theory of compulsive checking. Behaviour Research and Therapy 40(6):625--639

  15. [15]

    Redish, A. D. 2004. Addiction as a computational process gone awry. Science 306(5703):1944--1947

  16. [16]

    Scherer, K. R. 2001. Appraisal considered as a process of multilevel sequential checking. In Appraisal Processes in Emotion, 92--120. Oxford University Press

  17. [17]

    Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017. Proximal policy optimization algorithms. arXiv:1707.06347

  18. [18]

    S.; and Paiva, A

    Sequeira, P.; Melo, F. S.; and Paiva, A. 2011. Emotion-based intrinsic motivation for reinforcement learning agents. In ACII, 326--336

  19. [19]

    T.; and Zald, D

    Treadway, M. T.; and Zald, D. H. 2011. Reconsidering anhedonia in depression: Lessons from translational neuroscience. Neuroscience & Biobehavioral Reviews 35(3):537--555

  20. [20]

    S.; and Terry, J

    Chevalier-Boisvert, M.; Dai, B.; Towers, M.; de Lazcano, R.; Willems, L.; Lahlou, S.; Pal, S.; Castro, P. S.; and Terry, J. 2023. Minigrid & Miniworld: Modular & Customizable Reinforcement Learning Environments for Goal-Oriented Tasks. arXiv:2306.13831

  21. [21]

    Huang, S.; Dossa, R. F. J.; Ye, C.; Braga, J.; Chakraborty, D.; Mehta, K.; and Araújo, J. G. 2022. CleanRL: High-quality single-file implementations of deep reinforcement learning algorithms. JMLR 23(274):1--18

  22. [22]

    A.; Veness, J.; Bellemare, M

    Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015. Human-level control through deep reinforcement learning. Nature 518(7540):529--533

  23. [23]

    S.; Quinn, K.; Sanislow, C.; and Wang, P

    Insel, T.; Cuthbert, B.; Garvey, M.; Heinssen, R.; Pine, D. S.; Quinn, K.; Sanislow, C.; and Wang, P. 2010. Research domain criteria (RDoC): Toward a new classification framework for research on mental disorders. American Journal of Psychiatry 167(7):748--751

  24. [24]

    Beck, A. T. 1979. Cognitive Therapy of Depression. Guilford Press

  25. [25]

    K.; Rasmusson, A

    Pitman, R. K.; Rasmusson, A. M.; Koenen, K. C.; Shin, L. M.; Orr, S. P.; Gilbertson, M. W.; Milad, M. R.; and Liberzon, I. 2012. Biological studies of post-traumatic stress disorder. Nature Reviews Neuroscience 13(11):769--787

  26. [26]

    R.; Menzies, L.; Hampshire, A.; Suckling, J.; Fineberg, N

    Chamberlain, S. R.; Menzies, L.; Hampshire, A.; Suckling, J.; Fineberg, N. A.; del Campo, N.; Aitken, M.; Craig, K.; Owen, A. M.; Bullmore, E. T.; Robbins, T. W.; and Sahakian, B. J. 2008. Orbitofrontal dysfunction in patients with obsessive-compulsive disorder and their unaffected relatives. Science 321(5887):421--422

  27. [27]

    E.; and Berridge, K

    Robinson, T. E.; and Berridge, K. C. 1993. The neural basis of drug craving: An incentive salience theory of addiction. Brain Research Reviews 18(3):247--291