Pith. sign in

REVIEW 3 major objections 5 minor 86 references

An Affective-Taxis Hypothesis for Alignment and Interpretability

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper argues that affective valence is taxis navigation through an internal interoceptive landscape, so aligned AI must represent that landscape.

desk verdict The reward-as-directional-derivative identity is a promising but unproven hypothesis; the paper's honest limitations are its best feature. read the letter →

arxiv 2505.17024 v1 pith:LILBGQJD submitted 2025-05-03 cs.AI q-bio.NC

classification cs.AIq-bio.NC
keywords AIalignmentaffectivevalencetaxisnavigationinteroceptionfold-changedetectionrewardfunctionenergy-basedmodelsC.elegans
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the AI alignment problem is at root a problem about affect: human goals and values are not a scalar utility but the felt gradient of an internal taxis landscape, and an aligned AI must be able to model that landscape. The authors propose that reward in the brain is the directional derivative of a time-varying interoceptive energy function along the current direction of movement, making valence a measurable slope rather than an opaque number. If true, this would give AI alignment a concrete inductive bias to overcome the mathematical result that reward functions cannot be inferred from behavior alone, since behavior conflates beliefs and values. The paper grounds the claim in evolutionary-developmental neuroscience and in a computational model tested against C. elegans, a model organism whose taxis behavior lacks temporal associative learning and so exposes the gradient signal directly.

What carries the argument

The load-bearing object is the directional derivative of a log-density, $\nabla_z \log \gamma(z;\beta)\cdot v$, where $\gamma$ is the spatial density of attractants or, in the internalized case, the density of an allostatic energy function, and $v$ is the current velocity. The paper's central move is to equate this quantity with reward: instead of a scalar utility assigned to states, the agent reads the slope of its affective landscape in the direction it is moving. Fold-change detection is the neural mechanism said to implement this readout, since it computes the time-derivative of the log input, and energy-based models are proposed as the compositional representation that lets the same gradient logic pass from physical space to abstract interoceptive space. The identity turns taxis behavior into a Langevin-style gradient-biased random walk, so that exploration and exploitation are balanced by the same quantity that carries valence.

What would settle it

Take a preparation in which taxis-sensing neurons are recorded while the animal moves through a controlled attractant field; if the neural response is not proportional to $d\log \gamma(z(t))/dt$, encoding absolute concentration rather than fold-change, the proposed identity between reward and the directional derivative fails. A complementary check would test whether dopamine release in a vertebrate tracks the instantaneous gradient slope in the current movement direction rather than a temporally cached value prediction.

Watch

Extended reading notes

Core claim

The paper's central claim is that the brain's reward function consists of the negative directional derivative, in the present direction of movement, of a time-varying interoceptive energy function. Affect, on this view, is not a reaction to rewards but the navigation of an internal landscape: positive and negative valence are the signs of moving up or down gradients in a space of viscerosensory physiological indicators. The authors make this concrete by modeling reward as $R(s,a)=\nabla_z \log \gamma(z;\beta)\cdot v$ in a POMDP whose state includes the animal's location, orientation, and the spatial density of attractants, and they identify fold-change detection in C. elegans sensory neurons as the biological implementation of the log-gradient readout. They further argue that associative reinforcement learning evolved later to estimate the long-run reward rate over this taxis landscape, so present-oriented taxis can be studied in organisms that lack temporal associations. The consequence they draw for AI alignment is that agents must represent human affective states and the interoceptive facts grounding them, not merely observe choices or ratings.

Load-bearing premise

The paper's evolutionary grounding, that folding of the neural tube reoriented external taxis navigation toward an internal interoceptive landscape, is asserted without direct evidence, and the account of human affect depends on it.

Editorial extensions

If this is right

  • If reward is a directional derivative of an internal energy function, then reward functions are no longer unidentifiable in principle: the relevant quantity is the gradient field of an interoceptive landscape, not a free-standing scalar.
  • An aligned AI would need to infer and represent the human operator's time-varying interoceptive energy function, not just observe choices or ratings, because choices conflate beliefs and values.
  • The computational model can be tested in C. elegans, where taxis behavior is driven by present stimuli and no temporal associative learning confounds the reward signal.
  • Gradient-following by directional derivatives offers a biologically plausible alternative to backpropagation-based reward learning, since a single neuron can measure the derivative along its motion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: If reward really is a directional derivative of an energy field, inverse reinforcement learning could be reframed as estimating that field's gradient from trajectories, which may sidestep some non-identifiability results that assume a scalar reward.
  • Editorial inference: Interpretability would gain a visual meaning: an agent's values would be the geometry of its affective landscape, and aligning two agents would mean matching the shapes of their energy functions, not just their outputs.
  • Editorial inference: A direct human test would measure whether interoceptive or dopaminergic signals track the instantaneous directional derivative of an allostatic prediction error during movement, not just reward prediction error; such a signal would be the predicted valence readout.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This position paper proposes an 'affective-taxis hypothesis' according to which affective valence in humans and other bilaterians evolved from the internalization of taxis navigation. The authors formalize the core idea by identifying reward signals with the directional derivative of an interoceptive energy function: r(t) = ∇ log γ(z(t); β(t)) · dz/dt. They propose a POMDP model for taxis behavior in which rewards are exactly these directional derivatives, and they argue that fold-change detection (FCD) in C. elegans sensory neurons provides biological evidence for this quantity. The paper then argues that AI alignment requires agents to represent and model human affective states, and that the affective-taxis inductive bias could help overcome identifiability problems in inverse RL. The manuscript is primarily a conceptual and review contribution, with no experiments or implemented simulations; it explicitly lists its limitations and proposes directions for future work.

Significance. If the affective-taxis hypothesis is correct, it would give the brain's reward signal a concrete, measurable form (a directional derivative of an allostatic energy function), which could directly inform the design of interpretable and aligned AI systems by grounding reward in interoceptive physiology. The paper's strengths include a clearly articulated hypothesis, a simple mathematical formalization, a tractable model organism, and an explicit falsifiable prediction (that interoceptive or similar sensory neurons implement fold-change detection). The authors are also commendably transparent about the current lack of direct evidence and about the gap between C. elegans taxis and human affective experience. However, as the paper itself acknowledges, the central identity between reward and directional derivative remains a conjecture: the cited FCD evidence concerns chemosensory transduction, not reward or valence, and the evolutionary narrative connecting neural tube folding to internal taxis is speculative. The contribution is therefore a promising research program rather than an established result.

major comments (3)
  1. [Section 4 (Model Organism)] The central claim that the brain's reward function consists of the directional derivative of an interoceptive energy function is not supported by the evidence presented. The fold-change detection (FCD) results concern chemosensory neurons responding to environmental attractants, not reward or valence signals; the manuscript itself states that 'Direct evidence that interoceptive or similar sensory neurons... implement FCD would support the affective-taxis hypothesis.' This missing link is load-bearing because the alignment proposal depends on the assumption that the measured directional derivative is the actual reward/valence signal. The authors should explicitly separate the mathematical identity r(t) = d log γ/dt from the empirical conjecture that this quantity constitutes valence, and should propose a concrete experimental test (e.g., optogenetic manipulation of FCD neurons coupled with preference or approach/avoidance assays) that could validate the valence interpretation.
  2. [Section 2 (Across the Affective Landscape)] The evolutionary narrative that 'the folding of the neural plate into a tube inside the organism reoriented the combined apical and blastoporal nervous systems towards navigating an internal taxis landscape' is presented without direct evidence. This step is crucial for moving from external taxis navigation to internal affective valence, yet it is asserted rather than argued from comparative data. The paper should either marshal empirical or phylogenetic evidence supporting this transition (e.g., from chordate neuroanatomy) or explicitly label this as a speculative component of the hypothesis and discuss how it could be tested or disconfirmed.
  3. [Related Work (Section 5)] The statement that 'the reward function in the brain... consists of the (negative) directional derivative, in the present direction of movement, of a time-varying, interoceptive energy function' is written as an established fact, but it is the paper's own hypothesis and is not proven by the cited literature. This overstatement could mislead readers about the epistemic status of the claim. I recommend rephrasing to 'we hypothesize' or 'our proposal implies' and providing a clear account of how the hypothesis would be falsified.
minor comments (5)
  1. [Section 4 (Model Organism)] The notation 'r(t) :≈ d/dt log γ(z(t); β(t))' uses a nonstandard symbol '≈' to introduce a definition; the equality r(t) = d log γ/dt is a mathematical identity when v = dz/dt, and the approximation should be stated separately, e.g., with explicit assumptions about the FCD transduction.
  2. [Throughout] The organism name should be consistently capitalized as 'C. elegans' rather than 'c. elegans' to conform to standard biological nomenclature.
  3. [Title and headers] In the provided text, the title contains a line-break artifact 'Affective-T axis Hypothesis'; ensure the camera-ready version does not insert a space in 'Taxis'.
  4. [Section 3 (Computational Modeling Progress)] In Definition 1, the POMDP tuple uses 'pS' and 'pO'; while acceptable, the notation would be clearer with standard symbols such as 'T' and 'O' or with explicit subscripts, and the definition of the observation model as the gradient ∇ log γ should be connected to the FCD discussion in Section 4.
  5. [References] Several references are to unpublished preprints or online posts (e.g., [14], [65]); if the journal permits, please add preprint identifiers or DOI numbers to improve verifiability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the directional-derivative reward is an explicitly stated hypothesis, not a derived prediction.

full rationale

The paper's central claim is presented as a hypothesis rather than as a derived result. In Section 3, the POMDP reward R(s,a)=∇z log γ(z;β)·v is introduced as a modeling choice ('could consist precisely of the gradient's directional derivative'), and in Section 4 the authors explicitly say 'we include as part of our affective-taxis hypothesis that scalar "rewards" capture the directional derivatives'. The FCD identity r(t)=d/dt log γ(z(t);β(t)) = ∇z log γ·dz/dt is a mathematical consequence of fold-change detection, not of the reward definition; the two are linked by an explicit hypothesis, not by construction. The leap from FCD in C. elegans sensory neurons to the mammalian brain's reward function is acknowledged as under-supported: 'Direct evidence that interoceptive or similar sensory neurons... implement FCD would support the affective-taxis hypothesis' and 'direct experimental evidence remains limited.' This is an empirical gap, not a circular reduction. Self-citations (e.g., [36], [65], [68]) support auxiliary claims about interoception and active-inference formalism, but the load-bearing evolutionary and computational premises rest on external work (Cisek [21], Bennett [12], Karin & Alon [45], Shenhav [70]). The Limitations section further concedes that human affect requires temporal associations, situated conceptualizations, and theory of mind, confirming that the paper does not claim a closed derivation from FCD to human reward. No equation is shown to equal its own input by construction, and no fitted parameter is renamed as a prediction. Therefore the paper is not circular, though its central empirical claim is currently underdetermined.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The paper contributes no fitted parameters or invented entities. Its central claim depends on several domain assumptions taken from affective neuroscience, including the affective gradient hypothesis, the interoceptive basis of affect, and active inference, as well as on its own speculative evolutionary narrative. The axioms above capture these dependencies.

assumptions (7)
  • domain assumption Affective valence is a core evaluative process that guides motivated behavior via gradients (affective gradient hypothesis, ref [70]).
    The paper's opening premise is that affect, not rational calculation, is the brain's primary evaluative system, following the affectivist literature.
  • ad hoc to paper Taxis navigation predates and gives rise to affective valence through internalization of the taxis landscape (the paper's own hypothesis).
    This is the central proposal, stated in Section 2 as the affective-taxis hypothesis; it is not independently established.
  • domain assumption C. elegans sensory neurons implement fold-change detection, so their responses equal the time-derivative of log stimulus intensity.
    The paper cites Adler and Alon and Lang and Sontag for this; it is an established result in systems biology.
  • domain assumption C. elegans does not engage in temporal associative learning, so its taxis behavior reflects the immediate stimulus gradient, not learned predictions.
    The paper relies on this to claim its model bypasses the RL identifiability problem; it is stated in Section 4 but not demonstrated in this paper.
  • domain assumption The brain does not learn via backpropagation of errors.
    The paper cites Lillicrap et al. and Song et al.; this assumption motivates the need for alternative gradient computations.
  • domain assumption Active inference and energy-based models are appropriate normative frameworks for modeling affect and behavior.
    The paper adopts these frameworks in Sections 3 and 5 without deriving them from the hypothesis.
  • domain assumption Affective states are grounded in interoceptive signals (allostatic processes).
    The paper relies on Barrett's constructed emotion theory and Sennesh et al. to connect affect to body physiology.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Affective-Taxis Hypothesis for Alignment and Interpretability." pith.science (2026). https://pith.science/paper/LILBGQJD

@misc{pith2026250517024,
  author       = {Pith},
  title        = {Pith review of: An Affective-Taxis Hypothesis for Alignment and Interpretability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LILBGQJD}},
  note         = {Machine review of arXiv:2505.17024}
}
read the original abstract

AI alignment is a field of research that aims to develop methods to ensure that agents always behave in a manner aligned with (i.e. consistently with) the goals and values of their human operators, no matter their level of capability. This paper proposes an affectivist approach to the alignment problem, re-framing the concepts of goals and values in terms of affective taxis, and explaining the emergence of affective valence by appealing to recent work in evolutionary-developmental and computational neuroscience. We review the state of the art and, building on this work, we propose a computational model of affect based on taxis navigation. We discuss evidence in a tractable model organism that our model reflects aspects of biological taxis navigation. We conclude with a discussion of the role of affective taxis in AI alignment.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

86 extracted references · 63 canonical work pages

  1. [1]

    In: Advance s in Neural Informa- tion Processing Systems

    Abel, D., Dabney, W., Harutyunyan, A., Ho, M.K., Littman, M.L., Precup, D., Singh, S.: On the expressivity of markov reward. In: Advance s in Neural Informa- tion Processing Systems. p. 1–14 (2021), http://arxiv.org /abs/2111.00876, arXiv: 2111.00876

  2. [2]

    Current Opinion in Systems Biology 8, 81–89 (Apr 2018)

    Adler, M., Alon, U.: Fold-change detection in biological sys- tems. Current Opinion in Systems Biology 8, 81–89 (Apr 2018). https://doi.org/10.1016/j.coisb.2017.12.005 10 E. Sennesh and M. Ramstead

  3. [3]

    PLoS comput ational biology 10(8), e1003781 (2014)

    Adler, M., Mayo, A., Alon, U.: Logarithmic and power law in put-output relations in sensory systems with fold-change detection. PLoS comput ational biology 10(8), e1003781 (2014)

  4. [4]

    Advances in neural information processing systems 31 (2018)

    Armstrong, S., Mindermann, S.: Occam’s razor is insufficie nt to infer the prefer- ences of irrational agents. Advances in neural information processing systems 31 (2018)

  5. [5]

    Badman, R., Simmons-Edler, R., Berg, F., Lunger, J., Vast ola, J., Qian, W., Rajan, K.: Forageworld: RL agents in complex foraging arenas devel op internal maps for navigation and planning (2025)

  6. [6]

    arXiv pre print arXiv:2204.05862 (2022)

    Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasS arma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al.: Training a helpf ul and harmless assistant with reinforcement learning from human feedback. arXiv pre print arXiv:2204.05862 (2022)

  7. [7]

    M.: Learning the dynamics of realistic models of c

    Barbulescu, R., Mestre, G., Oliveira, A.L., Silveira, L. M.: Learning the dynamics of realistic models of c. elegans nervous system with recurr ent neural networks. Scientific reports 13(1), 467 (2023)

  8. [8]

    Social cognitive and aff ective neuroscience 12(1), 1–23 (2017)

    Barrett, L.F.: The theory of constructed emotion: an acti ve inference account of interoception and categorization. Social cognitive and aff ective neuroscience 12(1), 1–23 (2017). https://doi.org/10.1093/scan/nsw154

Show all 86 references
  1. [9]

    Advances in experimental social psychology 41, 167–218 (2009)

    Barrett, L.F., Bliss-Moreau, E.: Affect as a psychologica l primitive. Advances in experimental social psychology 41, 167–218 (2009)

  2. [10]

    Proceedings of the Royal Society B: Biological Scie nces 290(2002), 20230671 (Jul 2023)

    Barron, A.B., Halina, M., Klein, C.: Transitions in cogn itive evo- lution. Proceedings of the Royal Society B: Biological Scie nces 290(2002), 20230671 (Jul 2023). https://doi.org/10.1098/rs pb.2023.0671, https://royalsocietypublishing.org/doi/10.1098/rspb.2023.0671, publishe...

  3. [11]

    https://doi.org/10.485 50/arXiv.2202.08587, http://arxiv.org/abs/2202.08587

    Baydin, A.G., Pearlmutter, B.A., Syme, D., Wood, F., Tor r, P.: Gra- dients without backpropagation. https://doi.org/10.485 50/arXiv.2202.08587, http://arxiv.org/abs/2202.08587

  4. [12]

    Frontiers in Neuroanatomy 15, 693346 (2021)

    Bennett, M.S.: Five breakthroughs: a first approximatio n of brain evolution from early bilaterians to humans. Frontiers in Neuroanatomy 15, 693346 (2021)

  5. [13]

    Beshears, J., Choi, J.J., Laibson, D., Madrian, B.C.: Ho w are preferences revealed? Journal of public economics 92(8-9), 1787–1794 (2008)

  6. [14]

    https://doi.org/10.31219/osf.io/fe36n_v1, https://osf.io/fe36n_v1

    Byrnes, S.J.: Intro to brain-like agi safety (Mar 2025). https://doi.org/10.31219/osf.io/fe36n_v1, https://osf.io/fe36n_v1

  7. [15]

    In: Advances in Neural Information Processing Systems (202 1)

    Cao, H., Cohen, S.N., Szpruch, L.: Identifiability in inv erse reinforcement learning. In: Advances in Neural Information Processing Systems (202 1)

  8. [16]

    Chase, D.L., Koelle, M.R.: Biogenic amine neurotransmi tters in C. elegans. In: WormBook: the online review of C. elegans biology, pp. 1–15. WormBook (2007). https://doi.org/10.1895/wormbook.1.132.1, iSSN: 15518 507

  9. [17]

    elegans con nectome

    Chen, M., Feng, D., Su, H., Su, T., Wang, M.: Neural model g enerating klinotaxis behavior accompanied by a random walk based on c. elegans con nectome. Scientific Reports 12(1), 3043 (2022)

  10. [18]

    In: Proceedings o f the Annual Meeting of the Cognitive Science Society

    Chen, T., Houlihan, S.D., Chandra, K., Tenenbaum, J., Sa xe, R.: Intervening on emotions by planning over a theory of mind. In: Proceedings o f the Annual Meeting of the Cognitive Science Society. vol. 46 (2024)

  11. [19]

    Mathematics of Operations Research 10(4), 633–641 (1985)

    Chichilnisky, G.: Von neumann-morgenstern utilities a nd cardinal preferences. Mathematics of Operations Research 10(4), 633–641 (1985)

  12. [20]

    Attention, perception, & psychophysics 81, 2265–2287 (2019) An Affective-Taxis Hypothesis for Alignment and Interpreta bility 11

    Cisek, P.: Resynthesizing behavior through phylogenet ic refinement. Attention, perception, & psychophysics 81, 2265–2287 (2019) An Affective-Taxis Hypothesis for Alignment and Interpreta bility 11

  13. [21]

    Philosoph- ical Transactions of the Royal Society B: Biological Scienc es 377(1844), 20200522 (Dec 2021)

    Cisek, P.: Evolution of behavioural control from chorda tes to primates. Philosoph- ical Transactions of the Royal Society B: Biological Scienc es 377(1844), 20200522 (Dec 2021). https://doi.org/10.1098/rstb.2020.0522

  14. [22]

    Neural Networks 15(4-6), 603– 616 (Jun 2002)

    Daw, N.D., Kakade, S., Dayan, P.: Opponent interactions be- tween serotonin and dopamine. Neural Networks 15(4-6), 603– 616 (Jun 2002). https://doi.org/10.1016/S0893-6080(02) 00052-7, https://linkinghub.elsevier.com/retrieve/pii/S0893608002000527

  15. [23]

    In: Proceedings of the 41st International Conf erence on Machine Learning

    Du, Y., Kaelbling, L.: Compositional generative modeli ng: a single model is not all you need. In: Proceedings of the 41st International Conf erence on Machine Learning. vol. 235. Proceedings of Machine Learning Resear ch, Vienna, Austria (2024)

  16. [24]

    In: Advances in Neural Information Pro- cessing Systems

    Du, Y., Mordatch, I.: Implicit generation and modeling w ith en- ergy based models. In: Advances in Neural Information Pro- cessing Systems. vol. 32. Curran Associates, Inc. (2019), https://proceedings.neurips.cc/paper/2019/hash/378a063b8fdb1db941e34f4bde584c7d-Abstract.html

  17. [25]

    Nature human behaviour 5(7), 816–820 (2021)

    Dukes, D., Abrams, K., Adolphs, R., Ahmed, M.E., Beatty, A., Berridge, K.C., Broomhall, S., Brosch, T., Campos, J.J., Clay, Z., et al.: Th e rise of affectivism. Nature human behaviour 5(7), 816–820 (2021)

  18. [26]

    Neuro science & Biobehavioral Reviews 144, 104977 (2023)

    Emanuel, A., Eldar, E.: Emotions as computations. Neuro science & Biobehavioral Reviews 144, 104977 (2023)

  19. [27]

    Trends in Cogni- tive Sciences (2024)

    Feldman, M.J., Bliss-Moreau, E., Lindquist, K.A.: The n eu- robiology of interoception and affect. Trends in Cogni- tive Sciences (2024). https://doi.org/10.1016/j.tics.2 024.01.009, https://www.sciencedirect.com/science/article/pii/S1364661324000093

  20. [28]

    In: The Thirt y-eighth An- nual Conference on Neural Information Processing Systems ( 2024), https://openreview.net/forum?id=8KPyJm4gt5

    Foster, D.J., Block, A., Misra, D.: Is behavior cloning a ll you need? understanding horizon in imitation learning. In: The Thirt y-eighth An- nual Conference on Neural Information Processing Systems ( 2024), https://openreview.net/forum?id=8KPyJm4gt5

  21. [29]

    Fournier, L., Rivaud, S., Belilovsky, E., Eickenberg, M ., Oyallon, E.: Can for- ward gradient match backpropagation? In: Proceedings of th e 40th Interna- tional Conference on Machine Learning. p. 10249–10264. PML R (Jul 2023), https://proceedings.mlr.press/v202/fournier23a.html

  22. [30]

    Cognitive neuroscience 6(4), 187–214 (2015)

    Friston, K., Rigoli, F., Ognibene, D., Mathys, C., Fitzg erald, T., Pezzulo, G.: Active inference and epistemic value. Cognitive neuroscience 6(4), 187–214 (2015)

  23. [31]

    PLOS Computational Biology 8(1), e1002327 (Jan 2012)

    Friston, K.J., Shiner, T., FitzGerald, T., Galea, J.M., Adams, R., Brown, H., Dolan, R.J., Moran, R., Stephan, K.E., Bestmann, S.: Dopami ne, affordance and active inference. PLOS Computational Biology 8(1), e1002327 (Jan 2012). https://doi.org/10.1371/journal.pcbi.1002327

  24. [32]

    arXiv preprint arXiv:2409.11733 (2024)

    Gandhi, K., Lynch, Z., Fränken, J.P., Patterson, K., Wam bu, S., Gerstenberg, T., Ong, D.C., Goodman, N.D.: Human-like affective cognition in foundation models. arXiv preprint arXiv:2409.11733 (2024)

  25. [33]

    Communications Biology 5(1), 1354 (Dec 2022)

    Gündem, D., Potočnik, J., De Winter, F.L., El Kaddouri, A ., Stam, D., Peeters, R., Emsell, L., Sunaert, S., Van Oudenhove, L., Vandenbulck e, M., Feldman Bar- rett, L., Van den Stock, J.: The neurobiological basis of affe ct is consis- tent with psychological construction theo...

  26. [34]

    In: Lee, D., Sugiy ama, M., 12 E

    Hadfield-Menell, D., Russell, S.J., Abbeel, P., Dragan, A.: Coop- erative inverse reinforcement learning. In: Lee, D., Sugiy ama, M., 12 E. Sennesh and M. Ramstead Luxburg, U., Guyon, I., Garnett, R. (eds.) Advances in Neura l Infor- mation Processing Systems. vol. 29. Curran A...

  27. [35]

    , Naud, R.: A prospective code for value in the serotonin system

    Harkin, E.F., Grossman, C.D., Cohen, J.Y., Béïque, J.C. , Naud, R.: A prospective code for value in the serotonin system. Na- ture pp. 1–8 (Mar 2025). https://doi.org/10.1038/s41586- 025-08731-7, https://www.nature.com/articles/s41586-025-08731-7, publisher: Nature Publish- ing Group

  28. [36]

    Neural computation 33(2), 398–446 (2021)

    Hesp, C., Smith, R., Parr, T., Allen, M., Friston, K.J., R amstead, M.J.: Deeply felt affect: The emergence of valence in deep active inferenc e. Neural computation 33(2), 398–446 (2021)

  29. [37]

    Philosophical Transactions of the Royal Society A 381(2251), 20220047 (2023)

    Houlihan, S.D., Kleiman-Weiner, M., Hewitt, L.B., Tene nbaum, J.B., Saxe, R.: Emotion prediction as computation over a generative theory of mind. Philosophical Transactions of the Royal Society A 381(2251), 20220047 (2023)

  30. [38]

    PLoS ge- netics 14(3), e1007305 (2018)

    Hussey, R., Littlejohn, N.K., Witham, E., Vanstrum, E., Mesgarzadeh, J., Ratan- pal, H., Srinivasan, S.: Oxygen-sensing neurons reciproca lly regulate peripheral lipid metabolism via neuropeptide signaling in caenorhabd itis elegans. PLoS ge- netics 14(3), e1007305 (2018)

  31. [39]

    In: 35th International Conference on Machine Learning , ICML 2018

    Icarte, R.T., Klassen, T.Q., Valenzano, R., McIlraith, S.A.: Using reward ma- chines for high-level task specification and decomposition in reinforcement learn- ing. In: 35th International Conference on Machine Learning , ICML 2018. vol. 5, p. 3347–3358 (2018), citation Key: Icarte2018

  32. [40]

    Itskovits, E., Ruach, R., Kazakov, A., Zaslaver, A.: Con certed pulsatile and graded neural dynamics enables efficient chemotaxis in c. elegans. N ature communications 9(1), 2866 (2018)

  33. [41]

    Proceedings of the National Academy of Sciences 120(20), e2219341120 (May 2023)

    Ji, H., Fouad, A.D., Li, Z., Ruba, A., Fang-Yen, C.: A prop rioceptive feedback circuit drives caenorhabditis elegans locomotor adaptati on through dopamine sig- naling. Proceedings of the National Academy of Sciences 120(20), e2219341120 (May 2023). https://doi.org/10.1073/pn...

  34. [42]

    Neural networks 15(4-6), 535–547 (2002)

    Joel, D., Niv, Y., Ruppin, E.: Actor–critic models of the basal ganglia: New anatom- ical and computational perspectives. Neural networks 15(4-6), 535–547 (2002)

  35. [43]

    PLoS computational biology 9(6), e1003094 (2013)

    Joffily, M., Coricelli, G.: Emotional valence and the free -energy principle. PLoS computational biology 9(6), e1003094 (2013)

  36. [44]

    Iscience 24(7) (2021)

    Karin, O., Alon, U.: Temporal fluctuations in chemotaxis gain implement a simulated-tempering strategy for efficient navigation in co mplex environments. Iscience 24(7) (2021)

  37. [45]

    PLoS computational biology 18(7), e1010340 (2022)

    Karin, O., Alon, U.: The dopamine circuit as a reward-tax is navigation system. PLoS computational biology 18(7), e1010340 (2022)

  38. [46]

    Journa l of theoretical biology 30(2), 225–234 (1971)

    Keller, E.F., Segel, L.A.: Model for chemotaxis. Journa l of theoretical biology 30(2), 225–234 (1971)

  39. [47]

    Journal of theoretical biology 30(2), 235–248 (1971)

    Keller, E.F., Segel, L.A.: Traveling bands of chemotact ic bacteria: a theoretical analysis. Journal of theoretical biology 30(2), 235–248 (1971)

  40. [48]

    eLife 3, e04811 (dec 2014)

    Keramati, M., Gutkin, B.: Homeostatic reinforcement le arning for integrat- ing reward collection and physiological stability. eLife 3, e04811 (dec 2014). https://doi.org/10.7554/eLife.04811, https://doi.org/10.7554/eLife.04811

  41. [49]

    In: Proceedings of the 38th Internatio nal Conference on Ma- chine Learning

    Kim, K., Shiragur, K., Garg, S., Ermon, S.: Reward identi fication in inverse rein- forcement learning. In: Proceedings of the 38th Internatio nal Conference on Ma- chine Learning. vol. 139. Proceedings of Machine Learning R esearch (2021)

  42. [50]

    Journal of the Royal Society Interface 10(81), 20121001 (2013) An Affective-Taxis Hypothesis for Alignment and Interpreta bility 13

    Kojadinovic, M., Armitage, J.P., Tindall, M.J., Wadham s, G.H.: Response kinetics in the complex chemotaxis signalling pathway of rhodobacte r sphaeroides. Journal of the Royal Society Interface 10(81), 20121001 (2013) An Affective-Taxis Hypothesis for Alignment and Interpreta ...

  43. [51]

    In: 2016 American Control Conference (ACC)

    Lang, M., Sontag, E.: Scale-invariant systems realize n onlinear differential opera- tors. In: 2016 American Control Conference (ACC). pp. 6676– 6682. IEEE (2016)

  44. [52]

    elegans chemotaxis

    Larsch, J., Flavell, S.W., Liu, Q., Gordus, A., Albrecht , D.R., Bargmann, C.I.: A circuit for gradient climbing in c. elegans chemotaxis. Ce ll Reports 12(11), 1748–1760 (2015). https://doi.org/10.1016/j.celrep.20 15.08.032

  45. [53]

    Proceedings of the Nati onal Academy of Sci- ences 108(33), 13870–13875 (2011)

    Lazova, M.D., Ahmed, T., Bellomo, D., Stocker, R., Shimi zu, T.S.: Response rescaling in bacterial chemotaxis. Proceedings of the Nati onal Academy of Sci- ences 108(33), 13870–13875 (2011). https://doi.org/10.1073/pnas .1108608108, https://www.pnas.org/doi/abs/10.1073/pnas.1108608108

  46. [54]

    Nature Reviews Neuroscience p

    Lillicrap, T.P., Santoro, A., Marris, L., Akerman, C., H inton, G.E.: Back- propagation and the brain. Nature Reviews Neuroscience p. 1 –12 (2020). https://doi.org/10.1038/s41583-020-0277-3, citation K ey: Lillicrap2020

  47. [55]

    Lockery, S.R.: The computational worm: spatial orienta tion and its neuronal ba- sis in c. elegans. Current Opinion in Neurobiology 21(5), 782–790 (Oct 2011). https://doi.org/10.1016/j.conb.2011.06.009

  48. [56]

    In: Advances in Neural Information Pro- cessing Systems

    Ma, Y.A., Chen, T., Fox, E.: A complete recipe for stochas - tic gradient mcmc. In: Advances in Neural Information Pro- cessing Systems. vol. 28. Curran Associates, Inc. (2015), https://papers.nips.cc/paper/2015/hash/9a4400501febb2a95e79248486a5f6d3-Abstract.html

  49. [57]

    , Payne, A., Sanborn, S., Schroeder, K., Tavares, Z., Tolias, A.: NeuroAI for AI Sa fety (Nov 2024)

    Mineault, P., Zanichelli, N., Peng, J.Z., Arkhipov, A., Bingham, E., Jara- Ettinger, J., Mackevicius, E., Marblestone, A., Mattar, M. , Payne, A., Sanborn, S., Schroeder, K., Tavares, Z., Tolias, A.: NeuroAI for AI Sa fety (Nov 2024). https://doi.org/10.48550/arXiv.2411.18526,...

  50. [58]

    elegans locomotory behavior

    Moy, K., Li, W., Tran, H.P., Simonis, V., Story, E., Brand on, C., Furst, J., Raicu, D., Kim, H.: Computational methods for tracking, quantitat ive assessment, and visualization of c. elegans locomotory behavior. PLOS ONE 10(12), e0145870 (Dec 2015). https://doi.org/10.1371/jo...

  51. [59]

    Physical Review Research 4(1), 013120 (2022)

    Nakamura, K., Kobayashi, T.J.: Optimal sensing and cont rol of run-and-tumble chemotaxis. Physical Review Research 4(1), 013120 (2022)

  52. [60]

    Journal of Mathematical Psychology 53(3), 139–154 (2009)

    Niv, Y.: Reinforcement learning in the brain. Journal of Mathematical Psychology 53(3), 139–154 (2009)

  53. [61]

    : Dopamine signaling is essential for precise rates of locomotion by c

    Omura, D.T., Clark, D.A., Samuel, A.D.T., Horvitz, H.R. : Dopamine signaling is essential for precise rates of locomotion by c. elegans. PLO S ONE 7(6), e38649 (Jun 2012). https://doi.org/10.1371/journal.pone.0038 649

  54. [62]

    Neuroscience and Biobehavioral Reviews 63, 124–142 (2016)

    Pool, E., Sennwald, V., Delplanque, S., Brosch, T., Sand er, D.: Measuring wanting and liking from animals to humans: A sys- tematic review. Neuroscience and Biobehavioral Reviews 63, 124–142 (2016). https://doi.org/10.1016/j.neubiorev.2 016.01.006, http://dx.doi.org/10.1016/j...

  55. [63]

    The philosophical review 95(2), 163–207 (1986)

    Railton, P.: Moral realism. The philosophical review 95(2), 163–207 (1986)

  56. [64]

    Emotion R eview 9(4), 335–342 (2017)

    Railton, P.: At the core of our capacity to act for a reason : The affective system and evaluative model-based learning and control. Emotion R eview 9(4), 335–342 (2017)

  57. [65]

    Ramstead, M.: Ai alignment and theory of mind (Mar 2025), https://www.noumenal.ai/post/ai-alignment-and-theory-of-mind

  58. [66]

    , Pokala, N., et al.: A stochastic neuronal model predicts random search behavior s at multiple spatial scales in c

    Roberts, W.M., Augustine, S.B., Lawton, K.J., Lindsay, T.H., Thiele, T.R., Izquierdo, E.J., Faumont, S., Lindsay, R.A., Britton, M.C. , Pokala, N., et al.: A stochastic neuronal model predicts random search behavior s at multiple spatial scales in c. elegans. Elife 5, e12572 (...

  59. [67]

    , Bianciardi, M.: Decon- structing arousal into wakeful, autonomic and affective var ieties

    Satpute, A.B., Kragel, P.A., Barrett, L.F., Wager, T.D. , Bianciardi, M.: Decon- structing arousal into wakeful, autonomic and affective var ieties. Neuroscience Let- ters 693, 19–28 (Feb 2019). https://doi.org/10.1016/j.neulet.20 18.01.042

  60. [68]

    , Barrett, L.F., Quigley, K.S.: Interoception as modeling, allostasi s as con- trol

    Sennesh, E., Theriault, J., Brooks, D., Meent, J.W.v.d. , Barrett, L.F., Quigley, K.S.: Interoception as modeling, allostasi s as con- trol. Biological Psychology 167(December 2021), 108242 (2021). https://doi.org/10.1016/j.biopsycho.2021.108242, citation Key: Sennesh2021

  61. [69]

    Mit Press (2020)

    Shadmehr, R., Ahmed, A.A.: Vigor: Neuroeconomics of mov ement control. Mit Press (2020)

  62. [70]

    Trends in Cognitive Sciences (2024)

    Shenhav, A.: The affective gradient hypothesis: An affect -centered account of mo- tivated behavior. Trends in Cognitive Sciences (2024)

  63. [71]

    GoogleAPI preprint https://goo.gle/3EiRKIH (2025)

    Silver, D., Sutton, R.S.: Welcome to the era of experienc e. GoogleAPI preprint https://goo.gle/3EiRKIH (2025)

  64. [72]

    arXiv (Feb 2021), http://arxiv.org/abs/2101.03288, arXiv:2101.03288 [cs , stat]

    Song, Y., Kingma, D.P.: How to train your energy-based mo dels. arXiv (Feb 2021), http://arxiv.org/abs/2101.03288, arXiv:2101.03288 [cs , stat]

  65. [73]

    Nature Neuroscience p

    Song, Y., Millidge, B., Salvatori, T., Lukasiewicz, T., Xu, Z., Bogacz, R.: Inferring neural activity before plasticity as a founda tion for learn- ing beyond backpropagation. Nature Neuroscience p. 1–11 (J an 2024). https://doi.org/10.1038/s41593-023-01514-1

  66. [74]

    Physiology and Behavior 106(1), 5–15 (2012)

    Sterling, P.: Allostasis: A model of predictive regulat ion. Physiology and Behavior 106(1), 5–15 (2012). https://doi.org/10.1016/j.physbeh.20 11.06.004, http://dx.doi.org/10.1016/j.physbeh.2011.06.004

  67. [75]

    , Pierré, A., Schul- hoff, S., Tai, J.J., Tan, H., Younis, O.G.: Gymnasium: A stand ard interface for reinforcement learning environments (2024), https://arx iv.org/abs/2407.17032

    Towers, M., Kwiatkowski, A., Terry, J., Balis, J.U., Col a, G.D., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., Perez-Vicente, R. , Pierré, A., Schul- hoff, S., Tai, J.J., Tan, H., Younis, O.G.: Gymnasium: A stand ard interface for reinforcement learning environm...

  68. [76]

    Software Impacts 6, 100022 (Nov 2020)

    Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Boh ez, S., Merel, J., Erez, T., Lillicrap, T., Heess, N., Tassa, Y.: dm_contro l: Software and tasks for continuous control. Software Impacts 6, 100022 (Nov 2020). https://doi.org/10.1016/j.simpa.2020.100022

  69. [77]

    Proceedings of the Nation al Academy of Sciences 108(42), 17504–17509 (Oct 2011)

    Vidal-Gadea, A., Topper, S., Young, L., Crisp, A., Kress in, L., Elbel, E., Maples, T., Brauner, M., Erbguth, K., Axelrod, A., Gottschalk, A., S iegel, D., Pierce- Shimomura, J.T.: Caenorhabditis elegans selects distinct crawling and swimming gaits via dopamine and serotonin. ...

  70. [78]

    https://doi.org/ 10.31234/osf.io/be6nv , https://osf.io/be6nv

    Weber, L., Yee, D., Small, D., Petzschner, F.H.: Rethink ing reinforcement learn- ing: The interoceptive origin of reward. https://doi.org/ 10.31234/osf.io/be6nv , https://osf.io/be6nv

  71. [79]

    Trends in cognitive sciences 27(3), 246–257 (2023)

    Westlin, C., Theriault, J.E., Katsumi, Y., Nieto-Casta non, A., Kucyi, A., Ruf, S.F., Brown, S.M., Pavel, M., Erdogmus, D., Brooks, D.H., et al.: I mproving the study of brain-behavior relationships by revisiting basic assum ptions. Trends in cognitive sciences 27(3), 246–257 (2023)

  72. [80]

    Journal of Personality and Social Psychology 41(1), 56 (1981)

    White, G.L., Fishbein, S., Rutsein, J.: Passionate love and the misattribution of arousal. Journal of Personality and Social Psychology 41(1), 56 (1981)

  73. [81]

    elegans body cavity neurons are homeostatic sensors that integrate fluctuations in oxygen availability and internal nutrient reserves

    Witham, E., Comunian, C., Ratanpal, H., Skora, S., Zimme r, M., Srinivasan, S.: C. elegans body cavity neurons are homeostatic sensors that integrate fluctuations in oxygen availability and internal nutrient reserves. Cel l reports 14(7), 1641–1654 (2016)

  74. [82]

    Singul arity Institute for Artificial Intelligence (2004) An Affective-Taxis Hypothesis for Alignment and Interpreta bility 15

    Yudkowsky, E.: Coherent extrapolated volition. Singul arity Institute for Artificial Intelligence (2004) An Affective-Taxis Hypothesis for Alignment and Interpreta bility 15

  75. [83]

    In: NeurIPS 2024 Workshop on Scientific Methods for Understanding Deep L earning (2024), https://openreview.net/forum?id=vgXUoCrHmp

    Zhao, B., Okawa, M., Bigelow, E.J., Yu, R., Ullman, T., Ta naka, H.: Emergence of hierarchical emotion representations in large language models. In: NeurIPS 2024 Workshop on Scientific Methods for Understanding Deep L earning (2024), https://openreview.net/forum?id=vgXUoCrHmp

  76. [84]

    , Ma, L., Huang, T.: An integrative data-driven model simulating c

    Zhao, M., Wang, N., Jiang, X., Ma, X., Ma, H., He, G., Du, K. , Ma, L., Huang, T.: An integrative data-driven model simulating c. elegans brain, body and envi- ronment interactions. Nature Computational Science 4(12), 978–990 (2024)

  77. [85]

    Philosophical Stud- ies (Nov 2024)

    Zhi-Xuan, T., Carroll, M., Franklin, M., Ashton, H.: Be- yond Preferences in AI Alignment. Philosophical Stud- ies (Nov 2024). https://doi.org/10.1007/s11098-024-022 49-w , https://link.springer.com/10.1007/s11098-024-02249-w

  78. [86]

    In: Advances in Neural Information Processing Systems

    Zhu, J., Sanborn, A., Chater, N.: Mental sampling in mult imodal representations. In: Advances in Neural Information Processing Systems. vol . 31. Curran Associates, Inc., Montreal, Quebec, Canada (2018)

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.