{"id":"7d9267bd-f117-48be-bb86-6c43a0063bac","arxiv_id":"2505.17024","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper hypothesizes that affective valence evolved from taxis navigation and that reward in the brain is the directional derivative of an interoceptive energy function, with implications for AI alignment.","lead":"This paper proposes that the brain's reward signal is the directional derivative of an interoceptive energy landscape, an idea the authors call the affective-taxis hypothesis. It applies this idea to AI alignment, arguing that aligned agents must model human affective states grounded in body physiology.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central identity—reward as directional derivative of an interoceptive energy function—is asserted without the direct FCD-in-interoceptive-neurons evidence the paper itself identifies as needed.","rationale":"The reader's weakest assumption was the evolutionary story about neural tube folding reorienting taxis inward. That is a real gap, but it is not the most load-bearing point: the paper's strongest claim is an identity about the brain's reward function, and the evolutionary narrative is only one motivation for that identity. The more decisive weakness is the missing evidence that interoceptive neurons implement FCD and that such signals function as reward or valence. The paper itself flags this missing support in Section 4 ('direct experimental evidence remains limited') and in the Limitations section, where it concedes that human affect requires far more than the C. elegans model provides. Because the central claim is a hypothesis and the paper is transparent about its evidential status, the reader's CONDITIONAL verdict is appropriate: the proposal is valuable and worth testing, but the strongest claim is not yet established. My concern does not move the verdict; it sharpens the reason the verdict is conditional rather than accept. Agreement with the reader is partial because I identify the missing FCD/interoceptive-reward evidence as more load-bearing than the evolutionary narrative, though both are related aspects of the same evidentiary gap.","tokens_in":13643,"tokens_out":4570,"duration_ms":53809,"concrete_test":"Analyze existing calcium-imaging data from C. elegans interoceptive neurons (e.g., URX oxygen-sensing neurons, Witham et al. 2016; Larsch et al. 2015) and test whether their responses to step changes in O2 are fold-change (scale-invariant) and whether the response tracks d/dt log gamma(z(t)) rather than absolute concentration. Then test whether optogenetic silencing of these neurons abolishes gradient-following in a controlled oxygen gradient. If the responses are not FCD, or if silencing does not disrupt taxis, the FCD-based reward identity lacks its direct empirical anchor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim, stated in Related Work, is that 'the reward function in the brain... consists of the (negative) directional derivative, in the present direction of movement, of a time-varying, interoceptive energy function.' This identity is the foundation for the alignment proposal: if reward is not this directional derivative, then the proposed inductive bias and the claim that aligned AI must represent affective states lose their grounding. The only direct empirical anchor offered is fold-change detection (FCD) in C. elegans sensory neurons (Section 4). That evidence shows that some chemosensory neurons signal a relative change in attractant concentration, approximately d/dt log gamma(z(t)). Two gaps separate this from the central claim. First, FCD is a sensory transduction property, not a demonstrated reward or valence signal; the paper does not show that the FCD signal is used as a reward by the organism. Second, the hypothesized link to interoception is explicitly unverified: the paper states 'Direct evidence that interoceptive or similar sensory neurons... implement FCD would support the affective-taxis hypothesis,' and acknowledges 'direct experimental evidence remains limited.' The paper's own Limitations section further concedes that human affect requires temporal associations, situated conceptualizations, and theory of mind, none of which the C. elegans model captures. Thus, the leap from a chemotaxis control signal in a worm to the human brain's reward function is a stipulative reinterpretation rather than an empirically supported identity. The evolutionary narrative about neural tube folding is also speculative, but it is secondary: even if that narrative were accepted, the reward-as-directional-derivative claim would still lack direct evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper proposes an 'affective-taxis hypothesis' according to which affective valence in humans and other bilaterians evolved from the internalization of taxis navigation. The authors formalize the core idea by identifying reward signals with the directional derivative of an interoceptive energy function: r(t) = ∇ log γ(z(t); β(t)) · dz/dt. They propose a POMDP model for taxis behavior in which rewards are exactly these directional derivatives, and they argue that fold-change detection (FCD) in C. elegans sensory neurons provides biological evidence for this quantity. The paper then argues that AI alignment requires agents to represent and model human affective states, and that the affective-taxis inductive bias could help overcome identifiability problems in inverse RL. The manuscript is primarily a conceptual and review contribution, with no experiments or implemented simulations; it explicitly lists its limitations and proposes directions for future work.","tokens_in":13894,"tokens_out":4808,"duration_ms":48201,"significance":"If the affective-taxis hypothesis is correct, it would give the brain's reward signal a concrete, measurable form (a directional derivative of an allostatic energy function), which could directly inform the design of interpretable and aligned AI systems by grounding reward in interoceptive physiology. The paper's strengths include a clearly articulated hypothesis, a simple mathematical formalization, a tractable model organism, and an explicit falsifiable prediction (that interoceptive or similar sensory neurons implement fold-change detection). The authors are also commendably transparent about the current lack of direct evidence and about the gap between C. elegans taxis and human affective experience. However, as the paper itself acknowledges, the central identity between reward and directional derivative remains a conjecture: the cited FCD evidence concerns chemosensory transduction, not reward or valence, and the evolutionary narrative connecting neural tube folding to internal taxis is speculative. The contribution is therefore a promising research program rather than an established result.","major_comments":[{"comment":"The central claim that the brain's reward function consists of the directional derivative of an interoceptive energy function is not supported by the evidence presented. The fold-change detection (FCD) results concern chemosensory neurons responding to environmental attractants, not reward or valence signals; the manuscript itself states that 'Direct evidence that interoceptive or similar sensory neurons... implement FCD would support the affective-taxis hypothesis.' This missing link is load-bearing because the alignment proposal depends on the assumption that the measured directional derivative is the actual reward/valence signal. The authors should explicitly separate the mathematical identity r(t) = d log γ/dt from the empirical conjecture that this quantity constitutes valence, and should propose a concrete experimental test (e.g., optogenetic manipulation of FCD neurons coupled with preference or approach/avoidance assays) that could validate the valence interpretation.","section":"Section 4 (Model Organism)"},{"comment":"The evolutionary narrative that 'the folding of the neural plate into a tube inside the organism reoriented the combined apical and blastoporal nervous systems towards navigating an internal taxis landscape' is presented without direct evidence. This step is crucial for moving from external taxis navigation to internal affective valence, yet it is asserted rather than argued from comparative data. The paper should either marshal empirical or phylogenetic evidence supporting this transition (e.g., from chordate neuroanatomy) or explicitly label this as a speculative component of the hypothesis and discuss how it could be tested or disconfirmed.","section":"Section 2 (Across the Affective Landscape)"},{"comment":"The statement that 'the reward function in the brain... consists of the (negative) directional derivative, in the present direction of movement, of a time-varying, interoceptive energy function' is written as an established fact, but it is the paper's own hypothesis and is not proven by the cited literature. This overstatement could mislead readers about the epistemic status of the claim. I recommend rephrasing to 'we hypothesize' or 'our proposal implies' and providing a clear account of how the hypothesis would be falsified.","section":"Related Work (Section 5)"}],"minor_comments":[{"comment":"The notation 'r(t) :≈ d/dt log γ(z(t); β(t))' uses a nonstandard symbol '≈' to introduce a definition; the equality r(t) = d log γ/dt is a mathematical identity when v = dz/dt, and the approximation should be stated separately, e.g., with explicit assumptions about the FCD transduction.","section":"Section 4 (Model Organism)"},{"comment":"The organism name should be consistently capitalized as 'C. elegans' rather than 'c. elegans' to conform to standard biological nomenclature.","section":"Throughout"},{"comment":"In the provided text, the title contains a line-break artifact 'Aﬀective-T axis Hypothesis'; ensure the camera-ready version does not insert a space in 'Taxis'.","section":"Title and headers"},{"comment":"In Definition 1, the POMDP tuple uses 'pS' and 'pO'; while acceptable, the notation would be clearer with standard symbols such as 'T' and 'O' or with explicit subscripts, and the definition of the observation model as the gradient ∇ log γ should be connected to the FCD discussion in Section 4.","section":"Section 3 (Computational Modeling Progress)"},{"comment":"Several references are to unpublished preprints or online posts (e.g., [14], [65]); if the journal permits, please add preprint identifiers or DOI numbers to improve verifiability.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a thought-provoking position piece that could generate useful discussion, but its central empirical claim is acknowledged by the authors to lack direct support. The review process should focus on whether the authors can strengthen the separation of mathematical identity, empirical hypothesis, and speculative narrative, and on whether they can articulate a more concrete experimental program for testing the valence interpretation. I see no reason to doubt the authors' good faith, and the manuscript's self-critical limitations section is a positive sign."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a position paper, and the weight rests on a hypothesis: that reward and affective valence are the directional derivative of an interoceptive energy function. The authors state it as a hypothesis, and they are honest throughout that direct evidence is missing. The key gap is real: fold-change detection in C. elegans sensory neurons shows some neurons compute d/dt log γ, but nothing yet shows that signal is used as a reward, and no interoceptive neuron has been shown to implement FCD. The paper names that exact gap in its Limitations and in Section 4. So the central claim is not a demonstrated result; it is a research bet.\n\nWhat is genuinely new is the synthesis. I have seen the affective gradient hypothesis and active inference treated separately, but this framing—taxis navigation as the evolutionary origin of valence, reward as a directional derivative, FCD as the biological mechanism—is a fresh combination. The prose is clear, the model organism choice is justified, and the authors do real work connecting their proposal to the existing literature. The C. elegans dopamine/serotonin discussion is a useful example of how they'd mechanistically test the story. I also give them credit for flagging what they have not done: the Limitations paragraph admits the model organism cannot do temporal association, situated conceptualization, or theory of mind, all needed for human affect.\n\nThe soft spots are proportionate. The biggest is the evidence gap I already mentioned. Second, the evolutionary narrative about neural tube folding reorienting taxis inward is speculative and hard to test, though it does not carry the whole argument. Third, the alignment payoff is described rather than demonstrated; the paper says an aligned AI 'must' represent affective states, but that is an inference from the affectivist premise, not a proved theorem. If you need a new empirical result, this is not that paper. If you want a clearly-argued hypothesis pointing to a testable model—the POMDP with reward equal to ∇logγ·v is concrete and implementable—this deserves a serious referee.\n\nMy verdict: the math is correct but trivial, the biology is suggestive but incomplete, and the value is in the framing. I would send it to peer review, because interdisciplinary position papers like this benefit from expert scrutiny. I would not cite it as evidence, but I would cite it as a proposal.\n\nBest,","headline":"The reward-as-directional-derivative identity is a promising but unproven hypothesis; the paper's honest limitations are its best feature.","tokens_in":14426,"tokens_out":2845,"would_cite":false,"duration_ms":26565,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that affective valence is taxis navigation through an internal interoceptive landscape, so aligned AI must represent that landscape.","keywords":["AI alignment","affective valence","taxis navigation","interoception","fold-change detection","reward function","energy-based models","C. elegans"],"falsifier":"Take a preparation in which taxis-sensing neurons are recorded while the animal moves through a controlled attractant field; if the neural response is not proportional to $d\\log \\gamma(z(t))/dt$, encoding absolute concentration rather than fold-change, the proposed identity between reward and the directional derivative fails. A complementary check would test whether dopamine release in a vertebrate tracks the instantaneous gradient slope in the current movement direction rather than a temporally cached value prediction.","tokens_in":13426,"feed_emoji":"🧭","tokens_out":8623,"duration_ms":80468,"temperature":0.7,"pith_summary":"This paper argues that the AI alignment problem is at root a problem about affect: human goals and values are not a scalar utility but the felt gradient of an internal taxis landscape, and an aligned AI must be able to model that landscape. The authors propose that reward in the brain is the directional derivative of a time-varying interoceptive energy function along the current direction of movement, making valence a measurable slope rather than an opaque number. If true, this would give AI alignment a concrete inductive bias to overcome the mathematical result that reward functions cannot be inferred from behavior alone, since behavior conflates beliefs and values. The paper grounds the claim in evolutionary-developmental neuroscience and in a computational model tested against C. elegans, a model organism whose taxis behavior lacks temporal associative learning and so exposes the gradient signal directly.","feed_headline":"Brain reward may be a directional gradient, not a utility score","feed_subtitle":"Valence would be a measurable slope of an inner landscape, giving AI alignment a concrete target: human interoception.","key_machinery":"The load-bearing object is the directional derivative of a log-density, $\\nabla_z \\log \\gamma(z;\\beta)\\cdot v$, where $\\gamma$ is the spatial density of attractants or, in the internalized case, the density of an allostatic energy function, and $v$ is the current velocity. The paper's central move is to equate this quantity with reward: instead of a scalar utility assigned to states, the agent reads the slope of its affective landscape in the direction it is moving. Fold-change detection is the neural mechanism said to implement this readout, since it computes the time-derivative of the log input, and energy-based models are proposed as the compositional representation that lets the same gradient logic pass from physical space to abstract interoceptive space. The identity turns taxis behavior into a Langevin-style gradient-biased random walk, so that exploration and exploitation are balanced by the same quantity that carries valence.","core_discovery":"The paper's central claim is that the brain's reward function consists of the negative directional derivative, in the present direction of movement, of a time-varying interoceptive energy function. Affect, on this view, is not a reaction to rewards but the navigation of an internal landscape: positive and negative valence are the signs of moving up or down gradients in a space of viscerosensory physiological indicators. The authors make this concrete by modeling reward as $R(s,a)=\\nabla_z \\log \\gamma(z;\\beta)\\cdot v$ in a POMDP whose state includes the animal's location, orientation, and the spatial density of attractants, and they identify fold-change detection in C. elegans sensory neurons as the biological implementation of the log-gradient readout. They further argue that associative reinforcement learning evolved later to estimate the long-run reward rate over this taxis landscape, so present-oriented taxis can be studied in organisms that lack temporal associations. The consequence they draw for AI alignment is that agents must represent human affective states and the interoceptive facts grounding them, not merely observe choices or ratings.","pith_inferences":["Editorial inference: If reward really is a directional derivative of an energy field, inverse reinforcement learning could be reframed as estimating that field's gradient from trajectories, which may sidestep some non-identifiability results that assume a scalar reward.","Editorial inference: Interpretability would gain a visual meaning: an agent's values would be the geometry of its affective landscape, and aligning two agents would mean matching the shapes of their energy functions, not just their outputs.","Editorial inference: A direct human test would measure whether interoceptive or dopaminergic signals track the instantaneous directional derivative of an allostatic prediction error during movement, not just reward prediction error; such a signal would be the predicted valence readout."],"forward_implications":["If reward is a directional derivative of an internal energy function, then reward functions are no longer unidentifiable in principle: the relevant quantity is the gradient field of an interoceptive landscape, not a free-standing scalar.","An aligned AI would need to infer and represent the human operator's time-varying interoceptive energy function, not just observe choices or ratings, because choices conflate beliefs and values.","The computational model can be tested in C. elegans, where taxis behavior is driven by present stimuli and no temporal associative learning confounds the reward signal.","Gradient-following by directional derivatives offers a biologically plausible alternative to backpropagation-based reward learning, since a single neuron can measure the derivative along its motion."],"supporting_citations":[{"why":"Supplies the affective gradient hypothesis that gradients over affective hills and valleys point the way for motivated behavior.","marker":"[70]"},{"why":"Provides the evolutionary account that folding the neural tube turned external taxis toward navigating an internal interoceptive landscape.","marker":"[21]"},{"why":"Supports the reward-taxis reading by explaining aspects of dopamine-driven reward seeking in mice as taxis navigation.","marker":"[45]"},{"why":"Defines fold-change detection as the mechanism by which sensory neurons compute the log-derivative of attractant density.","marker":"[2]"},{"why":"Supplies the run-and-tumble chemotaxis model that grounds the gradient-biased random walk formulation.","marker":"[46]"},{"why":"Grounds the interoceptive energy function in allostatic control, linking bodily physiology to the model's internal landscape.","marker":"[68]"}],"fun_headline_variants":["Reward is a slope, not a utility score","Affect as taxis: navigating inner gradients for alignment","C. elegans exposes reward as a gradient readout","Alignment needs interoceptive gradients, not utility"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's evolutionary grounding, that folding of the neural tube reoriented external taxis navigation toward an internal interoceptive landscape, is asserted without direct evidence, and the account of human affect depends on it.","fun_headline_variants_meta":{"raw":{"variants":["Reward is a slope, not a utility score","Affect as taxis: navigating inner gradients for alignment","C. elegans exposes reward as a gradient readout","Alignment needs interoceptive gradients, not utility"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000615,"raw_usage":{"total_tokens":2831,"prompt_tokens":890,"completion_tokens":1941,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":1887}},"tokens_in":506,"tokens_out":1941,"duration_ms":15089,"temperature":1.0,"reasoning_tokens":1887,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:12:09.858437+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a preparation in which taxis-sensing neurons are recorded while the animal moves through a controlled attractant field; if the neural response is not proportional to $d\\log \\gamma(z(t))/dt$, encoding absolute concentration rather than fold-change, the proposed identity between reward and the directional derivative fails. A complementary check would test whether dopamine release in a vertebrate tracks the instantaneous gradient slope in the current movement direction rather than a temporally cached value prediction.","supporting_citations":[{"cited_title":"Trends in Cognitive Sciences (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies the affective gradient hypothesis that gradients over affective hills and valleys point the way for motivated behavior."},{"cited_title":"PLoS computational biology 18(7), e1010340 (2022)","cited_arxiv_id":null,"evidence_quote":"Supports the reward-taxis reading by explaining aspects of dopamine-driven reward seeking in mice as taxis navigation."},{"cited_title":"Current Opinion in Systems Biology 8, 81–89 (Apr 2018)","cited_arxiv_id":null,"evidence_quote":"Defines fold-change detection as the mechanism by which sensory neurons compute the log-derivative of attractant density."},{"cited_title":"Journa l of theoretical biology 30(2), 225–234 (1971)","cited_arxiv_id":null,"evidence_quote":"Supplies the run-and-tumble chemotaxis model that grounds the gradient-biased random walk formulation."},{"cited_title":", Barrett, L.F., Quigley, K.S.: Interoception as modeling, allostasi s as con- trol","cited_arxiv_id":null,"evidence_quote":"Grounds the interoceptive energy function in allostatic control, linking bodily physiology to the model's internal landscape."}],"review_version":1}